September 8, 2026

Don’t Lose Old Habits Because You Have New Tools

Coding agents change how fast you ship, but they don’t change the discipline of software.

Don’t Lose Old Habits Because You Have New Tools

What was discussed

We sat down with Priya, Head of Engineering at ASAPP, to talk about what carries over when you move from deterministic systems to probabilistic ones.

Priya spent years on risk and trading platforms in investment banking before joining ASAPP, where she built the platform team and now leads engineering. ASAPP serves enterprise customers in regulated industries, so shipping fast and staying trustworthy are both non-negotiable.

The conversation covers:

  • Why correctness is domain-specific, and how teams build proxy metrics when the golden metric is too noisy to act on
  • Why platform engineering is still the highest leverage team in an organization, and how to make that work visible at the leadership level
  • The ROI problem with agent spend: the investment is easy to measure, the return is not, so most teams only optimize cost
  • Managing velocity without losing system understanding, and why Priya tracks zero customer-reported incidents instead of zero incidents
  • Securing agents like software, managing them like employees, and budgeting them like productive capacity
  • Where autonomy should expand next, and what has to be in place before it does

If you’re figuring out how to maintain engineering discipline as output increases, check it out.

Watch the full session below.

Transcript

Shahram (00:00)
Hello everybody. I’m your host Shahram and I’m joined with my co-founder Willem and very, very special guest, Priya. I’ll share a funny story, versus doing like the regular intro. The way Priya and Willem and I met was actually at a dinner we had in New York. And she had a prior engagement, so I think she came in on a little bit late. So we’d already gotten into this big discussion, and she sits down. I think within the first five minutes she’s already peppering us with questions like, What how does AI work? And like, what’s this and what’s that? And it was great. It was like so refreshing.

And so to me, she’s a professional who’s gone from building high scale systems at banking, from Bank of America to Goldman Sachs. What I think really stands out is that I think she jumped on the AI wave way before all of us, including Willem and I did. ASAPP has been building AI for enterprise way longer than, you know, before even ChatGPT was released. So I’m very excited to have her. If I had one word to describe her, I would say thoughtful. So I’m really looking forward to this conversation. Welcome, Priya.

Priya (00:57)
Thank you, Shahram. That was one of the most thoughtful compliment I have got. So I think I can use that word right back for you. And yeah, for me also the moment we’ve met was really great. So I’m happy to be here with both of you. Just the discussions that we’ve had along the way, they have been great. So looking forward to this one as well.

Shahram (01:16)
Wonderful. Let’s just start with that intro, right? You’ve been building high scale systems for a long time. So I’m always curious. You went from clearly like a very regulated space, high pressure, right? There’s a whole bunch of different risk. But the way I describe those kind of systems, and correct me if I’m wrong, you’re almost trying to be as deterministic as possible. And so that’s a very clear way of building things. But then you move to ASAPP and you’re taking another sort of a leadership engineering role. And AI is probabilistic at best, right? So how does that transition go? What was helpful from the banking world and what did you have to unlearn to do well in the AI space?

Priya (01:53)
I actually think that background in the banking is serving me as I am serving my enterprise customers, and I’ll get to it afterwards. But let me call out the distinctions, right? So when I worked in investment banking, I was working on the risk and trading platforms. And what that means is you’re trying to take a portfolio of, I don’t know, interest rate products or credit products, et cetera. And you run some Monte Carlo simulations on it, you have some stochastic models in the back, and it spits out some numbers that someone’s looking at, and you understand your risk exposure, right? And that was that. And you could run it daily, you could run it at a 15-minute chunk or real time, depending on the size of your input and the technology that was available to you at that time, right?

When I entered the AI space at that time, actually we did have for most part rule-based automation, and so the struggle was how do we make this rule-based system still act human-like, so how do I take a free text input? How do I make sense of it? Et cetera. And then of course, like generative AI came into play and automation was put back on the table, which was great. Though to your point, when it comes to building the system, like it’s very probabilistic. So your notion of what’s correct is kind of changing in terms of building the product itself.

But let’s think about outcome here. So for example, ASAPP builds automation products for customer service. A very well-understood metric in customer service is containment. Now, what is containment? I want the conversation to be fully serviced by automation instead of a human agent, right? And when I say fully serviced, what does that mean? 100%, right? So that’s the number that we are chasing. Now there are many ways of getting to that 100%, and in our world, that’s the conversation design. So the path to correctness are many, but the common understanding of correctness is decently well-defined or well understood.

So I think where I had to shift my focus when it came to building this product is how do we make sure that we are optimizing for those various paths to get to that correctness? And that’s sort of like the creativity and art and science sort of question, right? So that’s the main difference, I would say there was one way of reaching that value before. Now there are many ways of reaching that correctness.

Shahram (04:12)
And correctness is like depending on the domain. And you know, Willem and I can talk all day about correctness and like how hard it is and it depends on each domain. How do you define it? In my layman’s view, it’s almost like it’s correct or it’s contained if there was no human escalation requested. Is that the way or like how do you how do you think about it?

Priya (04:30)
I think we think of it that way as well. It’s a well-defined or well understood concept in the industry, but with some caveats. So for example, I would say that within a contained conversation, you have different goals, right? A person might have really wanted some extension on their credit line. Now you could contain the conversation by declining that credit line to them, right? But does that satisfy the customer? Probably not. So the CSAT comes into play as well, right? So it’s one metric to look at, but in terms of servicing your customers, you kind of look at like a few of them. I was simplifying.

Shahram (05:07)
Yeah, yeah, that makes sense. I mean we I don’t know, Willem, maybe, you know, we can talk a little bit about how you’ve been looking at correctness, right? Like again, like to Priya’s point, each domain, we all say AI, AI, AI, but actually it’s quite domain specific in enterprise. What works for one domain does not work for the other, right?

Willem (05:24)
Yeah, I’m kind of curious, Priya just to give it a little bit more color. In our space, defining correctness is if you’re diagnosing a production failure, obviously a root cause, even that can be subjective. Like how far upstream do you go? If it’s a natural language response, how detailed do you have to be? So even that required a lot of debate internally. Correct me if I’m wrong, but you were saying earlier it seems relatively standardized in the industries that you were in both conversational AI as well as prior or not?

Priya (05:49)
So, in terms of what the customer service business users will care about, the buyers, the business buyers will care about is certain numbers, right? They care about understanding whether I can contain that conversation. They care about their CSAT, right? So these are the numbers. Of course, it’s a threshold for CSAT, but these are the numbers that are well understood. What I was trying to say is that depending on how I design my conversation. I have multiple ways of getting to that number that brings the ROI that my users care about. Yeah.

Willem (06:22)
Right. Yeah. So the challenge in our space is define one of the metrics and then build the system to target those and then ideally the buyer or the user understands those things you’ve targeted. But I think in your space it’s already understood what the target is, whether it’s CSAT or something else. The goal is just how do you make a system as efficient at getting to a high number.

Priya (06:43)
Do you find that you have to define the metrics because there wasn’t anyone solving this before? Because I feel like in every organization, people were trying to solve this, right? Through humans, of course, like you’re trying to debug, you’re trying to make sure you recover from a production incident, right? And we always track the mean time to recovery, etc., right? Or in case of bugs. SLA, right, to fix them. So when you say that you kind of had to define the metrics, get that understanding across the industry, and then work towards them, I’m trying to understand how they differ from the traditional ones.

Willem (07:19)
I can give you a bit more color there. So MTTR is a good example. A lot of the teams you work with, they say MTTR is great in principle, but we have very few black swan events where the MTTR is really critical and we have not enough data on that side, even internally, to accurately know whether that you know, if you’re actually making a move a difference there. One month’s MTTR varies wildly to the next month. So they want something that’s got higher volume of higher volume of data and is more accurate and has less noise.

And so we’re looking for proxy metrics. Let’s say how aligned are we with the way your engineers debug? So you can look at the almost like the debugging steps that an engineer takes and how closely an AI maps to that, or you’re trying to find proxy metrics that give you a sense of our accuracy.

The other thing is you don’t always have the final root cause as ground truth to compare against, because that would require an engineer to go and actually solve the problem by hand and compare it to the AI. And so you’re looking for different ways to in an unsupervised way compare your performance to what they’re doing and have the other person believe it. I mean, it should be substantiated somehow. So MTTR of course is gold standard,

Shahram (08:25)
The takeaway I took from what Priya said was that yes, there are these very sort of scope metrics, but the more interesting challenge is that there’s a subjectivity involved in getting to the right one because you could get perfect containment to quote her, but the way you do it matters and that’s how you get sort of that user love, which I assume is much harder to measure. Like was that person satisfied with the conversation?

Priya (08:48)
Yeah, and the one way that people generally do this in or have tried to do this before is surveys, right? But now you don’t need them. You can just tell AI to judge for itself if it did what the customer asked to do. So those are goal evaluators, right? So in terms of like when it comes to our buyers, they’ll understand that these are golden metrics, right? To your point, Willem, these are the MTTRs of the world. But then what are the proxy metrics to actually understand you’re making the difference, which is the goal evaluators in our case.

Shahram (09:23)
Yeah. I think staying on the platform side, and your transitions, Priya, you did another transition, I gather, within ASAPP too. Correct me if I’m wrong, but I think you started off on a more platform side of things and now you’re obviously leading all of engineering. So maybe that’s one of the reasons we connected because you know we’re kind of platform people at heart. But talk us through that. You’re obviously thinking very much at the business level right now, which I’d love to focus the conversation on. But what was that transition for you like? You know, do you feel like you’re more sympathetic to platform teams given you know, this transition?

Priya (09:55)
So I definitely think that I have a soft spot for platform engineering in my heart. And in fact, when we talk about leveraging AI tooling within organizations, I feel like platform engineering still remains the highest leverage team that you can have in any organization. Because I was in a conversation yesterday where so we have a platform team, obviously, the one that I created at ASAPP and I was talking to my CISO and he was talking about a conversation he was having with someone and it just sounded like what they were missing was someone whose job is to think about adoption of these tools in their company, right?

I see that a lot. Because this gives you power, this gives every person the power to sort of create, you see duplication, like. You can easily find three tools in the company doing the same thing, right? And then you don’t know which one is going to be adopted, which one should you maintain, give TLC to for the long term, right? So that’s kind of what you end up seeing if you don’t have an adoption that is driven through the platform engineering. So that I definitely will say is still my strong belief because of where I came from also. But because of what I truly believe, right? In terms of what that change meant for me.

So I love thinking about infrastructure platform. I think it gives me a lot more to think about actually in this world. So for example, when we talk about inference, like I have to think about, do I run this with bedrock or do I go to any of the Neoclouds right? And what do I get from one versus the other? What does that mean for me in terms of pricing? So all of these things are still relevant, right? In terms of developing the product. But I do think that allows me to accelerate the product development.

So ultimately providing initially it was about providing value to the internal users, right? And now it is about providing value to the external users as well. And I do think that being able to leverage this team allows me to do that job as a head of engineering much better because then I can create golden paths for all my engineers to move faster rather than just, you know, for some internal work.

Shahram (12:09)
One of the things I’m curious about, having led platform teams, I feel like there’s always this I don’t want to say imposter syndrome, but there’s always like this very high difficulty which I felt where people felt about connecting to the mission because you’re always one step removed. You are doing something to make the engineers who are building the product more productive. So as a manager that was always like really hard where, you know, that’d be that’d be an all hands and the CEO would be calling out all the wins, and it’s almost always going to product engineering. It never really feels you know that you’re getting recognized.

And I think in some ways AI has worsened this because you’ve got the coding agent users on the product side and they’re just shipping so much faster. There’s all of this talk. I’m curious, how have you changed how you manage platform teams maybe even better, like are you seeing this? And if so, like how have you adopted your management style?

Priya (12:58)
Yeah, so I do think this is a classic problem of the wins are always on the product side. But it depends on what the people on the platform team are looking for, right? Where you derive your happiness from satisfaction from, that’s one, right? And then it’s the it’s the job of managers to like make sure that that work gets represented.

So one of the few things in terms of like having appreciation for the end user, one of the few things that I’ve seen work well and have done is bootstrap the team by using people from within the org that actually understand the challenge, that actually have built the product, right? The other thing is whenever you form the platform team, I think it is really good if you can form teams with people from different backgrounds. For example, you can bring in some application developers who, well, some of them actually have appreciation for the application development challenge within your company, right? Or you can bring some SRE folks or some QA folks, and you look at a problem from multiple dimensions and then design a solution for that, right?

The other thing that I think helps in terms of other people developing appreciation for platform teams is. We always adopted sort of like a open source software models, like so you’ll have special interest groups, they can like people can contribute, we’ll have contribution guidelines, and then I would work with map my peers to say that hey, in the career ladders, we can put in some language around hey, if you’ve contributed to a common tooling or platform, then that’s like something that is expected from a senior person, right? So there are ways like this where you can sort of give that work a lot more visibility. And then of course you have to provide value, right?

There are some other ways that you can do this from like a leadership perspective, which is work in some OKRs or goals that are directly connected to platform teams’ work.

So one example I could give you is we moved from a single tenant system to a multi-tenant system, which was a huge migration. If you think about boiling that problem down to some of the first principles, you’re basically talking about different deployment models. When you’re talking about different deployment models, you’re essentially talking about controlling your configuration. Because if I can control the configuration, I can deploy in any way that I want, right?

So what I mean is that if I want to deploy in a certain way, I build some configuration logic that hydrates the artifacts to be able to deploy in that model versus another, right? So if I make the configuration hydration a challenge for the CI CD team or the platform team, that is a huge impact on literally the gross margins of the product, right? Because now I’m running a multi-tenant system. So there are ways like this where you can actually create visibility for the platform teams at leadership level.

Shahram (16:00)
Anything that’s changed with AI specifically? Because you know, you mentioned bedrock. Obviously, like token costs are a big deal. I really like this point about you having very clear, almost like business orientation around the platform team’s work. So if you’ve hey guys, you improved the gross margin. You’ve actually made it easier for customers. There’s probably like some business goals around going from single tenant to multi tenant, which then you can actually attribute directly to the platform team’s work. Just curious like anything from the AI side that you’ve seen which could be helpful for people listening.

Priya (16:30)
Yeah, so I mean I’m not gonna get into like the token maxing and all that. Like I think we were much ahead of like the yeah, it’s like thing of the past, right? The ways that I think the tooling, right? You have to make the right thing to do, the easy thing to do.

So for example, when we wanted to make sure that we are not breaching or exploding our token costs, etc. We worked with our security and IT teams, like the platform team worked with the security and IT team and also our finance team to ensure that we had like the right budgets in place and we put in like a self-service workflow for people to be able to go and increase their token limits if they had to with like a valid business reason or whatever, right? And then just making all of that super easy, super self-service, creating awareness around it. Training about it, et cetera. Like those are like those are some of the things when it comes to like the token costs.

The one thing I do think where platform team can play a role. And I haven’t yet figured out how this can be done. And I’m curious about your thoughts here as well, which is how do you know that you’re getting the right value out of your token? So for example, Shahram might be using it in certain way and Willem might be using it in certain way. And Shahram’s way is much better. Sorry, Willem. How will you learn from him? Right?

That is not something that has been codified yet. It’s a lot of like, let me share success stories or failure stories, and we learn from each other. And I do believe that there is a lot of opportunity there in terms of tooling to be able to codify that, but I haven’t yet figured out how. So I’m curious how you guys are thinking about it because you also have a team, right?

Shahram (18:10)
Yes. Willem, you want to take this one?

Willem (18:13)
I think if you continue to think about that eventually you want to have a measure of business impact. And I think there’s the reason this is not a generalized like a vendor solution or office shelf solution is it’s kind of different for every business. And the level of integrations and data you need and decision making is all the way downstream closer to the customers.

But I think there is a gap in the middle. I think what you’re describing is all the way from like tokens to code to like usage, where there’s still nothing and we can probably do more there too. Like you know, I hate to name specific things like how many lines of code, like how many paths are actually being used, because you can ship a lot of code and none of it’s used. Or if it’s used it’s one customer.

Shahram (18:51)
You should talk about the engineering dashboard that you shipped. Remember where we saw that I think it was Dan using this new skill that was cool that, you know, sort of propped up you know, that was cool.

Willem (19:02)
Yeah. Well, I think it’s that’s a different problem, but I can talk about that. So there’s kind of two orthogonal ones. One’s is like what you’re shipping and what impact that has and how much work one token is doing, like right. And the other one is what we are seeing, even in a small team, and I’m sure it’s in a big company, there’s a lot of siloing of ideas. A lot of techniques are being developed at the individual level and it’s not disseminating through the organization. So the spreading that like a platform team that may has good judgment can do that. But you can also make a grassroots, like a software engineer can show something to another software engineer.

So the first thing we built is hooks into our agents that would process in a very safe way, like a PII-friendly way, the commands that are being run. So it wouldn’t show the actual data values. But if somebody’s invoking a skill, we’d store the skill’s name. If they’re running a CLI, a specific CLI, we’d store that and we’d have a like a leaderboard per person, like rates of just like percentages of specific skills being invoked. And then on a heat map level, you can see this person is really using the skill a lot. And then we’d ask the question, what does this do? And we’d have that person show it to the team. And maybe it’s a great thing that they can use. And so at least in our team, we can share ideas.

That doesn’t mean those ideas are efficient. You can still have waste at the token level. But I think at a bigger company, that waste becomes a lot more important. But I haven’t seen a good solution. Yeah.

Priya (20:17)
No, I love that idea because I love that idea. And I’m gonna take that back to my team because we’re kind of building a in house marketplace of skills too. So this this is good. I think a heat map of the most used skills and then showcase them. I I think that’s pretty cool. So I I love that.

Shahram (20:34)
I don’t I don’t think we’re close to at least like really definite ROI measurements. Like, you know, when you like one of the things that I think about is the reason we’re so focused on the cost aspect is if you think about ROI, the I is so easy to measure, the R is not. So all we’re doing is trying to reduce I, right? And that’s that’s the only lever we have.

Priya (20:51)
Yeah. I yeah.

Shahram (20:54)
And so I think about like how can you even start adding some proxy metrics around the R. Because if I spend five thousand dollars in tokens and I generated five hundred thousand dollars of value, who cares? Like it absolutely does not matter, right?

Priya (21:07)
Yeah, and that’s the challenge, right? You don’t know the R until and over what timeline. So I think that matters quite a bit to be able to quantify that. And those to your point, those metrics are very hard because you’re kinda I could say like, this feature will definitely get used by like, you know, hundred users and they will be willing to pay me like I don’t know, ten thousand each for this. But will that happen? And over how much time? I don’t know, right? I’m curious how you think about it. Like is your product roadmap obviously you have been in the space, you know what to build, so you bootstrap a product, but then how do you curate your feature roadmap?

Shahram (21:46)
That’s a great question. Honestly, I don’t know if we have the perfect answer yet. But what we’re trying to do is, you know, like if you think about the R question, I think about it as almost like three levels of metrics. You always have business metrics, which are primarily revenue driven, right? And then you have product metrics, which is like, okay, what’s getting used, things like that. Mostly adoption. And then you have engineering metrics, which I think we’re pretty good at, right? Like rate of PRs, lines of code, how many tokens you spend, things like that.

What we’re focused a lot on is that second layer right now. So we try to instrument every feature which we build with good product metrics to measure adoption. Like is this actually being used? We had all these assumptions before shipping something. And is that actually happening? So there’s obviously like a more change management side of things. Just because you release something doesn’t mean people use it. So you do need to let it sit.

We also try to do things like we don’t release new features to users immediately. We insist on dogfooding them so that we can actually try it and just get a feel for it. And when we feel like we actually really like it, we push it. I think what that does is again, like it’s something which I think there’s still a lot more maturity there where we wanna start getting good at actually deleting features, which I think is like more interesting than adding them. Right? Yeah.

Priya (22:54)
I I was about to ask you that. Because that’s where the timeline question comes in. Because have you given it enough time? Or you know, in six months you’re deciding, okay, that’s it, this one’s not used. So have you done a feature deletion?

Shahram (23:08)
I mean let me think. We’re we’re planning a big feature deletion. We need to announce that to you. But

Willem (23:12)
We’ve done many. I think Shahram is I mean, I can give you an example. We had a lot of intuitions around the work that users would do. For example, we had like this review dashboard back in the day where we’d do work and then we’d ask the engineer to have like help help us like tell us what was good and bad. And turns out, you know, we knew people didn’t want to do that, but they really didn’t want to do that. And so we had to make it more implicit. And so we deleted like a month’s worth of work because nobody was using it.

Priya (23:40)
Okay. I I’ll give you an example where

Shahram (23:40)
I I I think there’s a difference. Sorry, go ahead, Priya. I

Priya (23:43)
No, go ahead, go ahead. Yeah. It’s I was just saying that we actually did something similar, but it’s not more deletion rather than morphing. Like, I think you also do this, which is a thumbs up, thumbs down. So we kind of put it at an utterance level. And obviously no one wants to tell you at every utterance level whether you do the like it’s a good bad. They wanted more at the like the conversation level, but then you’re not really sure. How it helps you because they might just tell you, Okay, you solved my problem, right? So it’s sort of like you can morph the feature, but then is it doing the same thing? So I was just thinking about

Willem (24:18)
W we actually did that. So in Slack you can select a rating and you can actually leave a message. And I was surprised, I thought, okay, this is going to be the same, people are not gonna do it. But we got a lot of feedback from users and it’s been absolutely gold. Now it doesn’t tell you how accurate you are. Like initially we thought, well, this is an accuracy rating if they’re giving us five out of five, but it’s mostly I’m confident in this answer. It looks good at a quick glance. So we still had to build a system that’s independent to do proper correctness scoring, but it was still good to get user feedback.

Priya (24:45)
Makes sense. Makes sense. Yeah.

Shahram (24:46)
I wanted to make almost a distinction between refining a feature versus deleting it. Right? Because I think you did it a lot in refinement. You have an idea, it’s not great, you keep learning. So even like this sort of review thing, we’ve obviously refined it a lot. And trying to be making it better.

Deletion to me it’s like you told somebody that you have this functionality and then the next day it does not exist anymore, like in its entirety. That’s something which I think about. How do you know? Because I’m sure you’re seeing the challenge as well. As you get more users, it’s interesting like they all congregate around very different parts of the product, right? So it’s not easy to just take away something, but then that’s I think the discipline.

But just to connect to point to your question, I think the next step is to start connecting these product metrics to business. So that now that you actually have like the way I think about trying to make things more agentified is that how do you expose as much of this data back to the agent, so that we can start making better product decisions right? But sorry, go ahead. I think you wanted to say something.

Priya (25:41)
Yeah. Exactly. And no, I was just gonna talk about the same thing. We are thinking about the same sort of like the flywheel, right? Like how do you figure out what metrics you’re looking at and then bring it back to modify your agent and then it keeps doing better. And so that’s like the ultimate where you get to and you wanna make sure that you choose the right metrics, right? Because you could hill climb till the end degree, but then that’s the cost and outcome trade-off that you have to think about. So yeah, but we are also thinking about it, of course, in the customer service space.

Shahram (26:15)
Talk me through that a little bit because I think you’ve coined this golden pathways, not coined, but it’s a framework that I think you follow quite strongly. So you’ve obviously been thinking a lot about how to build really high functioning organizations, especially on the engineering side. What are changes that you are now making with just this insanely increased output from coding agents and sort of how you’re dealing with it. Because just connecting this back to our previous question, it’s now attempting to create hundreds and thousands of features. Like you can, it’s just easy. But that’s probably not the best thing to do. So, how do you manage this?

Priya (26:47)
First of all you have to look for signals where you know that this is creating challenges, right? So one of the things that I’m worried about right now is the depth of what we were talking about before, which is the RCAs, right? Like I want to make sure that the depth of our root causes don’t become shallow because that kind of would be an indication of how much people are not understanding what’s out there, right? So obviously you use coding agents to put out more code. If you use an automated PR review tool, fine. Someone definitely looks at it. And I don’t know if it is just a stamp of approval because you trusted the person or because you are biased by the automated review.

But ultimately you have to go and understand whether the people putting out these things are still in touch with what their system does, how the whole system works, right? And to an extent, you know, like everyone talks about Greenfield, Brownfield, I think we have been around for a while. We kind of have some of these challenges that we are dealing with. And as much as you can say that because of AI, getting context, getting onboarding is faster, easier, et cetera. At the same time, is it actually happening? Or there is just some shallowness to all of this.

And what are the leading indicators, or maybe not even leading, trailing, I guess. Like because if it’s an RCA, it’s probably a trailing indicator. So that’s like one of the things that I’ve been thinking about. Like let me look at that.

So far, and especially because we serve enterprise customers, I think you kind of have to manage the velocity with the stability and the trust that your customers put in you, right? I don’t know if you’ve noticed this, but somehow fall seems to be the time a lot of infrastructure challenges come up. Like something happens magically during fall, I feel like. So I kind of have to eliminate seasonality from it and still understand if we are getting any bugs or more incidents, is it a result of lack of system understanding and the velocity without the right guardrails? I don’t think I have fully formed an opinion about whether we have gone too much on one end of the spectrum versus the other. I wanna say not really. I think we have struck a decent balance. But yeah.

Shahram (29:05)
Let me dig into that a little bit because I like the theme of what you’re talking about. It’s something Willem and I have been debating internally as well, which is I think with coding agents, there’s almost a very tempting, like a very strong incentive to go broad. Because you can just expand very, very quickly. And what I’m hearing you talking about is actually focus on depth. So in this case you’re saying let’s stay focused on the RCAs, let’s stay focused on, you know, stability. And these are all things that you can still apply with your coding agents, but it’s almost like requires a different kind of discipline to just stay on the durable problem. I don’t want to put words in your mouth, but is that kind of similar to what you’re thinking about?

Priya (29:45)
I’m not asking us to build new habits. I’m saying not to lose old habits because you have new tools, if that makes sense. Right. Like the discipline of software is still the discipline

Shahram (29:52)
Right. Right. I like that.

Priya (29:54)
Of software. You’re still testing what you built, right? And making sure that you don’t put something into production that’s gonna affect your customers, right? So I think the principles are still the same. You’re moving with velocity, so you want to make sure you don’t compromise on your system understanding and how it affects your customers’ interaction. You mentioned that you and Willem have been talking about this. So how how does that show up for you? Because you’re obviously in that space. Like I talked about RC as a trailing indicator, sort of like of the lack of system understanding. Are you kind of seeing that? Are you seeing like proliferation of issues that your agent has to look at? And the answer simply is someone didn’t pay attention to something really, really minor.

Willem (30:36)
That does happen. We’re seeing this at some of the teams we work with, I mean it’s even in our team, like the rate of shipping is so high that at least some of the I wouldn’t call them incidents, but paper cuts all the way to bugs, all the way to incidents come from the code that’s being shipped. And so the question is, is it leading or is it a lagging indicator? And I think most incidents, in my experience, have some breadcrumbs, some leading indicators.

The problem with those breadcrumbs is do you believe it when you see it? Because often you want to see the failure and you’re like, now I know I need to do something about this. So the one school of thought is react very quickly, and the other school of thought is can we prevent this before it happens? And so a lot of our analysis right now is trying to see how much can we prevent these failures in the first place. Whether it’s 30 seconds, 10 minutes, a day, you know, if you shift all the way left into the dev environment, you can potentially do that.

And I think the gap that I’ve seen right now is a lot of code that’s being generated is done based on LLM’s assumption about what is happening in production, an assumption about the world out there, of the behavior of users. And it loves to create code and tests and reviews all on those assumptions, but it’s only in reality when it’s shipped that you you might realize something else is the case. Might maybe a user doesn’t use it. Maybe they have only a limited amount of attention and they don’t even try that or they use it in the wrong way. So that’s where we’re focusing right now.

Shahram (31:55)
We’re trying to lean into that though. I mean, you know, we released this new thing called change verification where it’s no longer just looking at alerts and incidents or things like that where it’s reactive. It’s taking the PR and trying to infer your intent and then try to spot in advance if any regressions happened or if you have any instrumentation gaps.

So the way I’m trying to think about this is as an engineer, it doesn’t work when you tell people to slow down. That never works. You want to lean into saying, yeah, keep going, but let me plug the gaps for you with the agent. So in this case the gaps could be, hey, it’s you’ve missed out some instrumentation. Or the gaps could be it’s humanly not possible, no matter how much you review this, that you might get a ninety percent correctness, but the last ten percent will always happen in production. How do you make that more okay so that you can recover really fast?

Priya (32:41)
The thing that you spoke about, which is instrumentation is key. Actually, one of the OKRs that I always talk about is not the ambitious have zero incidents, it’s more have zero customer reported incidents. What does that mean? Like make sure that you have instrumented your code well enough that if you didn’t account for something and it happened in production, you are the one to figure it out and not your customer. So I think that’s a good metric to track even as your speed increases, right? Because then that helps you maintain that discipline. So

Willem (33:14)
So what I’m hearing is ship to prod, prod is really the one we trust, and then test it in prod before your customers find it.

Priya (33:23)
No, we do test it, not in prod first. But

Shahram (33:26)
Don’t put words in her mouth, Willem. Come on.

Priya (33:29)
No, you are saying that. I think you are shifting to prod is what I heard, but you’re trying to figure out whether… Yeah, no. No.

Shahram (33:34)
We are, we are. Guilty as charged.

Willem (33:36)
Yeah. Guilty. I think there’s another related challenge that we face in our team that we’re trying to solve as well. If you take this change verification further, which is the cognitive load on engineers. It’s not just alerts and incidents. It’s I’ve shipped so many things now and my team has I don’t even know if I’ve overwritten somebody else’s work, whether my system integrates properly with their system. Is somebody using this feature I shipped? Is it healthy? Sometimes people remember a feature they shipped a few months ago and like, wait, I should have a look and see look at the metrics and the data.

And I think the ownership of features and any really functionality or any expectations you have about your product, that’s a scope that agents will take on over time. I’m a firm believer in that. So we’re internally have our own tools that monitor our features that we release. And I think we’ll be leaning into that area a bit more so that an agent will tell you if something degrades or an agent will tell you if you need to expand the coverage of a specific feature to your whole user base or something like that.

Priya (34:31)
I do like that because one of the things that we have dabbled with, actually, one of the things that we are trying is to figure out if our designers can ship, right? And not like features, but like let’s say the big box is made by the engineers, and then the designers are playing with an end, trying to adopt more AI native component system for this, etc. And the act the tension when this was first floated was obviously who’s on call for it, right? The ownership of the feature. So

Shahram (35:01)
The designers.

Priya (35:03)
The designers, of course. So there’s a balance, right? Because you kind of can get faster feedback and iteration, but then at the same time, who’s on call, right? And then again, to your point, anyone can sort of make modification. If you have a contract, you make the modification, you ship the whole end-to-end feature. And I might not even know that someone did that, right? And I find out when I’m touching it like a few months later that someone’s using this contract that I built, right? I mean that can happen otherwise too. It’s just it can happen more now. So yeah, no answers other than obviously watch your production system very closely, I guess.

Willem (35:44)
In a sense, the software engineer is becoming somewhat sort of like a platform engineer in that sense because they’re like the platform team is running the software of the software engineers, and the software engineers have a little sandbox for the designers to operate within, but it’s their concern to make that a safe area. You can almost have to assume that it’s hostile, whatever the designers are creating. Yeah.

Priya (36:00)
Yes, whatever can happen. Yes, exactly, exactly. Like it all you can do is probably move the button, but make sure the button doesn’t just disappear, right? Yeah. So that makes sense.

Shahram (36:12)
I read this quote. Tell me if it’s still something that’s accurate. But you were talking about how you felt that agents should be secured like software, managed like employees and budgeted like productive capacity. I really like that because it’s almost like a very clear way of just seeing how ‘cause the agents are clearly not just software systems. They’re almost like, you know, productive bandwidth that’s available. How have you implemented this? Like maybe just talk us through that little bit.

Priya (36:37)
Yeah, so I think there was a debate a long time back. I don’t know if it has settled or not in the industry, which is around like are agents like your employees or you know, like how do I think about them? Do I get one senior engineer and like give them like four agents or whatever it is, right?

And the way I started thinking about it was more from like the budget aspect of it, right? First of all, because there was this expectation that you can think of these agents as employees, but many a times it’s like a vendor contract that you’re actually signing, right? And so then that becomes sort of like a software cost, right? So when it’s a software cost, does it go into your R&D expense, or do you think of like, and how do you think of it? Is it giving you capacity or it’s like the actual just, you know, a vendor cost that you’re line item on your R&D expense. That is not something you can optimize, right?

So that’s something that we had to think about. And that is why I kind of thought of like, okay, think of this as like the productivity bandwidth, right? To your point. So that’s like the one element of it.

When I talk about managing them as employees, when you think about employees, right? If a new employee joins your team, you have to give them context. You have to onboard them. You have to tell them this is how we do things here, right? And it’s the same thing. You can’t let loose something in your environment and just hope that it figures it out. You have to give it some guardrails. This is our monolith, or this is our well, distributed monolith, right? Like you kind of have to give a lot more context than you think.

So that’s the onboarding part of it and then the scope of it right we just discussed about like how much can the designers do so even if I’m giving them a design system that is AI native, can they just move the button or can they delete the button? That’s kind of like the scope of what you allow the tool to do and then in terms of securing them is the same story. Minimum access that is required. Don’t give it more access because we think a lot in terms of just like data, because we serve enterprise customers, so we have to be very, very careful about it. Yes,

Shahram (38:39)
I know this very well.

Priya (38:42)
You are on the receiving end of it. That is right. You will see like just the amount of questions that we have to ask the vendors to ensure that there is no possibility of data leak at all. We take redaction very seriously, ZDR very seriously, all of those things. It sort of like frustrates people sometimes. We have a setting in our cloud enterprise where if you have access web in the same session, you cannot access any of your internal connectors. And that’s like limiting for people, of course, right? But there is a reason that was put in place.

And the absolute disciplines of a discipline of like make sure you’re logging everything so that you can actually watch the logs, right? We have a pretty healthy security tooling that allows us to sort of make sure that none of our principles as a security organization will get compromised. All of those things have to be in place for us to be able to confidently say to our customers that, hey, these are our subprocessors, but we have made sure that they are within the guardrails and the security principles that we follow as an organization. So that trust extends.

Shahram (39:45)
I think we covered pretty much governance and compliance and things like that, because obviously those are the biggest concerns at the enterprise level. Part of that consideration is almost like curtailing the autonomy of the agent, right? That’s the kind of the trade off that we’re talking about. So you know if you plot where things are today versus say six months or twelve months from now, what’s your view on autonomy? Like do you see this getting pretty much to a hundred percent where an agent is gonna be as autonomous as a employee? Obviously like even as an employee, we have certain like constraints and guardrails. You can’t just go and do whatever you want. But do you see this like almost blurring or you know or not? The gap.

Priya (40:24)
So the way I think about it is you have to choose the scope of the autonomy, right? Like what I mean by that is we kind of talk about how you can or initially everyone thought about, okay, the agents can take the busy work from you and then you spend all your brain cycles on system design and architecture and making sure everything’s like well and good.

I actually do hope that agents become autonomous when it comes to that scope of work. And the reason I say that is because the one thing that I think about a lot is how your system design changes a lot. I think we spoke about this a little bit, Shahram, one time when I we were talking about how this technology is changing so frequently, such that you have to kind of guess. What not to build and build at the edge of the technology and it kind of moves you forward, right? And that is such a subjective thing to bet on.

Everyone talks about the first system and you kind of just always throw it away because it’s a prototype that went into production. Then everyone talks about the second system where you overbuild and it’s possible that that’s the one that sustains for a long time, or you end up in the third system where do everything perfect, right? What I have found is that because of the change of technology, your systems are iterating much faster, and I don’t know if you’re seeing that as well, but that makes migrations, etc., sort of like a norm in how you serve your customers. And those are where I would want a lot of energy. I would want my engineers to use their judgment there, right? Building those systems faster, like figuring out what other building blocks.

Shahram (42:01)
I see, I see.

Priya (42:03)
And so I do hope I’m not saying I’m not making prediction about I think if we scope as an industry, if we scope that hey, this is a scope of work that let the agent be autonomous, such that we can prepare ourselves for how software designs gonna evolve. I think that would be a great place to be. How do you think about it?

Shahram (42:23)
I mean, we’re all in, right? I think one of the advantages we have, honestly, is that being a small team, there’s no ceiling in terms of like what we could be doing, right? Like for each one of us, there’s a thousand other things we could be doing, but we can’t because there’s this, you know, other tasks that’s right in front of me that needs to be done. So the more I can offload to the agent, the more I can sort of move with hopefully like a higher level of abstraction.

So you know, like my view is that if you’ve picked a domain, then ideally a hundred percent of your time is spent thinking about what your customers and users want from that domain. You’re trying to understand how do you actually add value. Every minute you’re not spending on that is busy work because it’s stuff that you need to do in order to provide that value. So ideally if you can design systems around where agents can actually take on more of that busy work, I think that’s good. And it kind of to me cleanly solves this problem of what do humans keep doing and what does AI keep doing. But that’s me being bullish.

Willem (43:22)
There’s a good post on why software fails or systems fail. Actually written by an anesthesiologist, but he was like a systems researcher at University of Chicago. But one point he made is that you can’t just build a complex system. You have to build a simple system that works and then you have to grow the simple system up to a complex system.

And I think with agents right now, there’s a lot of one-shotting where you add maybe thirty minute thirty percent too much complexity in the first pass. And it kind of works, but then it fails in a different way. And before that’s cleaned up, another pass comes over that. So you’re layering complexity at a very high rate. And some teams, if they can do that safely and clean it up, that’ll work. Some teams are gonna go over their skis and it’s gonna be too hard. So I think that’s a unsolved problem where the slop is just growing too quickly.

The other thing that I’m seeing is that agents love to use kind of like tests to drive development. I think teams will move away. I mean, many are already moving away from code review, where they only review like critical parts like migrations and key APIs. You can’t operate at the line of code level anymore long term. It’ll move too quickly. And so the only thing you can do is really operate at the like tests and expectation level at the high level and hope that the agent creates the right seams and boundaries and systems in place. And so then you need proper in end-to-end tests or integration tests or service tests.

But you also need to learn from customers in production. And so the flywheel of learning from prod back into the dev loop is going to be the next big thing to solve. And you know, we’re also looking at that space quite heavily.

But what the effect of that is, is that because everyone is so dependent on large language models and other models right now, so your code base uses models and there’s data flowing back from production into your dev environment, your code base almost mutates so quickly, it becomes a model. It almost is very model like in that it’s trained on test cases that are very largely dependent on data itself. And so previously it was very deterministic and it was very slow changing and now it’s morphing very quickly. And so I find that very exciting. But I think we need to step back and look at the like high level expectations and review at that level, like the spec almost.

Priya (45:23)
Yeah, and to your point, I think the muscle that everyone has to build is like that experimentation A/B testing muscle, right? Because like that’s the feedback back into your product development. And that’s basically how we we have to deploy now, right? We kind of have to constantly experiment and see what works, doesn’t work. It feeds back and you learn from it. So yeah, I mean what you watch in production is changing. We did to your point, beef up our E2E suite, but that’s not enough. You kind of have to watch in production.

You kind of asked me a question at the beginning about how my journey from fintech to here, right? And the probabilistic system. And given that we are talking about what to watch in production, one of the interesting things that is happening is when we deploy to these regulated industries, I find myself in conversations talking a lot about how we have built a product that will give you a combination of determinism and then the generative AI capabilities where it makes sense. I think that’s a language that a lot of regulated industries appreciate and understand a lot more because they do want to push the boundaries, but they do have to work within the framework of the regulations that they follow.

One idea is this, right? Like the mix of the determinism, but the other idea that is important to them is the ability to observe. The ability to observe what the product is doing in production as well. So I think for us as we develop the software, we want to observe for them as they develop watch the product work they have to observe. So I think this whole idea of the busy work becoming autonomous and then you’re trying to figure out what you should be building and then what you should be observing, right? To know that what you built works and learning from that. I think that’s becoming more important.

Shahram (47:08)
When effectively you’re creating the right primitives so that like observability is a huge one, so that you can experiment faster, but you’re also trying to figure out how do you do it in a way that customers can actually palate. Like obviously like regulatory industries are gonna be very different from non, right? I think that’s a good point to end this here. I’d say it’s a pretty positive note. Thanks so much, Priya. See you again soon.

Priya (47:30)
Yeah, that sounds good. We’ll get coffee now that I’m here. Sounds good.

Shahram (47:32)
We should.

Priya (47:33)
All right. Bye.


Want to see how Cleric works? Book a demo

We’re hiring too. See open roles