September 1, 2026

Roll for initiative: our week at AI Engineer World’s Fair 2026

Everyone knows agents can write code, but the open question is what happens after the code is merged.

Roll for initiative: our week at AI Engineer World’s Fair 2026

We spent four days at AI Engineer World’s Fair 2026. Our pitch was that the alert-driven AI SRE is only half of the product.

An agent that waits for an alert, investigates, and reports a root cause is useful, and it’s still at the core of what Cleric does, but it only starts working after something has already broken. Writing code is no longer the constraint. Code review, deployment, and catching regressions are.

So we asked people at the booth: how do you know the change you merged last week did what you meant it to do?

Most of them could not answer.

The Cleric team at the AI Engineer World’s Fair

The fair ran from June 29 to July 2 at Moscone West in San Francisco. Most of our team was there, so the people who actually build the product staffed the booth.

Conversations mostly started the same way: someone would stop, read the banner, and ask what Cleric is.

The short version, which we gave a few hundred times: Cleric is an AI SRE. Our core product is an on-call agent that takes an alert, or a half-formed problem statement, and runs an open-ended investigation to find the root cause. What we’re building now applies that same capability earlier in the lifecycle. Cleric monitors code as it ships to production and checks two things: did the change have its intended outcome, and did it introduce a regression? Merge and deploy are two different events, and on plenty of teams, a week can pass between them.

The part that got slow

Writing code used to be the most time-intensive part of building software, so the tools went there. Now that code generation is commoditized, the bottleneck is code review, walking a change through deployment and release, and monitoring for regressions that surface three days later.

We got the same three follow-ups all week:

  • How do you know it worked?
  • Who owns it when it breaks?
  • Does it fit the pipelines you already have, or does it go around them?

Most teams have a solid deploy pipeline and good alerting for the moment something falls over. Far fewer could tell us how they’d catch a change that shipped clean, passed every test, and then made checkout 200 milliseconds slower for one segment of users. Nothing pages you for that.

Cleric team talking with attendees at the booth

The answers were split by company size. Smaller teams were frustrated by how long the stretch between merge and confidence has become. Larger ones asked what evidence an agent would have to show before they’d let it own any part of that.

Your agent can’t tell if it’s right

Willem gave a talk with that title on Expo Stage 2 on the last day.

Coding agents aren’t reliable because they’re clever. They’re reliable because something quickly tells them they’re wrong. Write a function, run the tests, and the answer comes back red a few seconds later. The mistakes get caught before a human sees them.

That’s not how production works. If you ship a change that does the wrong thing, nothing comes back. There’s no difference between nothing and success until someone happens to look.

The booth

We went big on the Cleric theme. Behind us, a cleric looking out over a valley. In front, a hand-drawn map so people could find us, and a d20 big enough to worry a halfling.

We also sent an elf and a knight into the expo hall. They asked attendees how production had been treating them and brought the ones who winced to our booth.

We had loot too: custom stickers, gold coins, and lollipops shaped like 20-sided dice.

Our elf and knight in the expo hall, and the booth loot

Everyone got one roll on the big d20, and a natural 20 won. Two LEGO sets, two winners, and a lot of theatrical groaning from everyone who rolled a 3. Congratulations to Minghua (Ben) Jiang from Okta, who left with Rivendell, and Rohan Ramanath from Nubank, who took the D&D set, dragon included.

Our two LEGO winners at the booth

Moscone is four halls with no natural light and a loud demo around every corner, a dungeon in its own right. Thanks to everyone who found their way to our booth, rolled the die, argued with us, and stuck around to talk it through.

See you at the next one!

Get the Cleric newsletter

New writing on AI SRE, operational memory, and what we’re learning in production — straight to your inbox. No spam, unsubscribe any time.

By subscribing you agree to our Privacy Policy .


Want to see how Cleric works? Book a demo

We’re hiring too. See open roles