Willem Pienaar
5 posts
Our LLM judge scored worse than chance until we made it compare
We rank an agent's investigation traces by comparing them in pairs and turning the wins into an Elo score. Grading each trace on its own scored worse than chance. Comparison made the ranking work, and it catches regressions our other evals miss.
Why Your AI SRE Needs Memory
Investigation logic is becoming a commodity. What won’t commoditize is operational memory: the ability to capture and persist engineering judgment.
Cleric Named a Cool Vendor in the 2025 Gartner® Cool Vendors™ in AI for SRE and Observability
We’ve been named a Cool Vendor in the 2025 Gartner Cool Vendors in AI for SRE and Observability report. We see this as validation of our approach to building a self-improving AI SRE.
What is an AI SRE?
Engineering teams are deploying AI agents to handle production operations. This deep dive shows how AI SREs work: building system understanding, investigating issues, and driving resolution. You’ll learn their current capabilities and limitations, and how they will change the way engineering teams operate.
Introducing Cleric: The first autonomous AI site reliability engineer
We’re excited to announce that we’re building an autonomous AI SRE, called Cleric, backed by Zetta Venture Partners in a $4.3M seed round. Cleric is an AI teammate designed to autonomously manage, optimize, and heal software infrastructure.
Give your on-call a headstart.
Start for free, or talk to us about a plan built for your team’s scale and security needs.
Speak to an engineer