Back to blog
Peter Richens
3 posts
Our LLM judge scored worse than chance until we made it compare
We rank an agent's investigation traces by comparing them in pairs and turning the wins into an Elo score. Grading each trace on its own scored worse than chance. Comparison made the ranking work, and it catches regressions our other evals miss.
How we verify Cleric’s production fixes
A correct diagnosis isn’t the same as a fixed problem. We built a verifier that checks Cleric’s fixes against the best ground truth available: production itself.
The Hidden Complexity of Building an AI SRE
Building an AI SRE isn’t just connecting an LLM to dashboards. It means reasoning through hidden dependencies, conflicting signals, and cascading failures in live production systems.
Give your on-call a headstart.
Start for free, or talk to us about a plan built for your team’s scale and security needs.