Evaluating AI Agents Like You Ship Software
On the agent I work on now, the same question can produce two different paragraphs that both might be right. Vibes-based testing worked at zero scale. It did not survive a deploy I actually had to trust.
Deterministic software comes with an obvious test. Input A, expect output B. Run it in CI and a regression shows up as a red check before anyone merges.
An agent breaks that on the first call. Ask it the same question twice and you get two different paragraphs that both might be right. On the agent I work on now, we fell back on what I have started calling vibes-based testing. You run the agent, read the output, decide it looks fine, ship. It works at zero scale. It stops working the first time something breaks in production and you cannot tell whether your fix helped.
Treat it like a test suite, because it is one
The setup that made deploys feel safe for me was to stop treating evaluation as a research problem and start treating it as a software one. Four pieces do most of the work in this suite.
- A ground truth set. Real inputs paired with the output we know is correct, curated by hand, checked in with the code.
- Assertions that run without a human. Not “does this read well” but “does it cite a source that exists”, “is the JSON valid”, “did it call the tool it was supposed to”.
- Sampling. You cannot grade every production response, but you can grade a random slice and put a confidence interval around the pass rate.
- Regression tracking. When the model version changes the pass rate moves, and you see it as a number instead of a vague feeling that the agent got worse.
I am not saying this is the only evaluation design. It is the one that matched how we already ship software, and it is the one I would reach for again on a similar agent.
The hard part is speed, not cleverness
The evaluation suite runs on every deploy. It grades answer generation across four frontier models and blocks the release if the pass rate drops.
Getting the metrics right took a week. Getting the suite fast enough that people actually wait for it took a month. A correct eval that takes forty minutes gets skipped under a deadline, and a skipped eval is worth nothing.
The model call is the tenth of the system you can see. The other nine tenths decide whether you can trust the tenth.
The same lesson showed up when I stripped two notebook agents down to a while loop, and when a clinical write path needed a composite key more than it needed a better prompt.
The reason to build this is not to prove the agent is smart. It is to make it safe to change.
FAQ
- If the model can produce two right answers, how do you have ground truth?
- The suite does not grade prose quality. It grades things a machine can check: did it cite a source that exists, is the JSON valid, did it call the tool it was supposed to. Two different paragraphs can both pass.
- Why block the deploy instead of just logging the pass rate?
- A number nobody waits for is decoration. The suite is only doing its job if a drop actually stops the release.
- Isn’t sampling just hoping you got lucky with the slice?
- You cannot grade every production response. A random slice plus a confidence interval around the pass rate is the version of “we measured it” that we could actually run on every deploy.
Amisha
Filed under Agents · AI Systems