How do you test an AI agent?
You test an AI agent with evals: a fixed set of real inputs, the outcome each one should produce, and automatic checks that run on every change. Good evals check the actions the agent took as well as the text it wrote, and they include cases built to make it fail.
Run the suite before every change ships and after every production failure, so each failure you fix stays fixed.
What an eval is, and how it differs from a benchmark
A benchmark is a public test of general model ability, such as coding or math problems. It tells you which models are strong in general. It says nothing about whether your agent handles your refunds correctly.
An eval is a test of your agent on your work. Each case has an input taken from real use, the result a competent person would produce, and a check that decides pass or fail. The checks run automatically, so the whole suite runs automatically on every change. Benchmarks help narrow the list of models worth trying. Evals decide which prompt, model, and tool setup goes to production, and they catch the day a change makes the agent worse.
What goes into an eval suite
We build the suite from your history, such as past tickets, documents, and cases a person had to fix, and add to it every time production shows a new failure. A typical suite includes:
- Common cases taken from real tickets, emails, or documents, with the result a person produced.
- Edge cases: missing fields, unusual formats, requests that span two systems.
- Adversarial cases: instructions hidden inside a document, requests the agent should refuse, and questions its data cannot answer.
- Action checks: whether the right tool was called with the right values, and whether the call succeeded.
- Cost checks: how many steps and tokens a case took, so a change that doubles cost fails the suite.
- Regression cases: every production failure that has been fixed, kept so the suite fails if it comes back.
Checking actions as well as answers
A support agent that says "your refund has been processed" after the refund call failed has told the customer something false, even if the reply reads well. So each case checks what happened in the systems the agent touched, using exact checks: was the record created, did the payment call return success, is the date correct.
Some qualities, like tone or completeness, need a judgment. For those a second model grades the output against written criteria. We prefer pass or fail against specific criteria over a 1 to 10 score, because a numeric score drifts between runs and does not tell you what to fix. Exact checks and graded checks run side by side, and a case fails if either one fails.
How we run evals on our own agents
Our own products run eval suites with simulated users, including adversarial ones, against the production agents on real conversation threads. The suites cover resistance to injected instructions, honesty about what the agent can and cannot do, honesty when a write fails, and cost discipline.
When an eval finds a failure that code can catch, the fix ships as code. Our evals showed a model pairing dates with the wrong weekday, so code now checks every date and weekday pair in that product's output before it reaches a user. We have also built adversarial evals and feedback loops for a production agent that analyzes cloud spending across AWS, Google Cloud, and Azure. For a client, the suite lives in your repository and runs automatically before every release.
Getting this set up
If you want this set up for an agent you already run or one you plan to build, start with the one-week audit. It ends with a ranked plan and something working by Friday, and larger builds are quoted in writing after it.
Questions
How many test cases does an AI agent need?
Enough to cover each kind of request the agent handles, plus every failure you have already seen. Start with the cases people handle most often and the ones that went wrong, then add a case each time production shows a new failure.
Should you use an LLM as a judge in evals?
Yes, for qualities that need judgment, such as tone or completeness, graded pass or fail against written criteria. Anything that can be checked exactly, such as whether a tool call succeeded or a date is correct, should be checked with code instead.
Can you evaluate an AI agent that is already live?
Yes. We start from its production history: past runs, complaints, and cases a person had to fix. Those become the first suite, which then runs before every change to the agent.
