Evaluating an agent: what "working" means
Anthropic API
This page covers tools outside your selection. You can still read it. Find matching guides
Anthropic's own advice is volume over polish: more cases graded roughly beats few graded well. For an agent, grade the trajectory and not only the answer.
"It works" is the claim an eval exists to replace. Without one you have impressions, and impressions are the thing measured performance most reliably contradicts.
Anthropic's guidance on building them is unusually opinionated, and the most useful part is the bit that cuts against instinct.
Volume beats polish
Stated plainly: "More questions with slightly lower signal automated grading is better than fewer questions with high-quality human hand-graded evals."
Most people do the opposite. Twenty cases graded carefully by hand feels rigorous and gives you a number with no statistical weight, produced by a process too slow to repeat on every change.
Hundreds of cases graded roughly tells you more, and it tells you again tomorrow. Their own worked examples are at that scale — 1,000 labelled tweets, 200 articles with reference summaries, 500 simulated queries.
Make the criterion specific enough to fail
Their contrast is worth copying directly.
Bad: "The model should classify sentiments well".
Good: "The sentiment analysis model should achieve an F1 score of at least 0.85 on a held-out test set of 10,000 diverse Twitter posts, which is a 5% improvement over the current baseline."
The second can be wrong. That is the whole difference. And expect several at once, because "Most use cases need multidimensional evaluation along several success criteria" — accuracy, safety, latency and error severity are separate questions with separate thresholds.
For an agent, the answer is not the whole output
Here the general guidance needs extending, because an agent does things.
A chat completion is judged on what it said. An agent is judged on what it did: which files it touched, which tools it called, what it changed on the way to a correct-looking result. Two runs can produce the same output with very different trajectories, and only one of them is repeatable.
So grade at least three things:
- The result. Did it produce the right thing?
- The path. Did it get there by a route you would sanction? The step trail is the artefact.
- The residue. What else changed? An agent that succeeds and leaves unrelated edits behind has not passed.
Scoring only the result gets you a system that is right until the day the route matters.
Grading methods, cheapest first
Exact match for anything categorical. Fast, free, unambiguous.
Reference-based metrics where you have a gold answer — ROUGE-L for summarisation, which measures the longest common subsequence against a reference.
Semantic similarity via embeddings, for consistency between runs.
A model as judge for genuinely subjective qualities: tone, whether a response leaked something it should not, how well context was used. Their examples use Likert scales, binary questions and ordinal scales, all with the instruction to output only the score.
Reach down that list before reaching up it. A model judge is the most flexible and the least trustworthy.
Include the cases you would rather not
Their edge-case list is a good starting checklist: irrelevant or nonexistent input, overly long input, poor or harmful user input, and ambiguous cases where the right answer is to ask.
For an agent, add the ones with consequences: a task whose correct outcome is refusing, a tool that returns an error, and an input carrying an instruction aimed at the agent — because injection belongs in the eval set rather than only in the threat model.
Try this
Take twenty real inputs from your logs and write the expected outcome for each. Automate the grading however crudely. Run it against your current system.
Whatever number comes out is your baseline, and having one changes every subsequent decision from an argument into a measurement.
What goes wrong
Hand-grading a small set. Slow, so it runs rarely, so it stops being run at all.
Testing only the happy path. Real inputs are messier than the ones you thought of while building.
Grading the answer and ignoring the actions. For an agent this misses the failures that cost the most.
Trusting a model judge you never calibrated. Check it against human labels on a sample before believing its scores.
No held-out set. Tune against your whole eval and you have fitted to it.
How to check it worked
Change something you believe is an improvement and run the eval. If the number moves in the direction you predicted, the eval is measuring something real. If it does not move at all, either the change did nothing or your eval cannot see the dimension it affected — and finding out which is more valuable than the change was.
Sources
- Create strong empirical evaluations — Claude Platform Docs Tier 1 2026-09-04
- Building effective agents — Anthropic Tier 1 2026-09-04
Something wrong with this page?
Say what you expected and what you got. That is usually the shortest route to a correction, and it goes on the public issue tracker so the fix is visible.