Claude Code

Reviewing what the agent wrote

Claude Code

You are reviewing a diff whose author cannot tell you what it was unsure of. Read for the things it never flagged, and let a fresh context do the first pass.

Applies to
Claude Fable 5.1 Claude Opus 5 Claude Sonnet 5 Claude Haiku 4.5
Last verified
Reviewed by
Timothy Fehr

Reviewing a colleague's change, you get a signal for free: they tell you what they were unsure about, what they nearly did instead, where they guessed. An agent's diff arrives without any of that. Every line looks equally considered.

Anthropic names the resulting failure the trust-then-verify gap — "Claude produces a plausible-looking implementation that doesn't handle edge cases" — and its fix is blunt: always provide verification, and "if you can't verify it, don't ship it."

Make it show evidence before you read anything

The highest-leverage move happens before review starts. Anthropic's phrasing: "Have Claude show evidence rather than asserting success: the test output, the command it ran and what it returned, or a screenshot of the result."

"I've implemented and tested this" is an assertion. Pasted test output with the command above it is evidence. The difference costs one sentence in your prompt and removes the largest category of wasted review time, because "reviewing evidence is faster than re-running the verification yourself."

Use a fresh context for the first pass

A model that just wrote the code is a poor reviewer of it. Anthropic is explicit that "a fresh context improves code review since Claude won't be biased toward code it just wrote."

A review subagent "sees only the diff and the criteria you give it, not the reasoning that produced the change", so it judges the result on its own terms. The bundled /code-review skill does this in a fresh subagent and reports back into your session.

Read for what is absent

Automated review catches what is wrong in the diff. Your advantage is knowing what should have been there.

Ask of each change: which existing pattern in this codebase does it ignore? Which error path has no handling? What did the task require that the diff does not touch? Scope creep counts here too — files changed that the task never needed is worth a question, because it usually means the agent solved a problem you did not ask about.

Check what the diff does not show

The diff shows tracked file edits. It does not show what a bash command did, and checkpoints do not track those either. If the session ran a formatter, a codegen step or a migration, git status and git diff are the honest picture. The conversation's summary of itself is not. See git hygiene.

Try this

On your next agent-written change, before reading the diff, ask: "What in this were you least confident about, and what did you decide without being told?"

The answer is generated, never recalled, so treat it as a pointer and nothing stronger. It is still the fastest way to find the two or three places worth your attention first.

What goes wrong

Reviewing at the wrong altitude. Line-by-line reading of a large agent diff exhausts attention on formatting while the design decision goes unexamined. Read the shape first, then the lines that carry it.

Letting "the tests pass" close the question. Tests written alongside the implementation encode the same misunderstanding. Passing means consistent, and consistent is not the same as correct.

Accepting the summary as the diff. The conversation's account of what changed is a description. git diff is the change.

Reviewing while the author watches. Asking the model that wrote it whether it is right gets you agreement. Use a fresh context, which is the entire reason the subagent pattern exists.

How to check it worked

Pick one behaviour the change was meant to produce and exercise it yourself, outside the conversation — run it, hit the endpoint, open the page. If it does what you wanted, the review held. If it does not, note whether the diff contained the clue: a review that could have caught it teaches you where to look next time, and one that could not tells you the verification step was the missing piece rather than the reading.

Sources

  1. Best practices for Claude Code — Anthropic Tier 1 2026-09-02
  2. Checkpointing — Claude Code Docs Tier 1 2026-09-02