Workflow: reviewing a change
Claude + Claude Code + Anthropic API
This page covers tools outside your selection. You can still read it. Find matching guides
Three steps, three different models. Reading is not judging, and paying Opus rates to read a diff is the commonest waste in this workflow.
Reviewing a change looks like one task. It is three, and they have almost nothing in common: reading a lot of code, judging a small amount of it, and checking specific claims against specific files.
Running all three on the same model is the default, and it is why review feels expensive.
The steps
- 1 · Read diff + filesbulk intake, inventory only
- 2 · Judgecorrectness and design — what you pay for
- 3 · Verify findingsnarrow checks, fresh eyes worth more
1. Read the diff and the files it touches. Bulk intake. The output is an inventory: what changed, where, what it calls. No judgement is being exercised and none is wanted yet.
2. Judge correctness and design. The actual review. Does this solve the stated problem, is the approach sound, what breaks that no test covers. Small input relative to step 1, and the step you are actually paying for.
3. Verify each finding. "Line 42 has an off-by-one" is a claim about one place. Checking it is narrow, mechanical, and does not need the model that produced the claim — arguably should not be, since a fresh check is worth more than a confirmation.
What it costs
Assumption: A ~500-line diff across six files, reviewed in roughly eight turns. Figures estimate a representative run, not yours.
An illustrative split
| Step | Model | Input | Output | Cost |
|---|---|---|---|---|
| Read the diff and the files it touches | Claude Haiku 4.5 | 120K | 1.5K | $0.128 |
| Judge correctness and design | Claude Opus 5 | 45K | 3.0K | $0.300 |
| Verify each finding against the code | Claude Sonnet 5 | 60K | 1.2K | $0.132 |
| Total (mixed) | $0.559 | |||
The whole thing on one model
| Model | Cost | |
|---|---|---|
| Claude Haiku 4.5 | $0.254 | Peak context unknown; fit not established |
| Claude Sonnet 5 | $0.507 | Peak context unknown; fit not established |
| Claude Opus 5 | $1.27 | Peak context unknown; fit not established |
| Claude Fable 5.1 | $2.54 | Peak context unknown; fit not established |
The modeled split costs $0.559 versus $1.27 on Claude Opus 5. Measure result quality separately. This estimate excludes caching, tool charges, and other billing options.
Reading the table
Step 1 dominates the token count and contributes least of the value. That is the whole finding. The reading is 120K cumulative input; the judging is 45K. Most of what you spend on review is spent before any reviewing happens.
The mix is no compromise. It does not review slightly worse for less money. Step 2 still runs on the most capable model available; the saving comes entirely from declining to pay that rate for reading files and re-checking line numbers.
Cumulative input, not diff size. A 500-line diff is maybe 20K tokens. Step 1 is 120K because an agentic loop re-sends the conversation every turn, and by turn eight the earlier reading is in the bundle six more times. This is the number people get wrong, and it is why "the diff is small" does not mean "the review is cheap".
Try this
Run your next review twice: once as you normally would, once with the reading step delegated to a cheap model or a subagent whose context is discarded afterwards. Compare the findings, not the cost. If the findings are the same, the saving is free.
What goes wrong
Reviewing in the session that wrote the code. The cheapest possible review and the least useful one — same context, same blind spot.
Skipping step 3. An unverified finding sends someone to rewrite working code. Verification is the cheapest step in the table.
Loading the whole repository "for context". Step 1 is already the expensive one. Widening it is the fastest way to double the cost of a review without improving it.
Treating every model choice as a quality decision. For step 2 it is. For steps 1 and 3 it is a budget decision, and conflating the two is how the budget goes.
How to check it worked
Compare the findings from a mixed run against a single-model run on the same diff. If the mixed run misses something, it will be in step 2 — and the fix is to raise the model on that step only, not on all three.
Related: Reviewing agent-written code for what to look for, and Choosing a model for the general rule this workflow is an instance of.
Sources
- Claude models overview — Anthropic documentation Tier 1 2026-08-31
- Claude pricing Tier 1 2026-08-31
- Effective context engineering for AI agents — Anthropic Tier 1 2026-08-31
Something wrong with this page?
Say what you expected and what you got. That is usually the shortest route to a correction, and it goes on the public issue tracker so the fix is visible.