Developer workflows

Workflow: debugging a failure

Claude + Claude Code + Anthropic API

Narrowing costs four times what diagnosing costs. Put the cheap model on the search and the expensive one on the thirty seconds that decide the answer.

Applies to
Claude Fable 5.1 Claude Opus 5 Claude Sonnet 5 Claude Haiku 4.5
Last verified
Reviewed by
Timothy Fehr

Debugging has a distinctive cost shape: almost all of it is searching, and almost none of the value is. The moment that solves the bug is short, and it is surrounded by a great deal of reading.

The steps

  1. 1 · Reproducecapture the failure, deterministic
  2. 2 · Narrowbisect — most tokens go here
  3. 3 · Diagnoseshort, hard — best model earns it
  4. 4 · Fix and provesmall, well specified

1. Reproduce. Run the thing, capture the failure. Deterministic. If you cannot reproduce it you cannot know you fixed it, so this is not optional — but it is not thinking either.

2. Narrow. Bisect in time, in space, in data. This is where the tokens go: repeated reading as the search space halves, every round re-sent on top of the last.

3. Form and confirm the cause. The actual diagnosis. Short, hard, and the only step where a better model reliably produces a better answer.

4. Fix and prove it. Once the cause is known the change is usually small and well specified.

What it costs

Assumption: One reproducible failure in a medium codebase, found over roughly twelve turns. Figures estimate a representative run, not yours.

An illustrative split

StepModelInputOutputCost
Reproduce and capture the failureClaude Haiku 4.525K0.8K$0.029
Narrow the search spaceClaude Sonnet 5180K2.0K$0.380
Form and confirm the causeClaude Opus 540K2.5K$0.263
Fix and prove the fixClaude Sonnet 530K2.0K$0.080
Total (mixed)$0.751

The whole thing on one model

ModelCost
Claude Haiku 4.5$0.311Peak context unknown; fit not established
Claude Sonnet 5$0.623Peak context unknown; fit not established
Claude Opus 5$1.56Peak context unknown; fit not established
Claude Fable 5.1$3.12Peak context unknown; fit not established

The modeled split costs $0.751 versus $1.56 on Claude Opus 5. Measure result quality separately. This estimate excludes caching, tool charges, and other billing options.

Reading the table

Narrowing is the bill: 180K cumulative input against 40K for the diagnosis itself. The step that feels like progress is the step that costs, and the step that solves the problem is nearly free.

Which means the leverage is in searching less, not in thinking cheaper. Every bisection that halves the space halves this step. Two good bisections beat any model swap in the table.

The diagnosis step is where not to economise. It is 40K of 275K — a small share of the total. Downgrading it saves little and costs the answer.

The failure this workflow prevents

Guessing a fix and applying it. It is fast, it sometimes works, and when it works it is worse than failing — because the real cause is still there and now has an alibi.

Cost-wise, guessing looks cheap and is not: a wrong fix is paid for twice, once to apply and once to undo, plus the second debugging session when the symptom returns somewhere else.

Try this

On the next bug, timebox narrowing before theorising. Two minutes of bisection first. You will usually reach the same place with evidence instead of a hypothesis, and the token count will be a fraction — because you stopped reading the moment the search space collapsed.

What goes wrong

Theorising before narrowing. The theory sends you reading in one direction, which is the expensive direction if it is wrong.

Reading whole files during the search. grep to locate, then read the range. Step 2 is already the largest line in the table.

Skipping reproduction because the cause seems obvious. A fix for an unreproduced bug is a hypothesis you have shipped.

Changing several things at once. The symptom goes away and you have learned nothing, so the next occurrence starts from zero.

How to check it worked

Remove the fix. The bug should come back. If it does not, you changed something else as well and the cause is still unidentified — which is the one outcome that looks exactly like success.

Related: systematic-debugging in the skills library packages this as an agent skill.

Sources

  1. Claude models overview — Anthropic documentation Tier 1 2026-08-31
  2. Claude pricing Tier 1 2026-08-31
  3. Best practices for Claude Code Tier 1 2026-08-31