Developer workflows

Workflow: a refactor across several files

Claude + Claude Code + Anthropic API

The plan is 1% of the tokens and decides the other 99%. Everything else is mechanical, and mechanical work does not need the expensive model.

Applies to
Claude Fable 5.1 Claude Opus 5 Claude Sonnet 5 Claude Haiku 4.5
Last verified
Reviewed by
Timothy Fehr

A refactor is the workflow where the cheapest step carries the most risk. The plan costs almost nothing and determines everything after it, which is an unusual shape and one worth designing around.

The steps

  1. 1 · Map the codeevery call site and test
  2. 2 · Write the plansmall, and expensive to get wrong
  3. 3 · Implementmechanical, longest, cheapest per token
  4. 4 · Verifysuite, diff, nothing unrelated moved

1. Map the affected code. Find every call site, every import, every test that touches the thing. Search, not thought — but a lot of it.

2. Write the plan. What changes, in what order, what must not change, what is out of scope. Small, and the only step where being wrong is expensive.

3. Implement. Mechanical once the plan exists. Long, and every turn re-sends what came before, which is why it is the largest line in the table.

4. Verify. Run the suite, read the diff, check nothing unrelated moved. Deterministic.

What it costs

Assumption: A change touching around fifteen files, planned, implemented and verified. Figures estimate a representative run, not yours.

An illustrative split

StepModelInputOutputCost
Map the affected codeClaude Haiku 4.5200K3.0K$0.215
Write the planClaude Opus 540K3.0K$0.275
ImplementClaude Sonnet 5250K20.0K$0.700
VerifyClaude Haiku 4.560K1.5K$0.068
Total (mixed)$1.26

The whole thing on one model

ModelCost
Claude Haiku 4.5$0.688Peak context unknown; fit not established
Claude Sonnet 5$1.38Peak context unknown; fit not established
Claude Opus 5$3.44Peak context unknown; fit not established
Claude Fable 5.1$6.88Peak context unknown; fit not established

The modeled split costs $1.26 versus $3.44 on Claude Opus 5. Measure result quality separately. This estimate excludes caching, tool charges, and other billing options.

Reading the table

The plan is 7% of the input and governs the rest. A wrong plan means steps 3 and 4 run twice — once building the wrong thing and once undoing it. That is the real cost of skipping step 2, and it does not appear in any table because it looks like "the refactor took two days".

Implementation dominates and is mechanical, at 250K cumulative input for work that is largely applying a decision already made. This is the step to move down a tier, and the one people most resist moving because it is where the code gets written.

Verification is the cheapest step and the one that gets skipped. It costs less than the mapping step and it is what stands between the refactor and a silently broken build.

Why this workflow rewards planning more than any other

In review, a bad decision costs one wrong finding. In debugging, a bad theory costs one wrong search. In a refactor, a bad plan costs every subsequent step, because each one builds on the last.

That asymmetry is the argument for stopping after step 2 and getting the plan approved before implementing. It is also why the plan should be a file rather than a message: step 3 will run long, and a plan that only exists in the conversation is a plan that degrades as the context fills.

Try this

Before your next multi-file change, write the plan and stop. Read it back and ask what would have to be true for step 3 to go wrong. Most of the answers are things you can check in five minutes with the mapping you already have — and each one you check is an implementation run you do not pay for twice.

What goes wrong

Implementing from the mapping, with no plan. Now the decisions are being made incrementally, in the most expensive step, with no record of why.

Running the whole workflow on one model. Two of four steps are mechanical and one is search.

Letting implementation drift past the plan. If a step turns out wrong, stop and revise the plan. Improvising forward is how a fifteen-file refactor becomes a forty-file one nobody agreed to.

Skipping verification because the tests passed during implementation. They passed on the state at that moment, several edits ago.

How to check it worked

Read the final diff against the plan's out-of-scope section. Files that changed and should not have are the finding, and nothing else will surface them — particularly not a summary written by the same session that made the change.

Related: planning-before-code and verifying-before-done in the skills library package steps 2 and 4 as agent skills.

Sources

  1. Claude models overview — Anthropic documentation Tier 1 2026-08-31
  2. Claude pricing Tier 1 2026-08-31
  3. Best practices for Claude Code Tier 1 2026-08-31