Large refactors: scoping work an agent can hold
Claude Code
This page covers tools outside your selection. You can still read it. Find matching guides
The unit of work is however much fits in one window with room to think. Size for that, and the refactor becomes a series of small verified changes.
A refactor across forty files fails for a reason that has nothing to do with difficulty. Each file read costs 1,100–2,400 tokens, so by file fifteen the plan agreed at the start is competing for attention with everything read since. The change is not hard. It is too big to hold.
Which model for which step is a separate question with its own page. This one is about size.
The unit is what fits with room left over
A good unit of work is one the agent can read, change and verify inside a single window without approaching the limit. That is smaller than most people choose, and the tell is retrospective: when a session starts reintroducing an approach you rejected earlier, the unit was too big.
Two practical shapes, both effective for the same reason.
Split by layer. All the call sites, then all the tests, then the docs. Each pass is mechanically similar, so the context needed stays flat.
Split by module. One subsystem end to end, verified, committed, cleared. The next one starts fresh with only the plan and the interface it must honour.
Map first, in a subagent
Finding every call site is search, and search is the most expensive thing you can do in your main window. It reads many files and keeps almost none of the value.
Delegate it. A subagent reads in its own window and returns a summary — in Anthropic's published walkthrough, three files read for 420 tokens returned. You get the map without paying for the reading.
The map is what makes the next decision possible: with every site listed, the split into units stops being a guess.
Write the plan to a file, not into the conversation
The plan is the thing that must survive the whole refactor, and a conversation is exactly where it will not.
Put it in a file the agent can re-read. Anthropic's advice on specs applies directly: the useful ones "are self-contained: they name the files and interfaces involved, state what is out of scope, and end with an end-to-end verification step that proves the feature works."
Then each unit begins by reading the plan rather than by remembering it, which is the difference between a constraint that holds at step nine and one that faded around step four.
Verify and commit between units
Every unit ends with the check running and a commit landing. Two reasons, and the second is the one people miss.
The obvious one: a failure is attributable to the unit that caused it.
The other: /clear is only safe when the work is committed. Without a commit you cannot reset the context, so you carry a full window into the next unit and the problem compounds. Committing is what buys you a clean start. See git hygiene, and note that checkpoints do not cover changes made by bash commands, which a refactor produces constantly.
Try this
Take a refactor you would have run as one instruction. Ask for the map only, in a subagent, and have it list the affected files. Then pick the smallest coherent subset, do that alone, verify, commit, clear.
Compare the diff quality against the last large change you ran end to end.
What goes wrong
One instruction for the whole thing. It starts well. The failure appears around the point where the early context stopped being attended to, which is also the point where the diff got too large to review.
Mapping in the main window. Thirty file reads spend the context the work needed.
Keeping the plan in the conversation. It is competing with everything read since. A file does not compete.
Not committing between units. Without a commit you cannot clear, and without clearing every unit inherits the last one's noise.
Reviewing only at the end. A forty-file diff gets approved rather than read, which is automation bias with a deadline attached.
How to check it worked
Count the units where the check passed first time. A high proportion means the units were the right size and the plan was carrying. A low one usually means units were too large rather than that the work was hard — try halving the next one and see whether the rate moves before concluding anything about the codebase.
Sources
- Best practices for Claude Code — Anthropic Tier 1 2026-09-04
- Explore the context window — Claude Code Docs Tier 1 2026-09-04
Something wrong with this page?
Say what you expected and what you got. That is usually the shortest route to a correction, and it goes on the public issue tracker so the fix is visible.