Context editing, compaction, or memory
Anthropic API
This page covers tools outside your selection. You can still read it. Find matching guides
Three mechanisms for a conversation that outgrows its window. One clears, one summarises, one stores — and they compose rather than compete.
A long-running conversation eventually outgrows its window, and the platform offers three mechanisms with names close enough to blur. They do different things, and the choice is simpler once each is reduced to its verb.
Context editing clears. Compaction summarises. Memory stores.
Where each one runs, and what that means
Context editing is applied "server-side before the prompt reaches Claude", so "your client application maintains the full, unmodified conversation history". You keep everything; the model sees less. Content is removed, never rewritten.
Compaction replaces history with a generated summary. Lossy by design, and the loss is chosen by a model rather than by a rule. The documentation's current position: "For most use cases, server-side compaction is the primary strategy for managing context in long-running conversations. The strategies on this page are useful for specific scenarios where you need more fine-grained control over what content is cleared".
The memory tool gives Claude persistent storage alongside the conversation. It does not shrink the window at all — it is where information goes so that clearing it from the window is safe. The documented pairing: "Claude can write essential information from tool results to memory files before those results are cleared".
The composition follows from the verbs: store what must survive, summarise what only needs its gist, clear what has been consumed.
The clearing strategy, precisely
clear_tool_uses targets the material that dominates agentic conversations: "Older tool results (like file contents or search results) are no longer needed once Claude has processed them".
When the configured trigger is crossed, the API "automatically clears the oldest tool results in chronological order" and "replaces each cleared result with placeholder text indicating to Claude that it was removed". The placeholder matters: the model knows something was there, so it does not reason as if the tool was never called.
The knobs, with defaults:
| Parameter | Default | What it does |
|---|---|---|
trigger | 100,000 input tokens | when clearing starts |
keep | 3 tool uses | recent pairs preserved |
clear_at_least | none | minimum cleared per activation |
exclude_tools | none | tools never cleared |
The response reports what happened in applied_edits (how many tool uses and tokens were cleared), so the behaviour is observable rather than assumed.
The caching interaction is the design constraint
Clearing changes the prompt, and a changed prompt invalidates cached prefixes from the edit point. The documentation states both the cost and the mitigation: "You'll incur cache write costs each time content is cleared, but subsequent requests can reuse the newly cached prefix", and the advice is to "clear enough tokens to make the cache invalidation worthwhile", which is what clear_at_least exists for.
Frequent small clears are the failure shape: each one pays a cache rewrite for little relief. Rare large clears amortise.
Choosing, by workload
Tool-heavy agent, long sessions: clearing first. The bulk is spent tool results, and a rule-based removal is predictable in a way a summary is not. exclude_tools protects the results that stay load-bearing.
Long dialogue where the past matters as gist: compaction. What an early exchange established matters; its wording does not.
Anything that must survive verbatim across the whole run — decisions, constraints, identifiers: memory, written before the window pressure arrives. This is the API-side version of the plan lives in a file, never in the conversation.
The app-side intuition transfers directly: clearing is /clear with a scalpel, compaction is /compact, memory is the store — the same budget logic driving all three.
What goes wrong
Summarising what should have been stored. A compacted constraint is a paraphrase of a constraint. If the wording carries the meaning, it belongs in memory before compaction runs.
Clearing on a hair trigger. Every activation pays a cache rewrite. Size clear_at_least so each clear buys real headroom.
Excluding nothing. Some tool results stay load-bearing for the whole run — the file being edited, the spec being followed. exclude_tools is the list of things the placeholder must never replace.
Treating the three as rivals. They run at different layers and compose. The documented pattern is memory alongside editing, and compaction over both.
How to check it worked
Read applied_edits on responses after the trigger point, and compare eval results on long conversations with and without the configuration. Cleared tokens with unchanged task quality is the mechanism working. Quality dropping after clears means something load-bearing is being removed — and the fix is exclude_tools or an earlier memory write, which the next long run will test.
Sources
- Context editing — Claude Platform Docs Tier 1 2026-09-04
- Prompt caching — Claude Platform Docs Tier 1 2026-09-04
Something wrong with this page?
Say what you expected and what you got. That is usually the shortest route to a correction, and it goes on the public issue tracker so the fix is visible.