Anthropic API

Context editing, compaction, or memory

Anthropic API

Three mechanisms for a conversation that outgrows its window. One clears, one summarises, one stores — and they compose rather than compete.

Applies to
Claude Fable 5.1 Claude Opus 5 Claude Sonnet 5 Claude Haiku 4.5
Last verified
Reviewed by
Timothy Fehr

A long-running conversation eventually outgrows its window, and the platform offers three mechanisms with names close enough to blur. They do different things, and the choice is simpler once each is reduced to its verb.

Context editing clears. Compaction summarises. Memory stores.

Where each one runs, and what that means

Context editing is applied "server-side before the prompt reaches Claude", so "your client application maintains the full, unmodified conversation history". You keep everything; the model sees less. Content is removed, never rewritten.

Compaction replaces history with a generated summary. Lossy by design, and the loss is chosen by a model rather than by a rule. The documentation's current position: "For most use cases, server-side compaction is the primary strategy for managing context in long-running conversations. The strategies on this page are useful for specific scenarios where you need more fine-grained control over what content is cleared".

The memory tool gives Claude persistent storage alongside the conversation. It does not shrink the window at all — it is where information goes so that clearing it from the window is safe. The documented pairing: "Claude can write essential information from tool results to memory files before those results are cleared".

The composition follows from the verbs: store what must survive, summarise what only needs its gist, clear what has been consumed.

The clearing strategy, precisely

clear_tool_uses targets the material that dominates agentic conversations: "Older tool results (like file contents or search results) are no longer needed once Claude has processed them".

When the configured trigger is crossed, the API "automatically clears the oldest tool results in chronological order" and "replaces each cleared result with placeholder text indicating to Claude that it was removed". The placeholder matters: the model knows something was there, so it does not reason as if the tool was never called.

The knobs, with defaults:

ParameterDefaultWhat it does
trigger100,000 input tokenswhen clearing starts
keep3 tool usesrecent pairs preserved
clear_at_leastnoneminimum cleared per activation
exclude_toolsnonetools never cleared

The response reports what happened in applied_edits (how many tool uses and tokens were cleared), so the behaviour is observable rather than assumed.

The caching interaction is the design constraint

Clearing changes the prompt, and a changed prompt invalidates cached prefixes from the edit point. The documentation states both the cost and the mitigation: "You'll incur cache write costs each time content is cleared, but subsequent requests can reuse the newly cached prefix", and the advice is to "clear enough tokens to make the cache invalidation worthwhile", which is what clear_at_least exists for.

Frequent small clears are the failure shape: each one pays a cache rewrite for little relief. Rare large clears amortise.

Choosing, by workload

Tool-heavy agent, long sessions: clearing first. The bulk is spent tool results, and a rule-based removal is predictable in a way a summary is not. exclude_tools protects the results that stay load-bearing.

Long dialogue where the past matters as gist: compaction. What an early exchange established matters; its wording does not.

Anything that must survive verbatim across the whole run — decisions, constraints, identifiers: memory, written before the window pressure arrives. This is the API-side version of the plan lives in a file, never in the conversation.

The app-side intuition transfers directly: clearing is /clear with a scalpel, compaction is /compact, memory is the store — the same budget logic driving all three.

What goes wrong

Summarising what should have been stored. A compacted constraint is a paraphrase of a constraint. If the wording carries the meaning, it belongs in memory before compaction runs.

Clearing on a hair trigger. Every activation pays a cache rewrite. Size clear_at_least so each clear buys real headroom.

Excluding nothing. Some tool results stay load-bearing for the whole run — the file being edited, the spec being followed. exclude_tools is the list of things the placeholder must never replace.

Treating the three as rivals. They run at different layers and compose. The documented pattern is memory alongside editing, and compaction over both.

How to check it worked

Read applied_edits on responses after the trigger point, and compare eval results on long conversations with and without the configuration. Cleared tokens with unchanged task quality is the mechanism working. Quality dropping after clears means something load-bearing is being removed — and the fix is exclude_tools or an earlier memory write, which the next long run will test.

Sources

  1. Context editing — Claude Platform Docs Tier 1 2026-09-04
  2. Prompt caching — Claude Platform Docs Tier 1 2026-09-04