Anthropic API

Prompt caching that actually hits

Anthropic API

Reads cost a tenth of normal input, so caching pays from the second request. A prompt under the minimum is silently not cached, with no error.

Applies to
Claude Fable 5.1 Claude Opus 5 Claude Sonnet 5 Claude Haiku 4.5
Last verified
Reviewed by
Timothy Fehr

Caching a prompt prefix means paying to store it once and a tenth of the normal input price every time it is reused. On a workload with a large stable prefix — a system prompt, tool definitions, a document you ask many questions about — that is the difference between a viable product and an expensive one.

It is also easy to configure in a way that never hits, and the failure is silent.

The arithmetic, so you know when it pays

Three prices, all multiples of the base input rate:

Multiplier
5-minute cache write1.25×
1-hour cache write2×
Cache read or refresh0.1×

Work it through. With the 5-minute TTL you pay 1.25 on the first request and 0.1 on each reuse, against 1.0 per request uncached. You are ahead from the second request. With the 1-hour TTL the write costs 2, so you break even during the third.

That is a low bar, which is why caching pays for almost any repeated prefix. The bar people actually fail is the one below.

The minimum, and the silence

A prompt shorter than the model's minimum is not cached. The documentation is explicit about what happens next: "Any requests to cache fewer than this number of tokens will be processed without caching, and no error is returned."

Nothing tells you. cache_control is accepted, the request succeeds, and you pay full price forever.

The minimum varies by model across four tiers: 512, 1,024, 2,048 and 4,096 tokens. It does not track how new or capable the model is. Check the platform documentation for your model rather than assuming, because a prefix that caches on one model silently stops when you switch to another.

Order decides what can be cached

Prefixes are built in a fixed order: tools, then system, then messages. Each level builds on the ones before it.

That makes the hierarchy a constraint on layout rather than a detail. Anything that changes must sit after everything you want cached. Put a per-request value in the system prompt and you have invalidated the system and message caches on every call — the tools cache survives, and it is usually the smallest part.

You get up to four breakpoints. With automatic caching on, the automatic one consumes a slot.

What throws the cache away

Changes cascade downward. Editing tool definitions invalidates tools, system and messages. The list below matters because several of them are things people tune casually:

  • Tool definitions: invalidates everything.
  • Thinking parameters and the effort setting: always invalidate the message cache. This is why tuning effort per request inside a cached workload is quietly expensive.
  • Toggling web search or citations: invalidates system and messages.
  • Adding or removing images: invalidates messages.
  • tool_choice: invalidates messages only.

The pattern: pick a configuration and hold it for the life of the cached conversation. Caching rewards boring.

The TTL choice

Five minutes is the default and costs nothing extra. One hour costs 2× on the write and suits a prefix reused across a longer session.

One detail that catches people: the lifetime is measured "from the start of the request that writes or reads the cache entry, not from the end of its response." On a long-running request, the clock has been going the whole time.

Try this

Send the same large-prefix request twice and read the usage block on both. The first should show cache_creation_input_tokens and the second cache_read_input_tokens. If the second shows neither, you are under the minimum or something between the two calls invalidated it.

What goes wrong

Assuming it worked because nothing failed. The silent no-cache path is the default failure, and it costs money indefinitely without a symptom.

Putting a timestamp or a user id early in the prompt. Anything variable before the cached content invalidates everything after it.

Changing the model without rechecking the minimum. The tiers do not follow generation order, so a migration can turn a working cache off.

Tuning sampling or effort per request. Each change re-creates the cache, and the write costs more than the uncached request would have.

How to check it worked

Compare a day of cache_read_input_tokens against input_tokens for the same workload. A healthy cached workload is mostly reads. If reads are a small fraction, something is invalidating between calls — and the cascade order tells you where to look, because the highest level that changed is the one that took everything below it with it.

Sources

  1. Prompt caching — Claude Platform Docs Tier 1 2026-09-04
  2. Create a Message — Claude API reference Tier 1 2026-09-04