Batching and cost control
Anthropic API
This page covers tools outside your selection. You can still read it. Find matching guides
Half price for work that can wait an hour. The catch is a hard 24-hour expiry, and it is the only cost lever that needs no prompt changes.
The Message Batches API charges 50% of standard prices for every request in a batch. Not a smaller model, not a shorter prompt, not a worse answer. The same request, submitted asynchronously.
That makes it the cheapest cost lever available, and the one most often left on the table — because it requires noticing which of your work does not actually need an answer this second.
What it costs and what it demands
"All usage is charged at 50% of the standard API prices." Both input and output, on every model.
In exchange you give up immediacy. Most batches complete in under an hour, and the real constraint is the ceiling: results become available when everything finishes or after 24 hours, whichever comes first, and "Batches expire if processing does not complete within 24 hours."
Expired requests are not billed. So the failure mode is missing output rather than a surprise invoice, which is the better of the two.
The limits worth knowing before you design around it
The request-count, size and completion limits are in the batch processing documentation and move with the product, so they are not copied here. One of them shapes a design: results are retained for 29 days from creation. After that the batch is still visible and its results are no longer downloadable. If the output matters, store it somewhere you control rather than treating the API as your archive.
Two smaller things that bite. Each request needs max_tokens of at least 1, so cache pre-warming with max_tokens: 0 is not available inside a batch. And because batches run with high concurrency, they "may go slightly over your Workspace's configured spend limit" — the limit is not a hard stop here.
Batching and caching are different levers
Separate them, because they get discussed together and solve opposite problems.
Caching cuts the cost of a repeated prefix and needs the requests close together in time. Batching cuts the cost of everything and needs the requests to tolerate delay.
A large batch of requests sharing a system prompt can use both, subject to the max_tokens rule above. Where you can only have one, batching is the simpler win because it changes no prompt structure at all.
Try this
Take one recurring job — an eval suite, a nightly classification, a backfill — and move it to the Batches API without changing anything else. Compare the bill for that job week on week.
The saving is exactly half, which makes it one of the few changes where you can predict the result before running it.
What goes wrong
Batching something latency-sensitive. The 24-hour expiry is real, and under load "you may see more requests expiring." Anything a person is waiting on belongs on the synchronous API.
Treating results as durable. They are downloadable for 29 days.
Assuming the spend limit holds. Concurrency means a batch can overshoot it.
Splitting badly. One batch of 100,000 is a single expiry risk; many small batches are more requests to track. Size for the failure you would rather have.
How to check it worked
Compare cost per unit of work before and after, on the same job. If it did not land near half, something is running outside the batch — a synchronous retry path, or a pre-processing call you forgot was there. That gap is worth finding, because it is usually a request that could have been batched too.
Sources
- Batch processing — Claude Platform Docs Tier 1 2026-09-04
- Prompt caching — Claude Platform Docs Tier 1 2026-09-04
Something wrong with this page?
Say what you expected and what you got. That is usually the shortest route to a correction, and it goes on the public issue tracker so the fix is visible.