Anthropic API

Error handling that distinguishes retryable from not

Anthropic API

A 529 deserves a retry, a 400 never earns one. Sorting the API's errors into those two bins is most of production error handling.

Applies to
Claude Fable 5.1 Claude Opus 5 Claude Sonnet 5 Claude Haiku 4.5
Last verified
Reviewed by
Timothy Fehr

Retrying everything melts your rate limits and hammers a request that can never succeed. Retrying nothing turns every transient blip into a user-facing failure. The whole discipline is sorting errors into two bins, and the API's status codes do most of the sorting for you.

The two bins

Retry, with backoff: transient states that resolve on their own.

  • 529 overloaded_error: "The API is temporarily overloaded", a condition of "high traffic across all users", and nothing about your request.
  • 500 api_error: the documentation says it directly: "Retry the request with exponential backoff; if the error persists, contact support" with the request ID.
  • 429 rate_limit_error: retryable with a caveat below.
  • Connection failures and timeouts: the network's problem before it is the request's.

Never retry: the request itself is wrong, and it will be exactly as wrong the fifth time.

  • 400: malformed request, an unsupported parameter for the model, or a spend limit you set. Model migrations produce these in clusters.
  • 401 / 402 / 403: credentials, billing, permissions. Fix the account; no loop fixes these.
  • 404 / 413: wrong resource, or a request over the size limit (32 MB for Messages).

The SDKs already implement the right default: they "automatically retry transient failures (such as connection errors, rate limits, and 5xx server errors) with exponential backoff, twice by default, honoring the retry-after header when present". Before writing retry logic, check whether you are duplicating theirs.

Two 429s wear the same code

The caveat that breaks naive retry loops. An ordinary rate-limit 429 comes with retry-after and resolves. But a spend-cap 429 "has no retry-after header and keeps failing until access resumes". Retrying it is a loop that burns requests until the billing period turns over.

The header is the discriminator: retry-after present, wait and retry; absent, stop and alert a human, because the fix is money and never time.

Errors after the 200

A streaming response can fail after the API returns success: "an error can occur after the API returns a 200 response. In that case, error handling doesn't follow these standard mechanisms". It arrives as an error event inside the stream. A handler that only checks the status code calls a half-delivered response a success.

The related trap is the long non-streaming request. Networks "may drop idle connections after a variable period of time", so a large max_tokens without streaming fails in ways that look like API errors and are not. The documented remedies: streaming for anything long, the Batches API for anything over ten minutes, and TCP keep-alive for direct integrations. The SDKs refuse non-streaming requests they expect to exceed ten minutes, which is a guardrail, whatever it feels like the first time it fires.

Handle types, and keep the request ID

Two habits from the documentation that pay off in the first incident.

Catch the SDK's typed exceptions "rather than string-matching error messages", because the error type values are versioned to grow and a string match is a slow-motion breakage.

And log the request-id header on failures. It is the handle support needs, and the difference between "it failed sometimes yesterday" and a ticket that can be acted on.

What goes wrong

Retrying a 400. It cannot succeed, and after a model migration a retried 400 storm is the classic symptom.

Retrying a capless 429. No retry-after means the spend cap; the loop resolves nothing and spends your remaining goodwill with the limiter.

Trusting the 200 on a stream. The error event comes later, inside the stream, on its own schedule.

Retrying without backoff or jitter. Synchronized immediate retries are how a transient 529 becomes a longer one for everybody, yourself included.

Blind retries around tool side effects. If the request triggered a tool that acted, a retry can act twice. Idempotency is part of error handling, inseparable from it.

How to check it worked

Test both bins deliberately: send a malformed request and confirm your system fails it immediately without retries; then simulate a 529 (or rely on the SDK's retry counter) and confirm it retries with growing delays and gives up at the configured ceiling. A handler that has never been watched doing both is an assumption — and the middle of an outage is the expensive place to test it.

Sources

  1. Claude API errors — Claude Platform Docs Tier 1 2026-09-04
  2. Batch processing — Claude Platform Docs Tier 1 2026-09-04