Anthropic API

Migrating to a newer model

Anthropic API

A model swap is not a string change. Parameters break loudly, behaviour changes quietly, and the quiet half is where migrations actually fail.

Applies to
Claude Fable 5.1 Claude Opus 5 Claude Sonnet 5 Claude Haiku 4.5
Last verified
Reviewed by
Timothy Fehr

Changing the model ID feels like a config edit. Anthropic publishes a dedicated migration guide per model, with breaking changes, recommended changes and a checklist, which tells you how much more than a config edit it is. This page is the general shape, using the Opus 5 guide as the worked example; read the guide for your actual target, because the specifics move with every release.

The loud failures: parameters that now error

These fail fast, which makes them the easy half. From the guides this site has verified directly:

  • temperature, top_p, top_k: deprecated after Opus 4.6; non-default values return a 400. See why answers vary.
  • budget_tokens: the manual thinking budget errors on models after the 4.6 generation. The replacement is adaptive thinking with an effort level.
  • Combinations: on Opus 5, thinking: disabled with effort xhigh or max "returns a 400 error. Claude Opus 4.8 accepts this combination, so audit requests that disable thinking before you migrate".
  • Forced tool use: tool_choice any and tool are rejected on some models entirely; the guide's fallback is auto plus strict tool use.

A 400 in staging is a good outcome. The migrations that hurt are below.

The quiet failures: behaviour that changed

Nothing errors; the output is different. The Opus 5 guide documents four worth generalising from.

Defaults flip. On Opus 4.8, a request without a thinking field runs without thinking; the same request on Opus 5 runs with it. Same code, new billing profile: "a workload that ran without thinking on Claude Opus 4.8 can produce more output tokens per request on Claude Opus 5".

Response shapes move. With thinking on, replies can open with thinking blocks before the first text block. The guide is blunt about what that breaks: code reading "content[0].text" or treating the first stream event as text. Select blocks by type, never by position — that habit survives every future migration.

Prompts tuned for the old model misfire on the new one. The sharpest example: Opus 5 verifies its own work unprompted, so "remove explicit verification or self-check instructions carried over from prompts tuned for earlier models; leaving them in causes over-verification". Cost and latency for nothing. The fix is removal, which nobody's migration instinct suggests.

Adjacent limits shift. The minimum cacheable prompt dropped from 1,024 to 512 tokens on Opus 5, a free win, while the caching tiers in general do not follow generation order, so a working cache can silently stop on a different target.

The process that survives contact

  1. Read the target's migration guide end to end before touching code. The breaking list is short; the recommended list is where the quiet failures are named.
  2. Run your eval against the new model before switching traffic. A migration without a baseline is a vibe check, and felt quality diverges from measured.
  3. Compare cost per unit of work, not per token. Same per-token price with thinking on by default is a different bill.
  4. Re-read your prompt for instructions the new model makes redundant. Verification clauses, format nudges, length constraints — test with them removed.
  5. Keep the old model ID deployable until the eval and the bill both look right for a full cycle of real traffic.

In Claude Code, /claude-api migrate automates the mechanical half — the ID swap and parameter changes — and hands back "a checklist of items to verify manually". Use it for the swap; the eval is still yours.

What goes wrong

Treating it as find-and-replace. The ID swap is the only part that is.

Testing only that requests succeed. The quiet failures all return 200.

Carrying the whole prompt forward. Instructions written to compensate for an old model's weaknesses become active defects on the new one.

Ignoring the billing shape. Thinking-by-default changes output volume at the same per-token price; nobody notices until the invoice.

Migrating everything at once. One workload first, evaled and billed, then the rest, with the same unit discipline as any other risky change.

How to check it worked

Three numbers, before and after, same inputs: eval score, cost per unit of work, latency. If all three moved the way the release notes promised, the migration is done. If quality rose and cost doubled, you have a decision to make rather than a finished migration, and you can only make it because you measured.

Sources

  1. Migrating to Claude Opus 5 — Claude Platform Docs Tier 1 2026-09-04
  2. Extended thinking — Claude Platform Docs Tier 1 2026-09-04