Foundations

Thinking, reasoning and "effort" — what these controls do

Anthropic API

Thinking is billed output you mostly never read. The control moved from a token budget to an effort level, and on current models the old one errors.

Applies to
Claude Fable 5.1 Claude Opus 5 Claude Sonnet 5 Claude Haiku 4.5
Last verified
Reviewed by
Timothy Fehr

Before answering, a model can work through the problem in tokens you usually do not see. Those tokens are generated, billed, and counted against the response limit like any others. That is all "thinking" means here — no separate faculty, just the model writing to itself first.

Which makes the useful questions concrete. How much of it happens, who decides, and what it costs.

Two generations of control, and one of them now errors

The older mechanism was a token budget: you set a number and the model thought against it before answering. The newer one is adaptive thinking with an effort level, where the model decides whether and how much to think per request and you set the depth.

This is a hard boundary, and no longer a matter of preference. The documentation is explicit that the manual budget is deprecated on the 4.6 generation and that later models "do not support it and reject requests that use it, returning a 400 error." Code written against a budget stops working on a newer model rather than degrading.

Which models sit where is recorded in the model matrix, because it changes with each release and belongs in one place.

The behavioural difference is the part that matters

Migrating is a small syntax change and a real change in behaviour. The docs put it plainly: "With a fixed budget, Claude thinks on every request. With adaptive thinking, Claude decides whether and how much to think on each request, and at lower effort settings it may skip thinking entirely on easy inputs."

So effort is a ceiling and a disposition rather than an instruction. Asking for high effort on a trivial question does not force deliberation; the model may still answer directly, which is the behaviour you want.

One number worth knowing: effort: "high" matches the API default. Setting it changes nothing. Most people who think they turned reasoning up have written the default out longhand.

What it costs

Thinking tokens are output tokens. They count against the response limit, and they appear on the bill.

The API reports them separately in usage.output_tokens_details.thinking_tokens, which is the honest way to find out what a given effort level is actually costing, in place of guessing from total spend.

Anthropic's guidance on budgets, where budgets still apply, describes "diminishing returns that depend on the task, and at the cost of increased latency." More thinking helps until it stops helping, and the point where it stops is task-specific enough that they tell you to test rather than assume.

The caching trap

Changing the thinking configuration invalidates prompt cache breakpoints, because the value is rendered into the prompt. Their advice: "pick a budget and hold it stable for the life of a cached conversation." The same applies to effort in adaptive mode.

The failure is quiet and expensive. Tuning effort per request in a workload you built caching around means paying full price for the cached prefix every time, and nothing surfaces that except the usage numbers.

When more thinking earns its keep

Multi-step reasoning where a wrong early step invalidates everything after it: maths, proofs, constraint problems, debugging where symptom and cause are far apart. Anything you would want a person to work through on paper.

Where it does not: retrieval, formatting, summarising, classification, and anything whose difficulty lies in having the right context and never in the reasoning. Those are context problems, and more deliberation over the wrong material buys nothing.

Try this

If you use the API, run the same prompt at two effort levels and compare thinking_tokens against whether the answer actually improved. Do it on a task you can score.

Most people find one of two things: a task where the difference is dramatic, or a task where it is invisible and they had been paying for it. Both are worth knowing, and neither is predictable from the outside.

What goes wrong

Setting effort: "high" and believing something changed. It is the default.

Reading thinking as a transcript of reasoning. What surfaces is a summarised, sometimes encrypted artefact. It shows the model worked; it is not an audit trail of how.

Assuming more thinking means more correct. It buys deliberation, and deliberation over insufficient context reaches a confident wrong answer more thoroughly.

Carrying a budget forward to a new model. On current models that request returns a 400 rather than quietly doing something sensible, which is the better failure but still a surprise mid-migration.

Tuning effort inside a cached workload. Every change re-creates the cache.

How to check it worked

Compare answer quality against thinking_tokens on a task you can score, at two settings. If quality is flat while thinking tokens rose, you found the diminishing-returns point for that task and the lower setting is the right one. If it improved, you have a number to justify the cost — which is a stronger position than either guessing or leaving the default in place because nobody measured.

Sources

  1. Extended thinking — Claude Platform Docs Tier 1 2026-09-04
  2. Create a Message — Claude API reference Tier 1 2026-09-04