Foundations

Tokens: why cost and quota do not track word count

Shared methods · A shared method; linked tool guides explain the exact steps.

The unit everything is billed and rationed in. It is not words, and the gap between the two is where people's mental arithmetic goes wrong.

Applies to
Shared methods
Last verified
Reviewed by
Timothy Fehr

Everything you send and everything you get back is counted in tokens. Not words, not characters, not messages. If your intuition about cost or quota is built on any of those three, it will be wrong in ways that matter.

A token is a chunk of text. Common English words are often one token each; longer or unusual words split into several. Code, punctuation, non-English text and long identifiers all cost more tokens per visible character than ordinary prose.

Two things this explains immediately

Why a short question can be expensive. The question is not what gets counted. On every turn the model re-reads the entire conversation: the system instructions, every earlier message, every attached file, every tool result. A three-word follow-up in a long conversation with four PDFs attached costs all of that again. The three words are a rounding error.

Why non-English costs more. German compound nouns, accented characters, and most non-Latin scripts fragment into more tokens per word than English does. The same document translated will not cost the same to process.

Input and output are not priced alike

Output tokens cost several times more than input tokens across every current model. On the API that is a line item you can read off the model matrix. On a subscription it is invisible but still real: asking for a full rewritten document consumes more of your allowance than asking which three paragraphs need changing.

This is the cheapest optimisation available and almost nobody makes it. Ask for the diff, not the document.

Try this

Take a conversation you consider typical and count what is actually in the bundle: how many earlier turns, how many attached files, how large those files are. Then ask yourself how many of them the current question needs. The difference between those two numbers is your recurring waste, and it is being paid on every single turn.

What goes wrong

Estimating from message count. "I only sent twenty messages" says nothing. Twenty short messages in a fresh conversation and twenty in one carrying a 50-page PDF differ by orders of magnitude.

Assuming attachments are one-off costs. They are not. An attached file is part of the bundle for the rest of the conversation.

Optimising the prompt while ignoring the history. People agonise over phrasing a question in twenty words instead of thirty while re-sending thousands of tokens of stale context.

Asking for whole documents back. Output is the expensive direction. Ask for what changed.

How to check it worked

Watch Settings → Usage across two days of deliberately shorter conversations with fewer attachments. If consumption drops noticeably, context habits were your cost driver. If it does not, the work itself is genuinely large and the answer is a different plan or a smaller model, not better discipline.

Sources

  1. Effective context engineering for AI agents — Anthropic Tier 1 2026-08-31
  2. What is the Max plan? — Anthropic Help Center Tier 1 2026-08-31