Reference

Compare API cost and context correctly

Anthropic API + OpenAI API + Gemini API

Separate cumulative usage from peak context, account for the chosen pricing conditions, and measure quality independently.

Applies to
Anthropic API + OpenAI API + Gemini API
Last verified
Reviewed by
Timothy Fehr

Compare a complete, verified task under a stated access and pricing profile. A low token rate is useful only when the model and workflow produce an acceptable result.

The provider catalogs list selected API profiles: Anthropic, OpenAI, and Google. Each record has its own source and date.

Keep two different token measurements

Cumulative input is the sum of billed input over requests. Peak context is the largest context held for one request, counted according to that API's rules.

Consider three requests with 100,000 input tokens each. Cumulative input is 300,000. That does not establish that any single request needed a 300,000-token window. Conversely, a small cumulative estimate cannot prove that a large tool result plus the requested output will fit a particular request.

Caching, retrieval, truncation, and compaction change what is used or billed. Record the actual API's input, output, cached-input, and reasoning usage where available. Preserve unknown fields as unknown.

State the price profile

Before calculating, name the API provider, exact model ID, processing tier, input length band, modality, caching behavior, and verification date. Account for separately priced tools such as search or execution.

A simplified uncached text estimate is:

input tokens × input rate / 1,000,000
+ output tokens × output rate / 1,000,000

That formula is not a universal invoice. Long-context rates, cache writes and reads, regional processing, discounts, and tool charges can change the result. Use the relevant provider's current pricing terms.

The existing Claude workflow tables are illustrative uncached Anthropic estimates. They use cumulative usage for cost and mark context fit unknown when peak per-request context was not recorded. They do not establish savings without a quality change.

Measure a complete outcome

Keep the task, corpus, acceptance tests, and review criteria consistent. Record confirmed defects, false findings, corrections, human review time, elapsed time, and reported usage. Count failed attempts and repair loops.

For a pipeline, compare at least a manual or single-agent baseline and the proposed combination. A cheap extraction stage that loses a requirement can increase the total implementation and review cost.

Reasoning settings are provider-specific controls. Identically named levels are not calibrated units that make different models directly equivalent.

What goes wrong

Applying API rates to a subscription quota produces an invented conversion. Summing per-step input and comparing it to a context window confuses usage with capacity. Treating a promotional rate as permanent creates a stale budget. A provider switch can also change data recipients and access policy.

How to check

Recalculate one small task from its returned usage and the cited pricing profile. Compare the result with the provider's usage report and explain any missing charge categories. If peak context was not measured, do not label the model as fitting or overflowing.

The pipeline recovery guide distinguishes enforceable dispatch limits from incomplete spend estimates.

Sources

  1. OpenAI: API pricing Tier 1 2026-09-08
  2. Anthropic: API pricing Tier 1 2026-09-08
  3. Google: API pricing Tier 1 2026-09-08