Developer workflows

Workflow: researching and synthesising

Claude + Claude Code + Anthropic API

Reading thirty sources costs eight times what thinking about them costs. This is the workflow where model choice saves the most.

Applies to
Claude Fable 5.1 Claude Opus 5 Claude Sonnet 5 Claude Haiku 4.5
Last verified
Reviewed by
Timothy Fehr

Of the four workflows here, this is the one where splitting by model saves the most, because the ratio between reading and thinking is the most extreme.

The steps

  1. 1 · Read the sourceshuge input, trivial output
  2. 2 · Cross-referencethe judgement — conflicts, unsourced claims
  3. 3 · Synthesisequality is the whole deliverable

1. Read the sources. Thirty documents in, a few lines of notes out per document. Enormous input, trivial output, no judgement beyond "what does this say".

2. Cross-reference and resolve conflicts. Noticing that two sources disagree, working out which one governs, spotting the claim that everybody repeats and nobody sourced. This is the judgement.

3. Write the synthesis. The only step where output tokens matter, and the only one where quality is the entire deliverable.

What it costs

Assumption: Around thirty sources read, reduced to one sourced summary. Figures estimate a representative run, not yours.

An illustrative split

StepModelInputOutputCost
Read the sourcesClaude Haiku 4.5400K6.0K$0.430
Cross-reference and resolve conflictsClaude Opus 550K3.0K$0.325
Write the synthesisClaude Opus 530K8.0K$0.350
Total (mixed)$1.10

The whole thing on one model

ModelCost
Claude Haiku 4.5$0.565Peak context unknown; fit not established
Claude Sonnet 5$1.13Peak context unknown; fit not established
Claude Opus 5$2.83Peak context unknown; fit not established
Claude Fable 5.1$5.65Peak context unknown; fit not established

The modeled split costs $1.10 versus $2.83 on Claude Opus 5. Measure result quality separately. This estimate excludes caching, tool charges, and other billing options.

Reading the table

Step 1 is 83% of the input and almost none of the value. Nothing else in these workflow guides has a ratio this lopsided, which is why this is the workflow to fix first if you are fixing one.

Step 3 inverts the usual shape. Output tokens cost several times input tokens on every model, and this is the one step where output is large. It is also the step where downgrading is most visible — a synthesis is read by people, and a cheap synthesis reads cheap.

So the split is: cheap to read, expensive to think, expensive to write. Two of three steps stay on the good model and the workflow still costs a fraction, because the step you moved was the enormous one.

Delegate the reading, do not just downgrade it

The better version of step 1 is not "run it on a cheaper model". It is give it to a subagent whose context is discarded afterwards.

Downgrading the model reduces the rate. Delegating removes the tokens from the parent session entirely — the 400K never enters the conversation that does steps 2 and 3, so it is not re-sent on every subsequent turn either. On a long research session that compounds into more saving than the model swap.

The parent gets the notes. The documents stay in the subagent.

The part cost cannot tell you

Cheap reading is only worth having if the notes are faithful. A summariser that silently drops the qualification on a claim has saved you money and cost you the accuracy the whole exercise was for.

Ask for quotes and locations rather than paraphrase in step 1. It costs a little more output and it is the difference between a synthesis you can defend and one you cannot. See Researching from primary sources for why a paraphrase of a paraphrase is where fabricated citations come from.

Try this

Take a research task you have already done and re-run just step 1 on the cheapest model, asking for a quote and a location per claim rather than a summary. Compare its notes against what you originally found. You are testing whether the cheap step is faithful, which is the only question that decides whether the saving is real.

What goes wrong

Reading everything into one conversation. Now all thirty documents are re-sent on every turn of steps 2 and 3, and the workflow costs several times the table above.

Downgrading the synthesis step. It is the smallest input in the workflow and the thing everyone will read.

Accepting paraphrase from the reading step. Fabricated and faithful paraphrase look identical.

Skipping step 2. Thirty summaries stapled together is not a synthesis, and conflicts between sources are exactly where the interesting finding is.

How to check it worked

Pick three claims from the finished synthesis and trace each to the source document and the specific passage. If you cannot, step 1 gave you paraphrase and the saving was not free.

Sources

  1. Claude models overview — Anthropic documentation Tier 1 2026-08-31
  2. Claude pricing Tier 1 2026-08-31
  3. Effective context engineering for AI agents — Anthropic Tier 1 2026-08-31