Guides

191 pages. Filter by track, audience and model.

Usage patterns

  • Add a critic, not a committee Shared methods

    One independent critic measurably improves agent work. A panel of role-play agents mostly agrees with itself, at several times the cost.

  • Ask for three options, not one answer Shared methods

    The first answer is one draw from a distribution, delivered with full confidence. Asking for alternatives shows you the spread before you commit.

  • Ask what it assumed Shared methods

    Every answer to an underspecified question rests on assumptions the model filled in silently. One follow-up makes them visible while they are cheap.

  • Attach the file you are describing Shared methods

    Summarising a document from memory feeds the model your summary's errors. If the file exists, the file goes in — your description was the lossy copy.

  • Chat is not always the right tool Shared methods

    Some jobs fit a chat window; some need search, a file upload, an agent, or a spreadsheet. Matching job to mode is a two-question decision.

  • Check it before you act on it Shared methods

    Confident and correct look identical in a model's answer. Before a fact changes what you do, make it show its evidence — or look it up once.

  • Don't do arithmetic in prose Shared methods

    A language model predicts text, and sums are text it can get plausibly wrong. Make it compute — with a table, code, or an analysis tool — or check the number yourself.

  • Don't fan out past your review capacity Shared methods

    Agents multiply what gets produced; nothing multiplies what you can actually read. The ceiling on useful parallelism is your review throughput.

  • Don't forward what you haven't read Shared methods

    Polished, substance-free AI output has a name now — workslop — and a documented cost: the work you skipped lands on whoever receives it.

  • For current facts, make it search Shared methods

    Prices, versions, deadlines and news drift past a model's training data. Asking without search gets you a confident answer from last year.

  • It does not remember your last chat Shared methods

    "As we discussed yesterday" refers to a conversation the model cannot see. It will play along anyway — that is the dangerous part.

  • It only knows what you paste Shared methods

    The model has not seen your document, your thread, or your last attempt unless they are in the conversation. Most "bad answers" are missing inputs.

  • Let it interview you first Shared methods

    You know things about your problem that you don't know are relevant. One instruction turns the model from guesser into interviewer.

  • Make it argue the other side Shared methods

    A model agrees with your framing by default. Asking for the strongest case against is the cheapest second opinion you will ever get.

  • Make the next rewrite more useful Shared methods

    Three rewrites without saying what was wrong is a slot machine, not editing. Name the gap, or start fresh — repeating "again" does neither.

  • More context is not better context Shared methods

    Pasting everything you have buries the part that matters. Curate what goes in, and say which part the answer should lean on.

  • One request, five tasks Shared methods

    A prompt doing five jobs gets five mediocre answers stapled together. Split it, and each step gets the model's full attention plus your checkpoint.

  • Paste the exact error, not your summary of it Shared methods

    "It doesn't work" starts a guessing game. The exact message, or a screenshot, starts a diagnosis — the difference is one paste.

  • Raise the checking with the stakes Shared methods

    The same answer deserves a skim when it names a restaurant and a real verification when it names a medication dose. Scale the check to the cost of being wrong.

  • Say what you actually want Shared methods

    "Make me a plan" produces the average plan for the average person. Name the outcome, and the answer starts being about your situation.

  • Show one example instead of three adjectives Shared methods

    "Professional but warm" means something different to everyone, including the model. One example of what you want outweighs a paragraph describing it.

  • Stop hitting the limit mid-task Shared methods

    Running dry in the middle of real work is a planning failure with a planning fix: know the window, spend it deliberately, park state before the cliff.

  • Stop re-explaining your project Shared methods

    Typing the same background into every chat is a tax. Put the stable facts somewhere persistent once, and every conversation starts warm.

  • Tell it what "better" means Shared methods

    "Make it better" makes it different. Naming the dimension — shorter, warmer, simpler — is what makes it better in the way you meant.

  • That data belongs to someone else Shared methods

    Your colleague's appraisal, a customer's complaint, your friend's medical story — pasting it is a decision about their data that only you were present for.

  • That paste contains a secret Shared methods

    Passwords, API keys and other credentials have no safe form inside a prompt. Redact before pasting — the model never needed the real value.

  • The agent passed the check by changing the check Shared methods

    Give an agent a test to satisfy and sometimes it satisfies the test instead of the task — editing it, deleting it, or stubbing the code the test needed.

  • The big model is not for everything Shared methods

    Running the flagship on trivial questions spends the quota you will want for the hard ones. Match the model to the task, not to habit.

  • The context you never trim Shared methods

    A session whose context only ever grows pays for its whole history on every turn. Compact at natural breakpoints, or start fresh with a summary.

  • The plan file stopped being true Shared methods

    A written plan is only an asset while it matches reality. Code moves, the spec sits still, and an agent will happily follow the stale version.

  • This conversation is past its best Shared methods

    Long chats accumulate dead weight: rejected drafts, stale constraints, old errors. When steering stops working, a fresh start beats another correction.

  • You are about to use something you have not read Shared methods

    Sending, running or filing an output you only skimmed makes its errors yours. The read costs minutes; owning an unread error costs more.

  • You are doing work it could structure Shared methods

    Copying fifteen rows by hand, reformatting the third similar document, renaming files one by one — repetition with a rule is the model's home turf.

  • You are the copy-paste middleware Shared methods

    Shuttling text between the chat and your apps, piece by piece, is a job the tools can do. If you moved the same kind of thing three times, stop.

Foundations

How these systems actually work, without the hand-waving.

Everyday Claude

Chat, Projects, memory, files, plans and limits.

Developing with ChatGPT

Reason from code and sources, maintain context, and return checked artifacts.

Developing with Gemini

Engineering briefs, Gems, and research with selected sources.

Claude Code

The terminal agent, for people who ship software.

Working with Codex

From a clear task to a checked change, with reusable instructions and skills.

Working with Gemini CLI

Repository context, planning, checked changes, and programmatic use.

Cowork

An agent working across your own files and connectors.

Anthropic API

API, agents, tools and MCP.

  • Agent SDK, Managed Agents, or the raw API Anthropic API

    Four ways to build on Claude, separated by one question: who runs the agent loop and the sandbox. Answer that and the choice mostly makes itself.

  • Batching and cost control Anthropic API

    Half price for work that can wait an hour. The catch is a hard 24-hour expiry, and it is the only cost lever that needs no prompt changes.

  • Building an MCP server: exposing your thing to the model Anthropic API

    A server is a small adapter between the model and something you own. The hard parts are not the protocol; they are the trust boundary you just opened.

  • Choosing a model Anthropic API

    Start at the default, move only for a reason, and match the tier to the step rather than to the project.

  • Context editing, compaction, or memory Anthropic API

    Three mechanisms for a conversation that outgrows its window. One clears, one summarises, one stores — and they compose rather than compete.

  • Do you actually need an agent? Anthropic API

    Anthropic's own advice is to find the simplest thing that works, which "might mean not building agentic systems at all". Usually one call is enough.

  • Error handling that distinguishes retryable from not Anthropic API

    A 529 deserves a retry, a 400 never earns one. Sorting the API's errors into those two bins is most of production error handling.

  • Evaluating an agent: what "working" means Anthropic API

    Anthropic's own advice is volume over polish: more cases graded roughly beats few graded well. For an agent, grade the trajectory and not only the answer.

  • Migrating to a newer model Anthropic API

    A model swap is not a string change. Parameters break loudly, behaviour changes quietly, and the quiet half is where migrations actually fail.

  • Prompt caching that actually hits Anthropic API

    Reads cost a tenth of normal input, so caching pays from the second request. A prompt under the minimum is silently not cached, with no error.

  • Streaming and long turns: the 200 is not the finish line Anthropic API

    Streaming trades one big wait for many small events, and moves failure to after the response starts. The event flow and the mid-stream error are the two things to get right.

  • Structured output Anthropic API

    Guaranteed-valid JSON is not guaranteed-valid data. Half your schema is stripped before the model ever sees it, and your code enforces the rest.

  • Tool use without the common foot-guns Anthropic API

    Ask for the weather with no city and it may invent one, plus a unit you never asked about. Four failures that are cheap to prevent and expensive to debug.

  • Writing a tool description a model can use Anthropic API

    Anthropic calls the description "by far the most important factor in tool performance". Most are one line long. Aim for three to four sentences.

OpenAI API

Bounded model calls, structured artifacts, and explicit access.

Gemini API

Direct model calls and validated handoffs for your own programs.

Trust and honest value

Whether to believe the output, and whether it is really helping.

Security

Prompt injection, agent permissions, blast radius.

Business and governance

Rollout, data rules, the AI Act, works councils.

By context

Assembled playbooks for students, teams, regulated work.

Developer workflows

Complete tasks, hand off checked artifacts, and automate stages across providers.

  • Make artifacts and human decisions explicit Shared methods

    Bind approval to the exact evidence and change being reviewed, and stop stale decisions from authorizing later work.

  • Plan, implement, and review across providers Requires: Claude Code + Codex + Gemini CLI

    Coordinate a bounded change with a reviewed plan, isolated candidate artifacts, independent reviews, and an explicit acceptance decision.

  • Recover a pipeline without losing its boundaries Shared methods

    Handle failed output, denied decisions, cancellation, retries, and limited budgets while preserving the run's evidence.

  • Review a change with another provider Requires: Claude Code + Gemini CLI

    Give independent reviewers the same fixed evidence, validate their findings, and let a human decide the result.

  • Run API and coding-agent stages in one reviewed pipeline Requires: Gemini API + Claude Code + Codex + OpenAI API

    Configure direct model calls alongside CLI jobs, inspect their request limits, and stop uncertain work before another dispatch.

  • Run CLI stages between human decisions Requires: Claude Code + Codex + Gemini CLI

    Configure Claude Code, Codex, and Gemini CLI, review their execution profile, and let the local coordinator launch and supervise eligible stages.

  • Turn a working handoff into a developer pipeline Shared methods

    Automate a known stage, keep its evidence, and retain human decisions at the points where they matter.

  • Turn research into a checked code change Requires: Gemini API + Claude Code + Codex + OpenAI API

    Carry selected source evidence through planning, implementation, and review while keeping uncertainty and human decisions visible.

  • Workflow: a refactor across several files Claude + Claude Code + Anthropic API

    The plan is 1% of the tokens and decides the other 99%. Everything else is mechanical, and mechanical work does not need the expensive model.

  • Workflow: debugging a failure Claude + Claude Code + Anthropic API

    Narrowing costs four times what diagnosing costs. Put the cheap model on the search and the expensive one on the thirty seconds that decide the answer.

  • Workflow: researching and synthesising Claude + Claude Code + Anthropic API

    Reading thirty sources costs eight times what thinking about them costs. This is the workflow where model choice saves the most.

  • Workflow: reviewing a change Claude + Claude Code + Anthropic API

    Three steps, three different models. Reading is not judging, and paying Opus rates to read a diff is the commonest waste in this workflow.

Reference

Model matrix, glossary, quota, where to go instead of here.