Anthropic API

Writing a tool description a model can use

Anthropic API

Anthropic calls the description "by far the most important factor in tool performance". Most are one line long. Aim for three to four sentences.

Applies to
Claude Fable 5.1 Claude Opus 5 Claude Sonnet 5 Claude Haiku 4.5
Last verified
Reviewed by
Timothy Fehr

The description is the only thing the model knows about your tool. It has not read your code, it cannot see your backend, and everything it decides — whether to call, when to call, what to pass — comes from the text you wrote in that field.

Anthropic's guidance is unambiguous about the weight this carries: "Provide extremely detailed descriptions. This is by far the most important factor in tool performance". And most tool descriptions in the wild are one line.

What a description has to answer

The documentation lists what a description must answer: what the tool does, when it should be used "(and when it shouldn't)", what each parameter means and how it affects behaviour, and "any important caveats or limitations", including what the tool does not return.

The stated floor: "Aim for at least 3–4 sentences for each tool description, more if the tool is complex".

The "when it shouldn't" and "does not return" halves are the ones people skip, and they do the most work. A model that knows a tool's limits stops guessing at the gaps — the documented failure where a missing parameter gets invented is fed by descriptions that never said what happens without it.

The worked contrast, worth copying

Anthropic's own poor example: "Gets the stock price for a ticker."

Their good one, for the same tool: "Retrieves the current stock price for a given ticker symbol. The ticker symbol must be a valid symbol for a publicly traded company on a major US stock exchange like NYSE or NASDAQ. The tool will return the latest trade price in USD. It should be used when the user asks about the current or most recent price of a specific stock. It will not provide any other information about the stock or company."

Read what each sentence buys: validity constraints on the input, the exact return format, the trigger condition, and the scope boundary. Five sentences, and every ambiguity a model could act on is closed.

This is examples beat adjectives applied to tools — and for complex inputs the parallel is literal: input_examples accepts schema-validated example calls, at a documented cost of roughly 20 to 50 tokens for simple examples and 100 to 200 for nested ones.

Fewer, better-named tools

The same page gives structural rules aimed at the choosing problem rather than the calling problem:

"Consolidate related operations into fewer tools": one tool with an action parameter over create_pr / review_pr / merge_pr, because "Fewer, more capable tools reduce selection ambiguity".

And namespace across services (github_list_prs, slack_send_message), which matters most under tool search, where the model finds tools by name before it reads anything else.

The response is part of the interface

The description gets the call made; the response decides what the model can do next. The guidance: return "only high-signal information", prefer "semantic, stable identifiers" over opaque internal references, and include only the fields needed for the next step — "Bloated responses waste context", and that context is the budget everything else competes for.

A tool that answers with forty fields of internal state taxes every turn that follows it.

Try this

Take your shortest tool description and rewrite it to answer all four questions: does, when (and when not), parameters, limits. Run the same prompts against both versions and compare when the tool gets called and what gets passed.

Wrong-tool selections and invented parameters usually drop visibly, which is a better argument than this page.

What goes wrong

One-line descriptions. The model fills the missing detail with assumptions, and the assumptions are invisible until a call goes wrong.

Describing only the happy path. Without "when it shouldn't" the tool gets called for adjacent jobs it cannot do, and the failure surfaces downstream.

A tool per verb. Twelve narrow tools force twelve-way selection on every turn. Consolidation moves the choice into a parameter the schema can constrain.

Stating limits nowhere. What the tool does not return is exactly what the model will otherwise promise the user it can get.

Forgetting descriptions bill. Every tool's full definition rides on every request or its cached prefix. Detail pays for itself; padding does not.

How to check it worked

Log a day of tool calls and read the misses: wrong tool chosen, right tool with wrong arguments, tool skipped where it should have fired. Each miss maps to a question the description failed to answer, and that mapping tells you which sentence to add — which beats rewriting the whole thing on instinct.

Sources

  1. Define tools — Claude Platform Docs Tier 1 2026-09-04
  2. Tool use with Claude — Claude Platform Docs Tier 1 2026-09-04