Anthropic API

Tool use without the common foot-guns

Anthropic API

Ask for the weather with no city and it may invent one, plus a unit you never asked about. Four failures that are cheap to prevent and expensive to debug.

Applies to
Claude Fable 5.1 Claude Opus 5 Claude Sonnet 5 Claude Haiku 4.5
Last verified
Reviewed by
Timothy Fehr

Tool use works on the first try often enough to feel solved. The failures arrive later, in production, and they are mostly four specific things that the documentation warns about plainly.

It will fill in parameters you did not give it

The one to internalise. Anthropic's own example: a get_weather tool requiring a location, asked "What's the weather?" with no city. The model may return

{
  "type": "tool_use",
  "name": "get_weather",
  "input": { "location": "New York, NY", "unit": "fahrenheit" }
}

Two invented values. A location nobody mentioned, and a unit that was not even part of the question. Your handler receives a well-formed call and runs it.

The documentation is specific about the model dependence: Opus "is much more likely to recognize that a parameter is missing and ask for it", while Sonnet "might ask… But it might also infer a reasonable value." And the behaviour is "not guaranteed, especially for more ambiguous prompts and for less capable models."

So a cheaper model in a tool loop is not only slower to reason. It is more willing to guess, and a guessed argument looks identical to a supplied one by the time it reaches your code.

strict: true exists and most schemas do not use it

Adding strict: true to a custom tool definition makes calls conform to your schema exactly. Without it, conformance is likely and not guaranteed.

Turning it on costs nothing and removes a class of runtime parsing failure. Note what it does not do: it constrains the shape of the call, and says nothing about whether the values were invented. Both defences are needed and they cover different failures.

Tool definitions cost tokens on every single request

Three things bill, and only the first is obvious:

  1. The tools parameter itself — names, descriptions and schemas.
  2. tool_use and tool_result blocks in the conversation.
  3. A hidden system prompt the API adds automatically to enable tool use.

That third one is invisible and per-model. On Claude Opus 5 it is 286 tokens with tool_choice of auto or none. The surprise is what forcing a tool costs: any or tool raises it to 406 tokens on the same model.

The spread across models is wider than the numbers suggest. The platform documentation carries the full table, and it does not track generation order, so a migration can change your per-request floor without anything else changing.

Two consequences. Forcing tool use has a standing price, so use it where you need the guarantee and not as a default. And a large tool catalogue is a cost on every call — which is what prompt caching is for, since tools is the first thing in the cache hierarchy and the most stable.

Steering is a prompt problem before it is a parameter problem

With the default tool_choice of auto, the model decides per turn. The documentation says that boundary is steerable in the system prompt, and gives three strengths:

  • "Use the tools to investigate before responding." — nudges toward calling.
  • "Always call a tool first before responding." — pushes harder.
  • "Use your judgment about whether to call a tool or respond directly." — keeps it conservative.

Try the prompt before reaching for tool_choice, because forcing carries the token cost above and removes the model's ability to answer directly when it already knows.

disable_parallel_tool_use: true on tool_choice caps a turn at one call, which is worth setting while you are still debugging a loop.

Client tools and server tools bill differently

Client tools run in your application and cost only tokens. Server tools run on Anthropic's infrastructure and several add usage-based charges on top — web search bills per search, code execution has its own rate.

The distinction matters for cost modelling, because a server tool in a loop has two meters running.

Try this

Call your tool with a deliberately underspecified prompt: leave out a required argument entirely. Look at what arrives in the tool_use block. If it is populated with something plausible, you have just watched the failure this page opens with, and your handler is the only thing standing between it and a real side effect.

What goes wrong

Trusting arguments because the call was well-formed. Schema validity says the shape is right. It says nothing about where the values came from.

Forcing tool use as a default. It costs more tokens per request and removes the option of a direct answer.

Returning errors as normal results. A failure reported as ordinary output teaches the model the call succeeded. Signal errors as errors so it can retry or change approach.

Growing the tool catalogue without measuring. Every definition is on every request. Cache the prefix, or prune.

How to check it worked

Log the arguments your handler receives alongside the user message that produced them, for a day. Then read for values that appear in the call and nowhere in the input. On a well-specified tool the count is zero; anything else tells you exactly which parameter needs a required-and-validated treatment rather than a hopeful one.

Sources

  1. Tool use with Claude — Claude Platform Docs Tier 1 2026-09-04
  2. Prompt caching — Claude Platform Docs Tier 1 2026-09-04