Examples beat adjectives: showing instead of describing
Shared methods · A shared method; linked tool guides explain the exact steps.
This page covers tools outside your selection. You can still read it. Find matching guides
"Conversational but authoritative" means something different to everyone, including the model. Three to five examples pin down what adjectives cannot.
Ask for something "concise but thorough, professional but warm" and you have handed over four words that each mean something different to everyone, including the model, whose reading of "professional" is an average over everything ever labelled professional.
Show it two paragraphs you consider right, and the ambiguity is gone. Not reduced. Gone, because the example carries the register, the sentence length, the formality and a dozen properties you would never think to name.
What Anthropic actually recommends
The prompting guidance is unambiguous about where examples rank: "Examples are one of the most reliable ways to steer Claude's output format, tone, and structure". A few well-crafted ones — few-shot or multishot prompting — improve accuracy and consistency.
The concrete number: include 3–5 examples for best results. Three properties make an example earn its place:
- Relevant: mirror your actual use case closely.
- Diverse: cover edge cases, and vary enough that the model "doesn't pick up unintended patterns".
- Structured: wrap them in
<example>tags so they are distinguishable from instructions.
That second one is the subtle trap. Three examples that all happen to start with a question teach "start with a question" whether you meant it or not. Everything consistent across your examples is read as intended.
Why this works when adjectives do not
An adjective is a compressed instruction that the model must decompress using its own defaults. "Formal" decompresses to its formal, and the gap between its default and your intent is exactly the part you cared about.
An example is not compressed. The properties travel intact, including the ones you could not articulate — which for tone and format is most of them. It is the same reason a style guide with samples beats one with only rules, and the same reason a spec names files rather than describing them.
Edge cases benefit most. "Handle missing data gracefully" is an adjective in disguise; one example showing a row with a missing field and the exact output you want for it settles the question permanently.
The costs, so you weigh them honestly
Examples are tokens. Three to five substantial ones can be several hundred each against your context budget. Cheap on a one-off question, a permanent per-request cost in an application, where they belong in the cached prefix.
And examples pin format hard. If you want the model to exercise judgement about structure, examples will suppress exactly that. Show five bullet-list answers and you will not get prose again, however much the eleventh question deserves it.
Try this
Take a prompt of yours that stacks three or more adjectives. Replace the adjectives with two examples of output you would accept, wrapped in <example> tags. Run both versions on the same input.
The difference is usually visible on the first response, and it is largest exactly where your adjectives were vaguest.
What goes wrong
Examples that are all alike. The model learns the accidental pattern along with the intended one. Vary everything you do not mean.
Examples of the input, not the output. Showing what you receive teaches less than showing what you want back. Pair them where the mapping matters.
Adjectives layered on top anyway. If the examples and the adjectives disagree, you have given contradictory instructions and will get an average. Let the examples carry it.
Stale examples in a standing prompt. An application's examples encode yesterday's format. When the format changes, they are the part everyone forgets to update — and they win against the updated instruction.
How to check it worked
Run the same request five times and compare structure across the runs. With good examples the format holds steady and the content varies; that is the consistency the technique buys. If format still wanders, your examples disagree with each other somewhere — find the property they fail to share, because that is the one the model is still guessing.
Sources
- Prompting best practices — Claude Platform Docs Tier 1 2026-09-04
Something wrong with this page?
Say what you expected and what you got. That is usually the shortest route to a correction, and it goes on the public issue tracker so the fix is visible.