Foundations

Why the same prompt gives different answers

Shared methods · A shared method; linked tool guides explain the exact steps.

Sampling makes the output vary by design, and on current models there is no setting that turns it off. The skill is designing work that survives it.

Applies to
Shared methods
Last verified
Reviewed by
Timothy Fehr

Ask the same question twice and you get two different answers. This surprises people who expect software to be a function: same input, same output. A model is not doing that. At every step it produces a probability distribution over possible next tokens and then draws from it, and a draw is a draw.

So variation is the design, working as intended. Which sounds like a problem until you know what it does and does not affect.

The dial is gone on current models

There used to be a temperature parameter that widened or narrowed the distribution. Low values made the output more predictable, high values more varied. The API reference now marks it deprecated: on models released after Claude Opus 4.6, only the default value is accepted and anything else comes back as a 400 error. top_p and top_k went the same way.

The model matrix records which models still accept them. Do not guess from a version number — the boundary sits mid-generation.

For anyone using Claude through the apps rather than the API, this changes nothing that was ever exposed to you. There was no dial in the interface to begin with.

What varies, and what mostly does not

Wording varies almost always. Structure varies a lot: you may get a table one time and three paragraphs the next, from an identical request.

Substance varies least, and it varies most on exactly the questions where you can least afford it. A well-evidenced factual answer tends to come back the same in content. A judgement call at the edge of what the model can support — which of two designs is better, whether a clause is enforceable, what a borderline test result implies — can genuinely land differently run to run.

That asymmetry is the useful part. Variation across runs is a rough signal of how well-supported an answer is. If three runs agree on substance, the answer sits somewhere solid. If they diverge, you have found a question the model is guessing at, and you learned that for the price of two extra prompts.

Try this

Take a question you have already asked and had answered. Open two fresh conversations and ask it again, identically, without the original thread's history. Compare the three on substance only, ignoring wording.

Agreement means the ground is firm. Divergence means you were shown one sample from a wide distribution and told nothing about its width.

What goes wrong

Treating one run as the answer. The output carries no marker for how confident the underlying distribution was. A guess and a well-supported fact arrive in the same tone, at the same length, with the same fluency.

Chasing reproducibility that was never on offer. Teams sometimes spend real effort trying to pin outputs down for audit or testing. The vendor documentation says plainly that this is not available. Version the prompt and the inputs, record the output you actually acted on, and drop the idea of replaying it byte for byte later.

Re-rolling until you like the answer. Running it again after a reply you disagreed with, then accepting the second one, is selecting for what you already believed. If you are going to compare runs, decide what would change your mind before you read them.

Blaming variation for a context problem. Answers that drift over a long conversation are usually context filling up, which has a different cause and a different fix. Sampling variance shows up between fresh runs, not as a slow decline within one.

How to check it worked

Pick a question where you acted on a single answer last week. Re-ask it in a clean conversation and compare the substance. If it holds, your confidence was warranted. If it does not, you now know one decision that rested on a draw rather than on the material, and that is worth knowing before someone else finds it.

Sources

  1. Create a Message — Claude API reference Tier 1 2026-09-02
  2. Effective context engineering for AI agents — Anthropic Tier 1 2026-09-02