Verifying output you are not qualified to check
Shared methods · A shared method; linked tool guides explain the exact steps.
This page covers tools outside your selection. You can still read it. Find matching guides
You cannot review an answer in a field you do not know. You can check whether it is internally consistent, externally corroborated, and willing to be wrong.
The advice to check the output assumes you can. Often the whole reason you asked is that you cannot: the contract clause, the statistics method, the language you do not read, the subsystem you have never touched.
Anthropic's own guidance states the rule without softening it — "If you can't verify it, don't ship it." That is correct and, on its own, not much help, because it does not say what to do when verifying is exactly the thing you lack. So here is what you can check without the expertise.
Verification is not the same as review
Reviewing means judging whether the answer is right. That needs the knowledge you do not have.
Verifying means establishing whether the answer behaves like a right one. That needs a method, and methods transfer between fields. Every technique below works without you understanding the subject matter.
Make it produce evidence, not assertion
The single highest-yield move: require the answer to carry something checkable alongside it. Anthropic's phrasing for code is "Have Claude show evidence rather than asserting success: the test output, the command it ran and what it returned" — and the principle generalises well past code.
Outside code that means quotes with locations, named sources for each specific claim, the arithmetic written out, and the standard or clause cited by number. You are not evaluating the reasoning. You are confirming the referenced things exist and say what they were reported to say. That needs literacy, and little else.
Ask it to argue the other side
Prompt for the strongest case against its own answer, and for the conditions under which it would be wrong.
An answer resting on solid ground produces a specific, narrow counter-case. A guess produces vague hedging or an unconvincing strawman, because there is no real opposing structure to describe. You are reading the shape of the rebuttal, not judging its content — and shape is legible to a non-expert.
Check it against itself, cold
Ask the same question in a fresh conversation with no history, then compare the two answers on substance. Divergence means you were shown one sample from a wide distribution, which is how sampling works and a decent proxy for how well-supported the answer is.
This catches a specific failure that re-reading never will: an answer that is internally coherent, fluent, and simply not stable.
Corroborate outside the model
Anthropic's guidance on web search says it plainly: "Cross-reference cited sources to understand the full picture. Use authoritative sources for critical decisions."
A second model is weaker corroboration than it appears. Models share training data and share incentives — the Nature work on evaluation incentivising confident guessing applies to all of them — so agreement between two can reflect a shared bias as easily as a fact. An independent human source, a standards document, or a primary text is worth more than a second opinion from the same kind of system. When the person you must convince is a colleague holding a fluent answer, the vocabulary for that is its own skill.
Know when none of this is enough
Some decisions need an actual expert, and the honest move is to buy one. Where being wrong is expensive, irreversible, or affects someone else's health, money or legal position, these techniques narrow your uncertainty without removing it.
Their real use is different and still valuable: arriving at the expert with a sharper question. An hour with a specialist goes considerably further when you can name what you do not understand.
Try this
Take an answer you accepted recently in an area outside your competence. Ask for the three specific claims it rests on, with a source for each. Then check one — the one you would be most embarrassed to repeat wrongly.
What goes wrong
Treating fluency as a signal. Well-formed prose is what these systems produce regardless of whether the content is sound. It carries no information about correctness.
Asking "are you sure?" It will revise, because that reads as dissatisfaction. Revision under pressure tells you about the pressure.
Accepting the second answer because it is second. Sequence is not evidence. If two answers differ, one of them is wrong and you still need to find out which.
Verifying the parts you understand. Attention naturally lands on the checkable bits, which are the bits least likely to be wrong. The unverified remainder is where the risk sat the whole time.
How to check it worked
Count how many of the answer's load-bearing claims you can now trace to something outside the model. If the answer is none, you have not verified it — you have read it twice. That is a fine place to stop for something low-stakes, and a clear signal to find a human for anything that is not.
Sources
- Best practices for Claude Code — Anthropic Tier 1 2026-09-02
- Enable and use web search — Anthropic Help Center Tier 1 2026-09-02
- Evaluating large language models for accuracy incentivizes hallucinations — Kalai, Nachum, Vempala & Zhang, Nature (2026) Tier 1 2026-09-05
- Same paper, accepted article preview (openly readable PDF) Tier 1 2026-09-05
Something wrong with this page?
Say what you expected and what you got. That is usually the shortest route to a correction, and it goes on the public issue tracker so the fix is visible.