Trust and honest value

Explaining to someone else why the output is wrong

Shared methods · A shared method; linked tool guides explain the exact steps.

Rejecting AI output is easy alone and hard in a meeting. The words that work attack the claim, not the tool, and put the burden where it belongs.

Applies to
Shared methods
Last verified
Reviewed by
Timothy Fehr

Sooner or later the review problem becomes a persuasion problem. A colleague pastes an answer that settles the debate; a manager forwards a generated analysis and asks you to "just sanity-check" it; a client arrives with a confident paragraph that contradicts your advice. You can see it is wrong. Now you have to say so, to someone who watched it be produced in seconds and finds it fluent, detailed and sure of itself.

This page is the vocabulary for that conversation.

Say why wrongness is normal, not scandalous

The strongest opening is not "the AI made a mistake". It is that confident error is how these systems work when uncertain. The clearest recent statement of why comes from research on hallucination: "Like students facing hard exam questions, large language models sometimes guess when uncertain, producing plausible yet incorrect statements instead of admitting uncertainty". And the guessing is not a bug awaiting a patch: "the training and evaluation procedures reward guessing over acknowledging uncertainty", because "language models are optimized to be good test-takers, and guessing when uncertain improves test performance".

That framing does something important socially: it removes the implication that the person who used the tool was foolish. The output looks exactly like a right answer because producing right-looking answers is the training target. Nobody was careless; the artefact is simply not evidence of its own correctness.

Move the debate off the text

Arguing with fluent prose is a losing game; it answers every objection with more fluency. Three moves relocate the argument to ground where evidence wins:

Isolate one checkable claim. Not "this analysis feels off" but "it says the deadline is in Article 30; Article 30 is about something else — here is the text." One verified error, concretely shown, does more than five suspicions, because it breaks the assumption doing all the work: if it were wrong, it wouldn't sound this sure. It always sounds this sure.

Re-ask and show the spread. Run the same question three times, or phrased two ways. Different answers make the point no lecture can: the output is a draw from a distribution, and you do not resolve a disputed fact by sampling it again. "Which of these three is the one we're trusting?" is a fair and devastating question.

Refuse the circular check. "I asked it and it confirmed" and the mid-conversation "you're right, actually…" reversal are the same trap: the system under social pressure updates toward the asker, in both directions. Verification has to leave the conversation — to the source, the code, the measurement — or it is not verification.

Put the burden where it belongs

The economics of the situation are lopsided: the paragraph cost seconds and the rebuttal costs an afternoon. Do not accept those terms. The productive sentence is: "whoever wants to rely on this claim owns checking it — here's what checking looks like." That is not obstruction; it is the same verification rule the site applies everywhere, stated at the moment it bites. A generated claim nobody will verify is not information; it is liability wearing information's clothes.

With a manager, the framing that lands is risk-shaped rather than truth-shaped: "if we act on this and it is wrong, here is what it costs; checking costs twenty minutes." You are not asking them to distrust their tool, only to price the verification the way they price any other input.

What goes wrong

Attacking the tool instead of the claim. "You can't trust AI" starts a culture war and loses it. "This citation does not exist — look" ends one.

Matching fluency with fluency. A long rebuttal essay gets skimmed. One falsified specific gets remembered.

Accepting the apology as progress. The model conceding your point is the same mechanism as the model confirming theirs. Neither settles anything.

Winning rudely. The person who pasted the answer decides whether your correction was help or humiliation, and that decision sets whether they check next time or just hide the tool use.

Being right without receipts. If you cannot show the error concretely, you have a hunch, and their fluent paragraph beats your hunch in every meeting. Get the receipt first.

How to check it worked

The test is what the other person does next time. If they arrive saying "it claims X, I have not verified it yet", the conversation worked: the artefact got reclassified from answer to draft. If they stopped using the tools or started hiding them, it failed in the other direction — the goal was never to win, it was to install the checking habit one level up from your own desk.

Sources

  1. Why Language Models Hallucinate — Kalai et al., arXiv 2509.04664 Tier 2 2026-09-04