Reviewing agent output you did not watch being made
Cowork
This page covers tools outside your selection. You can still read it. Find matching guides
The output looks the same whether it is right or wrong. Review the joins, the edges, and the things it did not say.
You came back to a finished document. It is well-organised, correctly formatted, and confident throughout. None of those properties tell you anything about whether it is right, because they would look identical if it were not.
Anthropic is direct about where responsibility sits: you remain responsible for all actions Claude takes on your behalf. Review is not optional diligence, it is the job.
Why confidence carries no signal
Models are trained and benchmarked in a way that rewards confident answers over admitting uncertainty — a leaderboard gives no credit for "I don't know", so guessing is the better strategy. That is a property of how these systems are evaluated, and no bug in your particular run.
The practical consequence: uniform confidence across a document means nothing, and the sections where it was least sure look exactly like the sections where it was most sure. Unless you asked it to mark them. Which is why the ambiguity rule in Writing a goal an agent can actually finish is a review technique disguised as a prompting technique.
Review the joins, not the middle
Long output is expensive to check end to end, and that is anyway the wrong place to look. Errors cluster at the seams:
Where it moved between sources. A figure that came from one document and ended up in a paragraph about another. Sample across sources; going down the page tells you less.
Where the input was ragged. The malformed file, the scanned page, the document in a different format from the rest. Find the three ugliest inputs and check only their rows.
Where it summarised. Any sentence beginning "overall" or "in general" is a claim it constructed. Those are the ones to test.
Where the format is uniform but the sources were not. A tidy column of dates from documents that formatted dates six different ways means something was normalised, and normalising is where a 03/04 becomes the wrong month.
Read the step trail for what is missing
Cowork shows each step: files opened, tools used, decisions made. On review, the useful reading goes beyond "did each step look sensible". Ask:
- Did it open everything it should have? Twelve files in the folder, nine in the trail, and the summary claims to cover all of them.
- Did it do anything you did not ask for? Helpful extra tidying is still an unreviewed change.
- Did it fall back? A connector failing and the browser being used instead is a different provenance for that data.
Check the filesystem, not the summary
Claude's account of what it did is generated the same way its other prose is. Sort the folder by modification time. Files that changed and should not have are the finding, and no summary will volunteer them.
Try this
Before reading the output, write down the three things that would have to be true for it to be right. Then check only those. It is faster than reading everything and it catches more, because you decided what mattered before the document's confidence had a chance to tell you.
What goes wrong
Reviewing for plausibility. Everything it produces is plausible; that is what it optimises for. Plausibility filters nothing.
Spot-checking the beginning. The start is where it had the most context and the least accumulated drift. Sample the middle and the end.
Accepting the summary as the record. It is another piece of output. A log it is not.
Reviewing once and then trusting the pattern. A run that went well tells you about that run. The next one has different inputs.
How to check it worked
Pick one claim in the output and trace it back to the exact source file and line yourself. If you cannot — because the provenance is not recoverable from the trail — that is a finding about the task setup, not about this particular claim, and the fix is to ask for sources alongside the answer next time.
Sources
- Use Claude Cowork safely — Anthropic Help Center Tier 1 2026-08-30
- Get started with Claude Cowork — Anthropic Help Center Tier 1 2026-08-30
- Why Language Models Hallucinate — Kalai et al., arXiv 2509.04664 Tier 2 2026-09-04
Something wrong with this page?
Say what you expected and what you got. That is usually the shortest route to a correction, and it goes on the public issue tracker so the fix is visible.