Automation bias: the failure mode of trusting a good tool
Shared methods · A shared method; linked tool guides explain the exact steps.
This page covers tools outside your selection. You can still read it. Find matching guides
The better the tool gets, the less you check it. The EU AI Act names this by its research name and obliges people to stay aware of it.
Most failure modes on this site come from the tool being wrong. This one comes from it being right.
A system that is correct nine times in ten teaches you to stop checking, which is a reasonable response to evidence and exactly what makes the tenth case expensive. The research name for it is automation bias: the tendency to over-rely on the output of an automated system.
It is named in law, which is unusual
The EU AI Act does not treat this as a soft concern. Article 14(4)(b) requires that people assigned to oversee a high-risk system be enabled
"to remain aware of the possible tendency of automatically relying or over-relying on the output produced by a high-risk AI system (automation bias)".
Notice what that implies. The legislator anticipated that adding a human reviewer might not work, because the reviewer would defer. Human oversight is written into the Act as a control, and this subparagraph is the Act admitting the control has a known failure mode.
If you have a "human in the loop" as your answer to AI risk, this is the provision that asks whether the human is actually in it.
Why competence makes it worse
The intuition is that expertise protects you. The evidence points the other way, because expertise is what makes deference rational.
METR's randomised trial found experienced developers were 19% slower using AI tooling on repositories they knew well, while believing they had been about 20% faster. The gap between measured and felt is the thing to sit with: they could not tell from the inside. Nothing about being good at the work supplied a signal that the tool was costing them.
That is automation bias operating on the people best placed to catch it.
What it looks like in practice
It shows up first as approving where you used to review. The diff, the draft, the summary gets a glance and a yes. Something egregious would still stop you; something subtle no longer does.
Then reading shifts from checking to confirming. You scan for whether the output looks like what you expected, in place of whether it is right, and the two feel identical from inside.
Then you start deferring on your own ground — accepting a suggestion inside your own specialism because it arrived confident and you were tired. Retrospectively indefensible, and extremely common.
The last stage is quietest. The verification step you set up months ago gets dropped, because the tool has been fine for a while and the step had stopped finding anything.
Controls that survive contact with a good tool
Resolutions fail here, because the failure is the erosion of a resolution. What works is making the check structural.
Make verification a step the work cannot skip. A test that must pass, a figure someone else confirms, a checklist attached to the task. Anthropic's framing for code generalises: give it a check it can run, so "looks done" stops being the stopping signal.
Sample deliberately. Pick a fixed fraction and check those properly. Broad shallow review is the form automation bias takes when it is trying to look diligent.
Track how often you change something. If your edit rate on AI output has fallen toward zero, either the tool improved or your reviewing did. Only one of those is visible in the number, which is why the number is worth having.
Rotate who reviews. Fresh attention finds what habituated attention has stopped seeing, the same argument as using a fresh context for review.
Try this
Take the last five outputs you accepted. For each, write down what you actually checked — the specific thing, and "I read it" does not count. If the honest answer for most of them is that you checked it looked reasonable, you have measured your own deference, which is the only way to see it.
What goes wrong
Treating it as a discipline problem. It is a predictable response to a reliable tool. Solutions that depend on trying harder fail on the timescale that matters.
Adding a reviewer and stopping there. The Act names the reason: the reviewer defers too. A human in the loop who approves everything is a rubber stamp with a job title.
Confusing familiarity with verification. Having used something for months tells you about its typical behaviour, and nothing about the output in front of you now.
Assuming a better model removes the problem. Higher accuracy raises the cost of the residual errors, because they arrive against a stronger prior that there will not be any.
How to check it worked
Count how many of the last twenty AI outputs you changed before using. A rate near zero over a long stretch means one of two things, and the useful move is finding out which: take five recent accepted outputs and verify them properly. If they hold, your trust is calibrated and you have earned the confidence. If one does not, you have found the interval at which checking needs to be structural.
Sources
Something wrong with this page?
Say what you expected and what you got. That is usually the shortest route to a correction, and it goes on the public issue tracker so the fix is visible.