Deskilling: keeping the judgement you are delegating
Shared methods · A shared method; linked tool guides explain the exact steps.
This page covers tools outside your selection. You can still read it. Find matching guides
Delegating work is fine. Delegating the judgement that checks the work is how you lose the ability to notice when the work is wrong.
Two different things get delegated when you hand work to a model, and only one of them is safe to lose.
The labour is the drafting, the searching, the first pass. Losing the habit of doing that yourself costs little, the way losing mental arithmetic to a calculator cost little.
The judgement is knowing what good looks like: whether the draft is right, whether the code is sound, whether the answer smells wrong. Every review practice on this site assumes you still have it. And it is a skill, which means it is maintained by use and decays without it.
Deskilling is what happens when delegating the first quietly takes the second with it.
Why this differs from the calculator
The calculator comparison fails on one point. A calculator's errors are rare and random; a model's are frequent enough to matter and fluent enough to miss. The arrangement only works while a person who can tell good from bad is checking — which is why the EU AI Act writes human oversight into law for high-risk systems, and why the same law concedes the overseer tends to stop overseeing.
Deskilling is the slower, deeper version of that concession. Automation bias is the reviewer not bothering. Deskilling is the reviewer no longer being able to, and the second is harder to reverse, because the capacity itself is gone rather than idle.
The mechanism is unremarkable
No one decides to lose a skill. The judgement was built by doing the work, the essay was exhaust from building the ability, and it is maintained the same way. Delegate the doing for a year and the maintenance stops. Nothing marks the day the atrophy starts, and the feeling of competence outlasts the competence, because felt and measured ability come apart: METR's developers believed they were 20% faster while being 19% slower, with every advantage in judging their own work.
Now run that forward. The gap those developers could not feel was a productivity gap. The deskilling version is a judgement gap, and the skill you would use to notice it eroding is the one eroding.
What to keep, deliberately
Not everything. Keeping every skill sharp means delegating nothing, which forfeits the value. The question is which judgement your role actually rests on, and the test is concrete: which outputs would be catastrophic to approve wrongly? The ability to evaluate those is the one to maintain by hand.
Three practices that hold it:
Do a fraction of the work cold, on purpose. Not as review; as production. A developer who writes one real change a week by hand, an analyst who builds one model unassisted, keeps the muscle the reviews depend on. This is the standing version of the retrieval check: close the tab and do the thing.
Review deep on a sample rather than shallow on everything. Full-depth review of a fixed fraction exercises real judgement; skimming everything exercises approval.
Keep score against ground truth. Predict what deep review will find before you do it. Calibration you can measure is the only counter to a confidence that outlasts its basis.
For teams, the stakes compound
An individual deskills in years. A team can deskill in one hiring cycle: the seniors keep judgement they built before delegation, the juniors never build it, and the pipeline that turns juniors into seniors was the work that got delegated.
That is a planning problem, and it belongs in rollout thinking now rather than in a retrospective later: decide which capabilities the team must retain the ability to do unaided, and cost the maintenance (some hand-done work, slower on purpose) as what it is: an insurance premium.
What goes wrong
Confusing comfort with competence. Fluent tool-assisted output feels like skill. The test is what you can still do with the tool closed.
Keeping the wrong skill. Maintaining the labour (typing speed, syntax recall) while losing the judgement (is this design sound) guards the cheap half.
Assuming seniority protects you. Judgement built over a decade still decays; it just starts from higher up.
Noticing at review time. The eroding skill is the noticing skill. External checks (the cold work, the scored predictions) exist because introspection reports competence right up until the exam.
How to check it worked
Twice a year, do one representative piece of work end to end with no assistance, at the difficulty you claim to operate at. Treat it like the calibration it is. If it goes fine, the delegation is resting on a maintained foundation. If it is harder than it used to be, you have caught the invoice early — while the fix is practice rather than retraining.
Sources
Something wrong with this page?
Say what you expected and what you got. That is usually the shortest route to a correction, and it goes on the public issue tracker so the fix is visible.