Does it actually make you faster?
Shared methods · A shared method; linked tool guides explain the exact steps.
This page covers tools outside your selection. You can still read it. Find matching guides
The best measurement we have found experienced developers were slower with AI, while believing they were faster. Both halves matter.
A help site that only told you how to use the thing would be leaving out the most useful finding in the field.
The study
In July 2025 METR published a randomised controlled trial. Sixteen experienced open-source developers, 246 real tasks, in repositories where they averaged five years of prior experience. Half the tasks allowed AI tools, half did not. Tooling was whatever they chose, mostly Cursor Pro with the frontier Claude models of the time.
They were 19% slower on the tasks where AI was allowed.
Beforehand, they forecast a 24% speedup. Afterwards — having done the work — they still estimated they had been about 20% faster.
Read the method before you use the number
This finding gets quoted badly in both directions, so:
- Sixteen developers. Small. Real, randomised, and carefully run, but small.
- Their own repositories, where they were already expert. This is close to the worst case for AI assistance: high existing context, high standards, deep familiarity. It says little about unfamiliar code, unfamiliar languages, or people earlier in their careers.
- Early-2025 tooling. Models and harnesses have moved considerably. METR attempted a replication with later tools; what happened to it is its own finding, below.
So: not "AI makes developers slower". Rather, "in the one careful measurement of expert developers on familiar code with early-2025 tools, the measured effect was negative while the perceived effect was strongly positive."
The replication collapsed, which is itself a finding
METR ran the study again in late 2025 with newer tools and more developers, and reported in February 2026 that "the data from our new experiment gives us an unreliable signal of the current productivity effect of AI tools". Not because the effect vanished or reversed — because the measurement broke. There was "a significant increase in developers choosing not to participate in the study because they do not wish to work without AI, which likely biases downwards our estimate of AI-assisted speedup", and "30% to 50% of developers told us that they were choosing not to submit some tasks because they did not want to do them without AI".
So the honest 2026 status of the 19% is: unreplicated and unrefuted. And the reason it could not be replicated says something the original number never could: AI use has become load-bearing enough that experienced developers refuse randomization away from it, at $50 an hour, on their own projects. Whatever the true speedup is, the dependence is measured.
The half that actually generalises
The speed result is contingent on tooling and will age. The perception gap probably will not.
A 39-point spread between measured and felt performance is not a fact about 2025 models. It is a fact about how this work feels. Delegating produces a sense of progress that is largely independent of whether progress occurred: output appears quickly, it looks competent, and the effort you did not spend is vivid while the review time you did spend is not.
What to do about it
Measure something. Anything crude beats intuition. Time a few comparable tasks. Count how often you rewrote the output. Note which tasks you finished and which you abandoned.
Notice where it clearly wins. Unfamiliar territory, boilerplate, first drafts, reading code you did not write, mechanical transformation. The evidence for these is much less ambiguous than for expert work on familiar code.
Notice where it clearly costs. Work where you are already fast, where standards are high, where the review takes longer than doing it. Anthropic's own research on Claude Code usage describes the working split as people making most planning decisions and Claude making most execution decisions — which is exactly the shape that goes wrong when you delegate the planning too.
Treat the good feeling as information about the feeling, not about the work.
Try this
Next time you finish a task with AI help and feel it went quickly, write down the elapsed time before you check the clock. Then check it. Repeat five times. You are not measuring the tool; you are calibrating yourself against it, which is the part you can actually improve.
What goes wrong
Quoting "19% slower" as a fact about AI. It is a fact about sixteen expert developers, familiar code, and early-2025 tools.
Dismissing it because the sample is small or the tools have moved. The perception gap is the finding that generalises, and dismissing the study is usually a way of avoiding it.
Evaluating a rollout on how it feels. That is precisely the measurement that failed.
Concluding you should not use it. Nobody's data says that. The conclusion is to know which of your tasks it helps with, which requires measuring rather than sensing.
How to check it worked
Pick one recurring task and time it, with and without, five times each. If the tool helps, you now have a number instead of a feeling. If it does not, you have found something worth knowing — and you found it the only way it can be found.
Sources
- Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity — METR (Jul 2025) Tier 2 2026-08-31
- Update on measuring AI speedup — METR (Feb 2026) Tier 2 2026-09-06
- METR research index Tier 2 2026-08-31
- How Claude Code is used in practice — Anthropic Tier 1 2026-08-31
Something wrong with this page?
Say what you expected and what you got. That is usually the shortest route to a correction, and it goes on the public issue tracker so the fix is visible.