When to not use it at all
Shared methods · A shared method; linked tool guides explain the exact steps.
This page covers tools outside your selection. You can still read it. Find matching guides
Five situations where the honest answer is to close the tab. A help site that never says this is selling something.
Most guidance answers how to use these tools. The question of whether gets skipped, which is convenient for everyone with something to sell and unhelpful if you are trying to work well.
Five situations where closing the tab is the right call. None of them is a statement about the technology being bad.
When you cannot verify and being wrong is expensive
Anthropic's rule for code generalises: "If you can't verify it, don't ship it."
The techniques in verifying what you cannot judge narrow uncertainty without eliminating it. Where the cost of being wrong is high — someone's health, someone's money, a legal position, an irreversible change — narrowed uncertainty is the wrong instrument. Buy an expert.
The tell is simple. If you would be unable to defend the answer when challenged, you are not in a position to use it.
When you already know how, in material you know well
METR ran a randomised trial with 16 experienced open-source developers across 246 tasks in repositories they already knew well. Allowed AI tools, they were 19% slower. They had forecast a 24% speedup, and afterwards still believed they had gained about 20%.
Read the conditions rather than the headline, because the conditions are where the lesson is: deep familiarity with the codebase, genuine expertise, early-2025 tooling. That is precisely the situation where you hold context the model has to be told, and telling it costs more than doing the work.
The inverse also follows. Unfamiliar material, boilerplate, or a language you do not know is where the trade runs the other way.
When the task is the learning
If the point is to end up able to do something, delegating the doing removes the point. A student writing an essay is building the argument-forming capability; the essay is a by-product. Same for reading a paper in your field, or the first pass through a codebase you will maintain.
Worth being honest here: this is a real cost with a delayed invoice. You will not notice the skill you failed to build for months, which is exactly what makes it easy to keep choosing against.
When the material may not leave
Some data cannot go into a prompt, and no amount of usefulness changes that. Client material under NDA, personal data you have no basis to process, secrets, regulated records. See what never goes in a prompt.
This one is not a judgement call, which makes it the easiest of the five to apply and the most costly to get wrong.
When you are in crisis, or need a diagnosis
The APA's advisory is direct: "Relying solely on an app during a mental health emergency can be dangerous." Its broader position is that these tools may be a supportive adjunct and never a replacement for a qualified provider.
The same reasoning covers medical, legal and financial decisions of consequence. Those fields have licensed professionals for a reason, and part of that reason is accountability — a model cannot carry any. More on wellbeing use.
The question that covers most cases
Before starting, ask: what would make this the wrong tool here?
If nothing comes to mind, either the task suits it — checkable output, material you can judge, nothing confidential, no skill you are trying to build — or you have not looked. The reflex of asking is most of the value.
Try this
Take the last three tasks you used it for. For each, ask whether you were faster, and how you know. If the answer for one of them is "it felt faster", put that one under the METR result and look again. Felt and measured came apart by 39 points in that study.
What goes wrong
Using it because it is open. The tab being there is not a reason. It is the single most common reason.
Mistaking speed for progress. Producing a draft quickly moves the artefact along. Whether it moves the work along depends on whether the draft is right.
Delegating the judgement along with the labour. Handing over the drafting is ordinary. Handing over which option to pick, without the material to check it, is the one to catch.
Assuming a bad fit becomes a good one with better prompting. Where the issue is verification, confidentiality or skill-building, prompt quality does not touch it.
How to check it worked
A week after deciding not to use it for something, ask whether the decision held up. If you did the work yourself and it went fine, you calibrated correctly. If you struggled with something the tool would have handled and which you could have checked, you were over-cautious — and that is worth knowing too, because the point is accuracy about where it helps rather than abstinence.
Sources
- Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity — METR Tier 1 2026-09-02
- Best practices for Claude Code — Anthropic Tier 1 2026-09-02
- APA Health Advisory on the Use of Generative AI Chatbots and Wellness Applications for Mental Health (November 2025) Tier 1 2026-09-02
Something wrong with this page?
Say what you expected and what you got. That is usually the shortest route to a correction, and it goes on the public issue tracker so the fix is visible.