What data may go into a model
Shared methods · A shared method; linked tool guides explain the exact steps.
This page covers tools outside your selection. You can still read it. Find matching guides
A classification scheme you can adopt, and the terms difference that most policies get wrong: consumer plans and commercial plans are not the same deal.
Most AI policies say "don't put confidential data in AI tools" and stop. That sentence fails in both directions: it forbids work people will do anyway, and it gives no way to decide the cases that actually come up.
A usable policy answers one question: given this material, and this account, may I paste it?
The terms difference that policies miss
Before any classification scheme, get this straight, because a scheme built on the wrong assumption is worse than none.
Consumer plans carry a training toggle. Decline it and roughly 30-day retention continues. Accept it and data may be used to improve models and retained in de-identified form in training pipelines for up to five years.
Commercial terms covering Work, Enterprise, Education and the API prohibit training on your inputs. That is a different contract, not a setting.
So "is this allowed" depends on which account, not only on which material. An employee using a personal account for work is operating under different terms than the same person on the company plan, possibly with the toggle on, and neither they nor you will notice from the interface.
An example scheme, to rewrite
Four tiers. Take the structure and replace every example with your own — the categories that matter are sector-specific, and a scheme whose examples do not look like your work will not be applied.
Green: no restriction. Already public, or trivially reconstructible. Published documentation, open-source code, marketing copy, general questions. Any account.
Amber: company accounts only. Internal but not sensitive. Internal documentation, non-personal operational data, most application code, drafts of material intended for publication. Requires commercial terms.
Red: company accounts, and only with a stated purpose. Personal data, customer data, unpublished financials, security-relevant configuration, legal material. Needs a recorded reason, a named owner, and, where it is personal data, a lawful basis you can point at.
Black: never. Credentials and keys. Special-category personal data under GDPR unless you have specifically assessed it. Anything under a confidentiality obligation that forbids third-party processing. Material you cannot lawfully send to a processor at all. For patient and clinical data specifically, healthcare-adjacent work walks the line in that setting.
The tiers are worth more than the examples. Publish the examples anyway, because "use your judgement" is how Red becomes Amber on a Friday afternoon.
Say it in terms people can apply
The classification only works if someone can apply it in ten seconds. Two questions do most of the work:
- Would I email this to a supplier under NDA? If no, it is at least Red.
- Would it embarrass us if it appeared in a training set? If yes, it is not Green, whatever the toggle says.
The second is the more useful one, because it survives changes in the terms.
Deletion does not mean gone
Deleting a conversation removes it from the account. It does not reach into a model that has already been trained. No provider can retroactively remove data already incorporated into model weights.
For a policy this matters twice: it means a mistake is not fully correctable, and it means "we'll delete it" is not a remediation you can promise a customer.
The GDPR position, briefly
The EDPB's opinion on AI models and personal data addresses when a model can be considered anonymous, whether legitimate interest can support development and deployment, and what follows if a model was built on unlawfully processed data.
The practical takeaway for a deployer: you are processing personal data when you paste it, your lawful basis has to cover that, and your records of processing should reflect it. The provider being compliant does not make your use of it compliant.
What goes wrong
One rule for everyone. A developer, a recruiter and a finance analyst handle different categories. One line cannot govern all three.
Forbidding without providing. The reliable way to create shadow usage.
Assuming personal and work accounts behave alike. Different contracts.
Writing the scheme and never surfacing it. It has to be reachable from where the decision is made.
Treating deletion as remediation. It is tidying.
How to check it worked
Ask three people in different roles to classify the same five realistic examples. Where they disagree, the scheme is ambiguous at exactly the point it needs not to be — and that disagreement is more useful than any review of the document.
Sources
- Updates to our Consumer Terms and Privacy Policy — Anthropic Tier 1 2026-08-31
- How long do you store my data? — Anthropic Privacy Center Tier 1 2026-08-31
- EDPB Opinion on AI models and personal data (Dec 2024) Tier 1 2026-08-31
Something wrong with this page?
Say what you expected and what you got. That is usually the shortest route to a correction, and it goes on the public issue tracker so the fix is visible.