Incident response when an agent did something
Claude + Cowork + Claude Code
This page covers tools outside your selection. You can still read it. Find matching guides
The playbook assumes a human actor and an agent breaks its first three questions. Decide before the incident what your records will be able to say.
A classic incident playbook opens with who did it, why, and whether they will do it again. An agent incident breaks all three: the actor is software acting under a person's authority, the "why" is a prompt plus whatever content the agent read, and "again" depends on grants that may still be standing.
None of that makes response harder in principle. It makes preparation different, because the evidence and the containment levers are not where the playbook expects them.
First moves, reordered for an agent
Stop the actor, keep the session. End the running task, but the conversation, trail and logs are your primary evidence, so stopping must not mean deleting. Deleting the chat is how retrieved data is removed in some products, which makes tidying up and destroying evidence the same click.
Freeze the grants. The permissions are the blast radius, and they are still live: connectors stay connected, allowlists persist, a schedule will fire again on time. Revoke or suspend before investigating, because the investigation takes hours and the next scheduled run does not wait for it.
Rotate anything readable. If the incident involved content reaching the model, treat every credential in scope as exposed: it was transmitted and sits in a transcript. Rotation is cheap; certainty about what was read is not.
Reconstructing what happened
The record depends on the surface, and this is the part to know before the incident.
Cowork: the step trail is the account of files opened, tools used and decisions made — and then check the filesystem against it, because the summary is the agent's account of itself.
Claude Code, local: the session transcript, plus git status and git diff for what actually changed — remembering that checkpoints do not track bash commands, so the transcript's edits are not the whole story.
Cloud sessions: the strongest position. In Anthropic-hosted environments, "All operations in cloud sessions are logged for compliance and audit purposes". The governance inversion from cloud vs. local bites here: local execution feels cautious, and it is the mode where session history sits on one laptop, outside central retention, invisible to whoever runs the response.
The question the post-mortem must answer
Not "why did the agent do it", where the honest answer is a probability distribution, but "what allowed it to matter?"
An agent misreading intent, or following an instruction hidden in something it read, is the expected failure. The blast-radius design question was always "what could this accomplish if it went wrong"; the incident is that question answering itself. So the durable fixes are grant-shaped, and "better vigilance next time" is not a corrective action, because the premise of the whole security track is that detection is imperfect.
One genuinely new question for the review: did the agent read hostile content? If injection is plausible, the source document, page or message is part of the incident, and whatever ingests from that source is still exposed.
Prepare the answers now
Answer these while nothing is wrong, because during an incident each unanswered one costs an hour:
- Where is the activity record for each surface we use, and who can read it?
- Who can revoke which grants, and how fast, off-hours?
- Which credentials are in scope of which agent, so rotation has a list?
- Does "delete the conversation" destroy evidence on our products?
- Who is told what — including, for business use, whether an incident touching personal data starts a regulatory clock?
Anthropic's own reporting channels are part of the answer: suspicious behaviour goes to /feedback, and vulnerabilities to their HackerOne programme rather than a public post.
What goes wrong
Deleting the evidence while containing. The chat, the trail, the transcript: gone in the same gesture that felt like cleanup.
Investigating with grants live. The schedule fires mid-investigation, and the second occurrence is on the response, not the agent.
Writing "the AI malfunctioned" as root cause. It names no fixable thing. The fixable things are grants, scopes, sources and checks.
Skipping rotation because "it only went to the model". It left the machine and it is in a transcript. That is two copies outside your control of a thing whose value was being secret.
Discovering the logging gap during the incident. Local-mode history on a departed employee's laptop is a fact to design around in advance, never one to learn at the worst moment.
How to check it worked
Run the drill on a harmless fiction: pick one real agent setup, pretend one task went wrong yesterday, and answer the five questions above from your actual records within an hour. Wherever you stall (no trail, no revocation path, no credential list) is the gap, found for the price of an hour instead of an incident.
Sources
- Security — Claude Code Docs Tier 1 2026-09-04
- Use Claude Cowork safely — Anthropic Help Center Tier 1 2026-09-04
Something wrong with this page?
Say what you expected and what you got. That is usually the shortest route to a correction, and it goes on the public issue tracker so the fix is visible.