Security

Data exfiltration through an agent's own tools

Claude Code

An agent that can read files and reach the network is already a channel. The control is egress, because detection is the part that cannot be relied on.

Applies to
Claude Fable 5.1 Claude Opus 5 Claude Sonnet 5 Claude Haiku 4.5
Last verified
Reviewed by
Timothy Fehr

Exfiltration usually means malware. Here it means something duller and harder to spot: two capabilities you granted on purpose, used together.

An agent that can read your files and reach the network can move the first through the second. Neither half is an exploit. The combination is the channel, and it exists from the moment you grant both.

Where the instruction comes from

You do not need a compromised machine. You need one document, page, issue or dependency README that the agent reads, carrying text aimed at the agent rather than at you.

That is prompt injection, and the reason it matters here is what the agent does next. A wrong answer is recoverable. A wrong action, taken with tools that reach outside your machine, is not.

The channels are the useful ones

Every capability you granted for a reason can carry data out:

  • curl and wget, or any command that makes a request
  • Web fetch, pointed at a URL with your data in the query string
  • MCP tools that talk to a remote service
  • git push, to a remote you did not intend
  • A connector that sends mail, which is exfiltration with a delivery receipt

None of these look wrong in a transcript. A request to an unfamiliar host among forty legitimate ones is not something anyone catches by reading.

What Claude Code already does

More than people assume, and you should know it before building your own controls on top.

Commands that fetch from the web — curl, wget and their relatives — are not auto-approved. In Manual mode they prompt like any other non-read-only command. To remove them entirely, put them in permissions.deny.

Web fetch runs in an isolated context window, specifically "to avoid injecting potentially malicious prompts". That breaks the path from fetched content straight into your session.

Most tools that make network requests require approval in Manual mode. And first-time codebases and new MCP servers require trust verification — though note that verification "is disabled when running non-interactively with the -p flag", which is exactly where nobody is watching.

The controls that actually bound it

Sandboxing gives filesystem and network isolation for bash commands. This is the strongest local control and the one people skip because permissions feel sufficient.

Cloud sessions restrict network access by default, and it "can be configured to be disabled or allow only specific domains". An allowlist of domains is the shape you want: the agent works, and there is nowhere unexpected to send anything.

Branch restrictions limit git push to the current working branch in cloud sessions.

Deny rules on network commands, if you never need them in a given repository. Deny beats ask beats allow, and the first match wins.

Anthropic's own advice for the sharp cases is blunt: "Use virtual machines (VMs) to run scripts and make tool calls, especially when interacting with external web services", and "Avoid piping untrusted content directly to Claude".

One documented bypass

Worth singling out because it defeats the permission system rather than squeezing past it. On Windows, the docs recommend against enabling WebDAV or letting Claude Code reach paths like \\*: doing so "may allow Claude Code to trigger network requests to remote hosts, bypassing the permission system."

A permission model has an edge, and this is a documented one.

Try this

List every tool in your current setup that can reach the network. Include MCP servers, and include bash. Then ask what is readable in the same session.

The product of those two lists is your exfiltration surface. Most people find it larger than expected, and usually because of one MCP server added months ago for a task that finished.

What goes wrong

Treating read-only as safe. Read-only means it cannot change your files. It says nothing about where their contents can go, and the combination with any network tool is the whole problem.

Approving network commands once and forgetting. An allowlist entry outlives the task that justified it.

Relying on review of the transcript. One request to an unfamiliar host, inside a long legitimate session, is not something reading catches reliably.

Running unattended without egress limits. -p disables trust verification, and nobody is present to answer a prompt. That combination deserves a sandbox or a VM rather than a permission rule.

How to check it worked

Try to exfiltrate something yourself. Ask the agent to send the contents of a harmless file to a host you control, in the configuration you actually use.

If it refuses or cannot reach the host, your egress controls are real. If it succeeds, you have measured the channel rather than assumed it — and you now know exactly which control was missing, which is a better position than inferring it from a permissions file.

Sources

  1. Security — Claude Code Docs Tier 1 2026-09-04
  2. Configure permissions — Claude Code Docs Tier 1 2026-09-04