Security

Limit what an agent can read, change, and send

Shared methods · A shared method; linked tool guides explain the exact steps.

Choose an execution boundary for the actual tool and host, keep credentials separate, and place human decisions before consequential actions.

Applies to
Shared methods
Last verified
Reviewed by
Timothy Fehr

Decide what a task may read, change, execute, and send before it runs. Prompt injection or an ordinary mistake can turn a broad permission into an unintended action. A human reviewing the final answer cannot undo data that was already sent elsewhere.

The NCSC discusses why prompt-injection risk needs management beyond filtering malicious-looking text. This page applies that principle to developer workflows.

Describe the actual boundary

For each stage, record:

CapabilityConcrete example
ReadA prepared snapshot of the affected files and contract
WriteA candidate artifact or one isolated worktree
ExecuteReviewed checks in a prepared environment
NetworkApproved provider and required development destinations
External actionsA separately authorized integration action, if present
CredentialsOnly those required by the executor, excluded from task artifacts

A folder selection, instruction file, or permission prompt is not automatically an operating-system sandbox. Inspect how the selected product and host enforce the boundary.

Keep instruction and authority separate

Repository files, tool output, web pages, and reviewer findings can contain instructions. Treat that content as evidence to interpret under the existing task policy. It cannot approve a new provider or change the coordinator's allowed commands.

Keep recipes, approval records, and baseline check definitions outside the agent's writable scope. Bind a decision to the exact artifact, source revision, policy, and target being reviewed.

For the downloadable fixture, generated output is validated JSON data. The checker does not execute generated code. That limited task makes its boundary easy to inspect.

Prepare repository execution deliberately

Use branches or worktrees to separate candidate changes. Use the configured sandbox and credential policy to control what running tools and tests can reach.

Tests are executable code. A test process can read files or environment values available to it, even if the agent's normal file tool is restricted differently. Keep provider tokens and production secrets out of the check environment.

Product controls differ. See Claude Code permissions, Codex programmatic use, and Gemini CLI programmatic use. For Cowork-specific architecture, use its own official architecture reference. Do not generalize one product's sandbox description to every cloud or computer-use surface.

Put the human decision before the action

Let already approved read and check stages proceed under their existing scope. When a proposed action needs additional authority, present the exact change and target to the person who can decide.

A timeout, silence, or another model's approval is not a human decision. If an external action's result is uncertain, reconcile the target before attempting it again.

What goes wrong

A worktree can be mistaken for a security boundary. A model's file restrictions may differ from a shell process's access. Shared credentials can escape through tests. Broad network access can permit data transfer before final review. An unverified managed policy can override the local configuration you expected.

How to check

Use a synthetic task and deliberately request one disallowed file, write, or tool action. Observe whether the actual execution environment denies it. Inspect effective policy and logs; do not accept the model's promise as proof.

Then run the intended task within the reduced boundary. Keep the smallest verified capability set that lets it finish.

Sources

  1. UK NCSC: thinking about AI security Tier 1 2026-09-08
  2. OpenAI: non-interactive Codex Tier 1 2026-09-08
  3. Google: Gemini CLI policy engine Tier 1 2026-09-08