Limit what an agent can read, change, and send
Shared methods · A shared method; linked tool guides explain the exact steps.
This page covers tools outside your selection. You can still read it. Find matching guides
Choose an execution boundary for the actual tool and host, keep credentials separate, and place human decisions before consequential actions.
Decide what a task may read, change, execute, and send before it runs. Prompt injection or an ordinary mistake can turn a broad permission into an unintended action. A human reviewing the final answer cannot undo data that was already sent elsewhere.
The NCSC discusses why prompt-injection risk needs management beyond filtering malicious-looking text. This page applies that principle to developer workflows.
Describe the actual boundary
For each stage, record:
| Capability | Concrete example |
|---|---|
| Read | A prepared snapshot of the affected files and contract |
| Write | A candidate artifact or one isolated worktree |
| Execute | Reviewed checks in a prepared environment |
| Network | Approved provider and required development destinations |
| External actions | A separately authorized integration action, if present |
| Credentials | Only those required by the executor, excluded from task artifacts |
A folder selection, instruction file, or permission prompt is not automatically an operating-system sandbox. Inspect how the selected product and host enforce the boundary.
Keep instruction and authority separate
Repository files, tool output, web pages, and reviewer findings can contain instructions. Treat that content as evidence to interpret under the existing task policy. It cannot approve a new provider or change the coordinator's allowed commands.
Keep recipes, approval records, and baseline check definitions outside the agent's writable scope. Bind a decision to the exact artifact, source revision, policy, and target being reviewed.
For the downloadable fixture, generated output is validated JSON data. The checker does not execute generated code. That limited task makes its boundary easy to inspect.
Prepare repository execution deliberately
Use branches or worktrees to separate candidate changes. Use the configured sandbox and credential policy to control what running tools and tests can reach.
Tests are executable code. A test process can read files or environment values available to it, even if the agent's normal file tool is restricted differently. Keep provider tokens and production secrets out of the check environment.
Product controls differ. See Claude Code permissions, Codex programmatic use, and Gemini CLI programmatic use. For Cowork-specific architecture, use its own official architecture reference. Do not generalize one product's sandbox description to every cloud or computer-use surface.
Put the human decision before the action
Let already approved read and check stages proceed under their existing scope. When a proposed action needs additional authority, present the exact change and target to the person who can decide.
A timeout, silence, or another model's approval is not a human decision. If an external action's result is uncertain, reconcile the target before attempting it again.
What goes wrong
A worktree can be mistaken for a security boundary. A model's file restrictions may differ from a shell process's access. Shared credentials can escape through tests. Broad network access can permit data transfer before final review. An unverified managed policy can override the local configuration you expected.
How to check
Use a synthetic task and deliberately request one disallowed file, write, or tool action. Observe whether the actual execution environment denies it. Inspect effective policy and logs; do not accept the model's promise as proof.
Then run the intended task within the reduced boundary. Keep the smallest verified capability set that lets it finish.
Sources
- UK NCSC: thinking about AI security Tier 1 2026-09-08
- OpenAI: non-interactive Codex Tier 1 2026-09-08
- Google: Gemini CLI policy engine Tier 1 2026-09-08
Something wrong with this page?
Say what you expected and what you got. That is usually the shortest route to a correction, and it goes on the public issue tracker so the fix is visible.