Claude Code

Subagents: delegating without losing the thread

Claude Code

A subagent reads in its own context and returns only a summary, so the investigation costs you a paragraph instead of thirty files.

Applies to
Claude Fable 5.1 Claude Opus 5 Claude Sonnet 5 Claude Haiku 4.5
Last verified
Reviewed by
Timothy Fehr

The expensive part of most agent work is reading. Answering "how does auth work here" might take thirty file reads, and every one of them stays in your context afterwards, crowding the work you actually wanted to do.

A subagent moves that reading somewhere else. It runs in its own context window and, in the documentation's words, "does that work in its own context and returns only the summary."

The economics, in one number

Anthropic's published walkthrough of a session shows a research subagent reading three files and returning 420 tokens to the main conversation. Its own reads cost your window nothing, because they happened in a different one.

Compare that against reading those files yourself at 1,100–2,400 tokens each. The delegation is not a stylistic preference; it is an order-of-magnitude difference in what the investigation costs.

What it starts with, and what it does not

This is the part that determines whether delegation works, and it surprises people.

A subagent starts fresh. It does not see your conversation history, files already read, skills already invoked, your output style, or the main conversation's auto memory.

It does receive its own system prompt, a task message summarising the work, the CLAUDE.md hierarchy, a git status snapshot from session start, any preloaded skills, and a roster of sibling agents. The built-in Explore and Plan agents skip CLAUDE.md and git status deliberately, to stay fast and cheap.

Defining one

A Markdown file with frontmatter, in .claude/agents/ for a project or ~/.claude/agents/ for yourself. Those are the two you will write by hand, and they sit in the middle of five scopes that resolve in a fixed order: managed settings first, then the --agents CLI flag, then project, then user, then a plugin's own agents/ directory. A name defined higher up wins.

Project subagents are found by walking up from your working directory, so every .claude/agents/ between there and the repository root is scanned, and the definition closest to where you are standing wins. Both directories are scanned recursively, so you can file them under agents/review/ without changing anything: identity comes from the name field alone, which is also why those names have to stay unique across the whole tree.

---
name: code-reviewer
description: Reviews code for quality and best practices
tools: Read, Glob, Grep
model: sonnet
---

You are a code reviewer. When invoked, analyze the code and provide
specific, actionable feedback on quality, security, and best practices.

name and description are required, and the description is what Claude reads when deciding whether to delegate — so write it as a routing instruction, not a job title.

Optional fields that earn their place: tools and disallowedTools scope what it can touch, model picks a cheaper one for mechanical work, maxTurns bounds a runaway, and isolation: worktree gives it its own git checkout so parallel edits do not collide.

The tool restriction that matters for review

A reviewer with tools: Read, Glob, Grep cannot edit anything. That is a stronger guarantee than asking it not to, and it makes the review independent of the work in a way that a prompt alone does not.

It also stops the reviewer delegating. With nesting on by default, Agent is a tool like any other, so a subagent holding it can spawn its own workers and the read-only guarantee stops at the first hop. Listing tools explicitly excludes it already; where you are not listing tools, disallowedTools is the lever.

Background subagents also lose most built-in tools by default, keeping file tools, Bash, web tools, skills and MCP tools. Foreground ones keep the full inherited set.

Limits worth knowing before you scale it

Subagents can spawn subagents up to three layers deep by default, and up to 20 run concurrently. Each invocation costs tokens for its own system prompt and context, so delegation is cheap for your window and not free overall.

Both numbers are settings you can change. CLAUDE_CODE_MAX_SUBAGENT_SPAWN_DEPTH sets the depth, and 1 turns nesting off altogether, which is the lever to reach for if a fan-out ever surprises you. The 20 caps how many run at once. Nothing caps how many a session spawns over its lifetime, so a long session can get through hundreds without ever meeting that limit.

One quiet trap: combined subagent descriptions over 15,000 tokens trigger a warning, because those descriptions load into every session. A large collection of agents is a startup cost you pay whether or not you use them.

When the main thread is the right place

Delegation suits work that is self-contained, verbose, and reducible to a summary. It suits it badly when the task needs iterative back-and-forth, when phases share a lot of context, when the change is small and targeted, or when latency matters.

The test: could you write down what you want back in a sentence? If yes, delegate. If the answer only emerges through conversation, keep it in the thread.

Try this

Next time you would ask a broad question about a codebase, say instead: "Use a subagent to investigate how X works and report back." Then run /context and compare against what the same question costs asked directly.

What goes wrong

Delegating with a prompt that assumes shared history. The subagent has none. A delegation that references "the approach we discussed" delegates nothing usable.

Expecting edits to be restored by rewind. Most subagent edits are not captured in your session's checkpoints, and the documentation's advice is to use git — see git hygiene.

Delegating the iterative part. Work that needs three rounds of correction costs more through a subagent than in the thread, because each round pays the startup cost again.

Accumulating agents. Every description loads every session. Prune them the way you prune CLAUDE.md.

How to check it worked

Compare what came back against what you would have needed to read to produce it. If the summary answers the question and your context barely moved, the delegation did its job. If you find yourself asking follow-ups the subagent could have answered had it known more, the delegation prompt was thin, and that is the thing to fix rather than the pattern.

Sources

  1. Subagents — Claude Code Docs Tier 1 2026-09-11
  2. Best practices for Claude Code — Anthropic Tier 1 2026-09-04