Claude Code subagents are a context tool

This research library uses AI-assisted source research and drafting. Linked sources support product claims; analysis and proposed exercises are our interpretation. Unless an article documents a test and its results, do not read it as a hands-on review or an independently verified benchmark.
A subagent runs in its own context window and reports back a summary. That is the mechanism, and everything useful follows from it. The parent conversation never sees the forty files the subagent read. It sees the conclusion.
The popular framing is a team of specialists: a security agent, a testing agent, a docs agent. That framing produces disappointment, when a specialist label is mistaken for validated expertise. The configured model, tools, instructions and evidence determine its capabilities; the name alone does not.
Which is still valuable. An empty context and a narrow instruction is often exactly what a task needs.
When Claude Code subagents pay off
The pattern is: the work produces far more intermediate output than final answer.
- "Where in this repo do we handle refund idempotency?" The search reads thirty files. You want one paragraph back.
- Reviewing a diff against a checklist, where the review reasoning is long and the verdict is short.
- Running an exploratory build or test sweep whose logs are enormous and whose signal is one failing case.
- Two genuinely independent pieces of work you want running at once, like a frontend change and an unrelated migration.
The parallel case has a hard constraint most people learn the expensive way: the tasks must not touch the same files. Two subagents editing the same module produce a mess neither of them understands, and you get to untangle it.
When they make things worse
Anything requiring back-and-forth. If a delegated task needs decisions from the parent or user, define how it should report the ambiguity and pause. Otherwise it may guess instead of returning useful evidence.
Anything where the detail is the point. If you delegate "figure out why this test is flaky", you get a summary saying it is a timing issue, and the specific stack trace may be missing unless you ask for it or inspect the recorded output.
And small tasks. Spinning up a subagent to read one known file is slower and costs more than reading it. Compare the setup and handoff cost with simply reading the known file.
Writing one that behaves
Two things determine quality: the description, which decides when it fires, and the tool list, which decides what it can do.
Make the description name real triggers, not a job title. "Use when the user asks where something is implemented, or needs a codebase-wide search across naming conventions" beats "code exploration expert".
Then restrict tools. A research subagent should not have write access. If it only reads, it cannot half-edit a file and leave you with a partial change nobody asked for. Give it read and search tools. An unrestricted shell can write files and call external systems, so do not describe that setup as read-only; enforce any needed shell restrictions separately.
Finally, tell it what to return. Subagent output tends toward a wall of prose unless you specify the shape: Return the file path, the function name, and one sentence on how it works. No code blocks. Specifying the return format is the line that changes subagent output most.