Subagents in Coding Agents, Explained
Published October 7, 2026 · by the AQ team
A subagent is a worker agent that a coding agent spawns to handle part of a task: it runs with its own context window, does its delegated work (exploring a codebase, running tests, reviewing a diff), and returns a summary to the parent agent instead of flooding the main conversation with raw output. As of October 2026 all four major harnesses ship them: Claude Code defines them as Markdown files and spawns up to 20 concurrently by default, Codex ships built-in explorer and worker agents plus custom agents as TOML files, Cursor runs built-in Explore, Bash and Browser subagents plus custom ones, and the Devin CLI delegates through explore and general profiles. The mechanics differ enough, and multiply token spend enough, that it is worth knowing exactly how each harness spawns, scopes and bills them.
What problem do subagents solve?
The main agent's conversation is where you state requirements, make decisions, and read final answers. Exploration output, test logs and stack traces are noise there: they bury what matters (context pollution) and degrade the model's performance as the window fills (context rot, a term OpenAI's own Codex documentation uses as of October 2026). A context window is the bounded amount of text a model can consider at once, and a long agent session consumes it fast; context windows for coding agents covers those mechanics.
Subagents move the noisy work into a separate context: the parent writes a task prompt, the subagent burns its own window on file reads and command output, and only a distilled result returns. The main conversation stays reliable longer, and independent subtasks can run in parallel. The cost: each subagent is its own model session, so five parallel subagents consume roughly five times the tokens of one agent doing the same work in sequence (Cursor's documentation states this multiplier plainly).
How Claude Code spawns and scopes subagents
Claude Code subagents are Markdown files with YAML frontmatter: a name, a description, an optional tools allowlist, an optional model, then a system prompt. Project subagents live in .claude/agents/ (checked into the repo so the team shares them) and personal ones in ~/.claude/agents/. Claude delegates automatically based on each description, or you force a specific one by naming it. Each subagent starts with a fresh, isolated context window: no conversation history, no files the parent already read, only its task prompt, the project's CLAUDE.md files and a git status snapshot. A fork is the exception: it inherits the whole conversation and shares the parent's prompt cache.
# .claude/agents/code-reviewer.md
---
name: code-reviewer
description: Reviews code for quality and best practices
tools: Read, Glob, Grep
model: sonnet
---
You are a code reviewer. Analyze the diff and report
specific, actionable findings.
Scoping, as of October 2026: subagents can spawn their own subagents up to three layers below the main conversation, 20 can run concurrently by default (both configurable), and each can pin its own model. Foreground subagents block the conversation; background ones run while you keep working, their permission prompts surfacing in the main session under the asking subagent's name. Agent teams are a separate, newer Claude Code feature in which teammates message each other peer to peer instead of reporting up to one parent; Claude Code agent teams covers when that shape beats a subagent tree.
How Codex spawns and scopes subagents
Codex subagent workflows reached general availability in March 2026 and are on by default in current releases as of October 2026. Three built-in agents ship (default, worker for implementation, explorer for read-heavy exploration), spawned when you ask directly ("spawn one subagent per review point, wait for all three") or when AGENTS.md or skill instructions request delegation. Custom agents are TOML files under ~/.codex/agents/ (personal) or .codex/agents/ (project), each requiring a name, a description and developer_instructions, optionally pinning a model, reasoning effort or sandbox mode:
# .codex/agents/reviewer.toml
name = "reviewer"
description = "PR reviewer focused on correctness and security."
sandbox_mode = "read-only"
developer_instructions = """
Review code like an owner. Lead with concrete findings.
"""
Scoping lives under the [agents] table in config.toml: max_concurrent_threads_per_session caps open agent threads, and a default subagent model and reasoning effort can be set fleet-wide. Subagents inherit the parent turn's sandbox and approval policy, and an approval request from a background thread names its source thread. OpenAI's own starting advice: parallelize read-heavy work (exploration, tests, triage), and be careful with parallel write-heavy work, where agents editing at once create conflicts.
How Cursor spawns and scopes subagents
Cursor shipped subagents in version 2.4 (January 2026). Three built-ins run automatically: Explore for codebase search (on a faster model, so many searches run in parallel cheaply), Bash for command output, and Browser for browser automation. Custom subagents are Markdown files with YAML frontmatter in .cursor/agents/ or ~/.cursor/agents/, and as of October 2026 Cursor also reads .claude/agents/ and .codex/agents/, so one checked-in definition can serve three harnesses. Fields include a model (inherit by default, or a pinned ID with per-model effort and context parameters), a readonly flag, and is_background.
Two scoping details stand out. Nesting: since Cursor 2.5, the main agent and its direct subagents can launch children, but a grandchild cannot spawn further ones. Isolation: subagents share the parent's checkout by default, so parallel editors can overwrite each other; ask for isolation and each subagent gets its own git worktree (a separate working directory on the same clone) or its own cloud VM, with changes kept on per-subagent branches until the parent merges them. Billing follows the model that actually ran, at that model's list price, even when the parent chat is on Auto.
How Devin spawns and scopes subagents
The Devin CLI delegates through profiles rather than free model choice. subagent_explore is read-only research on a router-selected default model; subagent_general can edit code and always runs on the parent's model, so a premium parent model makes every general subagent a full extra session at that rate. Custom profiles (experimental as of October 2026) are Markdown files under .devin/agents/ with a name, description, allowed-tools list and optional pinned model, the only way to run a write-capable subagent on a cheaper model. Nesting is off by default unless a custom profile opts in with a max-nesting depth, background subagents cannot ask for new permissions (a failed one resumes in the foreground to grant them), and administrators can pin the default subagent model or disable subagents entirely.
The four harnesses at a glance (October 2026)
| Harness | Custom definition | Parallelism | Nesting | Watching progress |
|---|---|---|---|---|
| Claude Code | Markdown + YAML in .claude/agents/ | 20 concurrent by default, configurable | 3 layers below the main conversation by default | /tasks panel; open any subagent's live transcript; transcripts persist on disk |
| Codex | TOML in .codex/agents/ | Capped by max_concurrent_threads_per_session | Allowed; depth configurable | /agent in the CLI switches between threads; IDE panel above the composer |
| Cursor | Markdown + YAML in .cursor/agents/ (also reads .claude/ and .codex/) | Multiple parallel launches; optional per-subagent worktree or VM | Children yes, grandchildren cannot spawn (since 2.5) | Task cards; background output written to ~/.cursor/subagents/ |
| Devin CLI | Markdown profiles in .devin/agents/ (experimental) | Foreground or background per subagent | Off by default; opt-in via max-nesting | Subagent panel with profile, status, elapsed time and tool count; Ctrl+B to background |
When parallel subagents help, and when they multiply cost
Subagents earn their tokens when the work is read-heavy, independent and summarizable: mapping an unfamiliar codebase, reviewing one diff from three angles, triaging a test suite, digesting a huge log. They cost more than they return when the task is small (a fresh subagent re-gathers context the parent already had, so it is often slower), when subtasks depend on each other (coordination overhead eats the parallelism), and when several write-capable subagents edit one checkout. Every vendor's docs say some version of the same thing: each subagent is its own session with its own context window and inference calls, so fan-out multiplies spend. Coding agent token costs for teams covers budgeting for that.
Subagents are not the same as several top-level agents
A subagent tree is one task, one owner: the parent decomposes the work, children report back, and you get one answer and one diff. Several top-level agents in parallel are several tasks with independent lifetimes: each has its own conversation you can steer directly, its own branch and pull request, and each survives the others failing. Use subagents to make one hard task tractable; use separate top-level sessions, each in its own git worktree, for genuinely separate tasks like three unrelated bug fixes, because a subagent's transcript is buried inside its parent session and its lifetime ends with the parent. Running multiple AI coding agents in parallel covers the top-level shape in depth.
Where AQ fits
AQ is the multiplayer coding harness where engineering teams run AI coding agents like Claude Code and Codex together: shared live terminals, a code editor, and app previews, in your own cloud. Subagents keep working exactly as each CLI defines them, because AQ runs the real CLIs (Claude Code, Codex, Cursor Agent, Kimi, Grok, or plain shells) in persistent tmux sessions on your team's VM, streamed live to the browser.
What AQ adds is the top-level layer this guide distinguishes from subagents. Each workspace gets its own isolated git worktree with dependencies auto-installed, so parallel tasks never collide; sessions survive a closed laptop and resume from any device, so a fan-out you started at 6pm is still running at 9am; and teammates open the same workspace and watch the same live session, subagent panels included. Agents authenticate with each person's own Claude or OpenAI login (AQ never marks up usage on your own subscriptions). The Free plan is a personal sandbox for one person with nothing to install and no time limit; the Team plan is $50 per user per month early access (standard $200, billed monthly), covering VMs you connect from your own cloud or a dedicated always-on AQ-managed VM, rate locked for your first 12 months.
Frequently asked questions
What is a subagent in Claude Code or Codex?
A worker agent the main agent spawns to handle part of a task in its own context window. In Claude Code it is defined as a Markdown file with YAML frontmatter in .claude/agents/; in Codex it is one of the built-in default, worker or explorer agents, or a custom TOML file in .codex/agents/. Either way it does its delegated work and returns a summary to the parent instead of raw output.
Do subagents cost more tokens?
Yes. Each subagent is its own model session with its own context window and inference calls, so the spend adds on top of the parent's. Cursor's documentation puts it directly: five subagents in parallel use roughly five times the tokens of a single agent. Devin's general subagents additionally run on the parent's model, so a premium parent model multiplies at the premium rate. Fan out when the parallelism or context isolation is worth it, not by default.
Can a subagent spawn its own subagents?
Depends on the harness, as of October 2026. Claude Code allows nesting up to three layers below the main conversation by default. Codex allows it with a configurable depth. Cursor (since 2.5) lets the main agent and its direct subagents launch children, but a grandchild cannot spawn further. The Devin CLI disables nested spawning by default; a custom profile can opt in with a max-nesting value.
How do I watch what a subagent is doing?
Every harness ships a panel. Claude Code lists running subagents below the prompt and in /tasks, and you can open a live transcript and even message a subagent to steer it. Codex uses /agent in the CLI to switch between agent threads, with a background-agent panel in the IDE. Cursor shows task cards and writes background subagent output to ~/.cursor/subagents/. Devin's panel shows each subagent's profile, status, elapsed time and tool call count, with Ctrl+B to move one to the background.
When should I use multiple top-level agents instead of subagents?
When the tasks are genuinely separate. Subagents decompose one task and report to one parent, so they share its lifetime and produce one combined result. Three unrelated bug fixes want three top-level sessions, each in its own git worktree and branch, each steerable and reviewable on its own. Subagents make one hard task tractable; parallel top-level agents make many tasks concurrent.
Are subagents the same as Claude Code agent teams?
No. Subagents report results only to the parent that spawned them. Agent team teammates are peer agents that message each other directly and claim tasks from a shared list, and you can talk to any teammate without going through the lead. Teams fit work that needs mid-task coordination between workers; subagents fit work that decomposes cleanly into report-back pieces.