Multi-Agent Systems
A multi-agent system splits work across several LLM-driven agents, each with its own instructions, tools, and context, coordinated by code or by another agent. A research system might have a lead agent that plans and several subagents that search different angles in parallel. A customer service system might route conversations between billing, technical, and account agents via handoffs. A coding agent might spawn subagents to explore a codebase without cluttering its own context.
Multiple agents can deliver real gains: parallelism, context isolation (each agent works in a clean window), and specialization. They also multiply cost, latency, and failure modes. A well-designed single agent often beats a poorly coordinated team, so the key skill is knowing when splitting work helps.
TL;DR
- Common architectures: orchestrator-subagents, handoffs (routing control between specialists), pipelines, and debate/voting.
- Main benefits: parallel exploration, isolated context windows, and specialized prompts and tools.
- Main costs: more tokens (often several times a single agent), coordination errors, and harder debugging.
- Subagents should return condensed results, not raw transcripts, to the orchestrator.
- Give subagents clear, self-contained task descriptions: objective, scope, output format, and boundaries.
- Use multiple agents for broad, parallelizable tasks; keep tightly coupled, sequential work in one agent.
Quick Example
An orchestrator that fans research out to parallel subagents and synthesizes the results:
Each subagent spends its own context on searching and reading; the orchestrator sees only the distilled briefs.
Core Concepts
Architectures
Frameworks support these directly: LangGraph (graphs of agents), the OpenAI Agents SDK (handoffs), the Claude Agent SDK (subagents), CrewAI, and AutoGen. See AI agent frameworks.
Why Multiple Agents Help
- Context isolation: each subagent's window holds only its task's material. The orchestrator avoids being flooded with search results, file contents, and dead ends, a direct answer to the limits of context windows.
- Parallelism: independent subtasks run simultaneously, cutting wall-clock time for broad tasks.
- Specialization: focused system prompts and small toolsets improve reliability on each subtask, and cheaper models can handle simpler roles.
- Separation of concerns: a reviewer agent that didn't write the code evaluates it more independently.
Anthropic reported that its multi-agent research system substantially outperformed a single agent on broad research tasks, while using many times more tokens. That trade is worth it for high-value tasks, and wasteful for simple ones.
Communication and State
- Task messages: the orchestrator's instructions to a subagent must be complete, since subagents don't share the orchestrator's context. Include the objective, constraints, what's already known, expected output format, and when to stop.
- Results: subagents return compressed, structured summaries (findings, sources, confidence, open questions), or write large artifacts to files or storage and return references.
- Shared state: a database, file system, or state object for artifacts, with clear ownership to avoid agents overwriting each other.
- Protocols: the Agent2Agent (A2A) protocol standardizes communication between agents from different vendors and frameworks, while the Model Context Protocol standardizes agent-to-tool connections.
Failure Modes
Debugging requires tracing every agent's prompts, tool calls, and outputs as one linked trace. See LLMOps.
When to Use Multiple Agents
Good fits:
- Breadth-first tasks with independent parts: research across many sources, auditing many files, comparing many options.
- Tasks whose exploration would overflow a single context window.
- Systems with clearly distinct domains and toolsets (customer service routing).
- Workflows benefiting from independent review.
Poor fits:
- Tightly coupled tasks where every step depends on shared, evolving context, such as most single-feature coding and multi-step reasoning chains.
- Low-value or latency-sensitive requests where token multiples aren't justified.
- Problems a single agent with good tools already solves reliably.
Best Practices
Start With One Agent
Build a strong single agent first. Split only when you hit concrete limits (context overflow, slow sequential exploration, conflicting tool needs) that multiple agents would solve.
Design Delegation Like an API
Treat subagent tasks as function contracts: inputs, expected output schema, scope limits, and effort guidance (how many tool calls are reasonable). Vague delegation is the most common failure.
Scale Effort to Task Complexity
Simple questions need one agent and a few tool calls; complex ones may need several subagents. Teach the orchestrator explicit heuristics for how many subagents to spawn, so it doesn't over-invest in easy tasks.
Evaluate End to End and per Agent
Measure final task success, cost, and latency against a single-agent baseline, and inspect individual agents' behavior to find weak links. See LLM evaluation.
Common Mistakes
Role-Play Teams Without a Reason
Creating "CEO", "PM", "engineer", and "QA" agents for a task one agent could do adds conversation overhead without new capability. Agents should exist because of isolation, parallelism, or specialization needs, not org-chart aesthetics.
Passing Full Transcripts Between Agents
Forwarding a subagent's entire conversation to the orchestrator defeats context isolation. Summarize, or store artifacts externally and pass references.
Parallel Agents Writing to the Same Resources
Concurrent subagents editing the same files or records produce conflicts and lost work. Partition ownership, or have subagents propose changes that one agent applies.
FAQ
Are multi-agent systems better than single agents?
For broad, parallelizable tasks, often yes: they explore more in less wall-clock time and keep contexts clean. For tightly coupled or simple tasks, a single agent is usually as good or better, and much cheaper. Decide with evaluations, not by default.
How do agents communicate with each other?
Most commonly through the orchestrating code: a lead agent calls subagents like tools, passing task descriptions and receiving results. Alternatives include shared state (files, databases), message queues, and standardized protocols like A2A for cross-system agent communication.
Why do multi-agent systems cost so much?
Each agent consumes its own tokens, re-reading instructions, calling tools, and reasoning, and orchestration adds more calls. Several agents working in parallel can easily use many times the tokens of one agent. Reserve them for tasks whose value justifies the spend, and use cheaper models for simpler roles.
What's the difference between a subagent and a tool?
Functionally, an orchestrator often invokes a subagent as a tool. The difference is that a subagent is itself an LLM loop with its own instructions, tools, and context, able to take many steps before returning. An ordinary tool executes a single deterministic operation.
Related Topics
- AI Agents — Single-agent foundations
- Agent Design Patterns — Orchestrator-workers and other patterns
- AI Agent Frameworks — Libraries for multi-agent orchestration
- Context Engineering — Why context isolation matters
- Model Context Protocol — Standard tool connections
- Tool Calling — Subagents invoked as tools