Multi-agent systems
A multi-agent system is a group of AI agents, each a language model using tools in its own loop, that work together on one job.
Here, an agent is a language model that works in a loop. It picks a next move. It calls a tool, such as web search, to act or look something up. Then it reads what came back and chooses again. A multi-agent system runs several of these loops and links them up. OpenAI’s Agents SDK, a toolkit for building agents, sets each one up with three things: its instructions, the tools it may call, and handoffs, which let it pass work to another agent.
There are a few common wirings. A manager can treat other agents like tools. It gets one piece back from a specialist, then writes the final reply itself. A handoff works the other way round. The specialist takes over and replies for the rest of that turn. Or a plain program can set the order. It can feed one agent’s output into the next. It can also run several side by side on tasks that do not depend on each other. Setting the order in code makes run time and cost easier to predict.
Research papers have tried other shapes. In a CAMEL demo, one chat agent plays a stock trader and keeps giving instructions to another that plays a programmer, and together they work on a trading bot. AutoGen split the work between two agents. An assistant writes code. A “user proxy” stands in for the person: it runs that code or asks a real human for input. In multiagent debate, several copies of one model each answer a question. Each reads the others’ answers and revises its own, and they usually end up agreeing. The authors found better maths and strategy reasoning, and fewer invented facts, known as hallucinations. A Berkeley study later sorted the ways these systems fail into three groups. They were system design, agents falling out of step, and checking whether the task was really done.
So when is one agent better? Think of a group project at school. Splitting it up helps when each person can research a different chapter on their own. It hurts when everyone needs the same notes or waits on each other. Agent teams work the same way. Models read and write text in tokens, small chunks of text. A single agent already uses about four times the tokens of a normal chat. A team uses about fifteen times, so the job has to be worth that bill. Agents are also not yet good at handing work to each other on the fly. Teams earn their cost on valuable jobs that split into many parallel parts, need more reading than one agent’s context window (the text it can hold in view at once) can take, or use many complicated tools.
One agent working alone runs into three walls on big, open-ended jobs.
Follow one research question through Anthropic's Research system.
- 1 · planA lead agent works out an approach and stores the plan in a separate memory, so it is not lost if the conversation outgrows the context window.
- 2 · splitIt starts several subagents at once and writes each a brief: what to find, how to report it, which tools to try and where its part of the job stops.
- 3 · searchEach subagent explores its slice with a fresh context window, then returns a short summary of what matters instead of everything it read.
- 4 · combineThe lead combines the summaries and, if something is still missing, sends out more subagents. A citation agent then matches each claim to the source it came from.
Tokens are the small units a model reads and writes: whole words, parts of words or single characters. Teams of agents spend far more of them than one agent does.
| Who | What they ask | What it works with |
|---|---|---|
| Research assistant | “Who sits on the boards of these companies?” | Parallel subagents, each taking a slice of the list |
| Blog writer | “Can you research, outline and draft this post, then critique it?” | A chain of agents, each taking the last one's output as its input |
| Developer | “Can you plot this stock data and fix the script if it fails?” | An assistant agent that writes code and a proxy agent that runs it |
| Maths homework check | “What is the right answer to this word problem?” | Several model copies that debate the answer over rounds |
- On Anthropic's internal research test, a Claude Opus 4 lead with Claude Sonnet 4 subagents scored 90.2% higher than a single Opus 4 agent working alone.
- It suits questions that spread out into many separate threads you can chase at the same time.
- Running 3 to 5 subagents at once, each calling 3 or more tools in parallel, shortened complex research runs by as much as 90%.
- A front-desk agent, called a triage agent, can pass each chat to the matching specialist, so every specialist keeps a narrow, focused set of instructions.
- It is expensive. Anthropic measured multi-agent runs spending about 15 times as many tokens as ordinary chats.
- Jobs where every agent needs the same context, or where agents depend on each other, fit poorly. Most coding tasks have fewer truly parallel parts than research.
- More agents do not guarantee success. A study of 7 open-source multi-agent systems found failure rates from 41% to 86.7%.
- Coordination is fragile. Early versions of Anthropic's system sometimes started 50 subagents to answer an easy question, and a small tweak to the lead agent could shift what the subagents did in ways nobody predicted.
Sources used
This explainer is written in original language. The links below support its factual claims.
- officialHow we built our multi-agent research system, Anthropic · read 28 Sept 2026
- docsAgent orchestration, OpenAI Agents SDK · read 28 Sept 2026
- paperAutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation, Wu et al., Microsoft Research · read 28 Sept 2026
- paperCAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society, Li et al., KAUST · read 28 Sept 2026
- paperImproving Factuality and Reasoning in Language Models through Multiagent Debate, Du et al., MIT and Google Brain · read 28 Sept 2026
- paperWhy Do Multi-Agent LLM Systems Fail?, Cemri et al., UC Berkeley, NeurIPS 2025 · read 28 Sept 2026
- docsGlossary, Anthropic · read 28 Sept 2026