Building with AI

Multi-agent systems

5 min readintermediateUpdated 28 Sept 2026
1 · In one line

A multi-agent system is a group of AI agents, each a language model using tools in its own loop, that work together on one job.

1 · What it is

Here, an agent is a language model that works in a loop. It picks a next move. It calls a tool, such as web search, to act or look something up. Then it reads what came back and chooses again. A multi-agent system runs several of these loops and links them up. OpenAI’s Agents SDK, a toolkit for building agents, sets each one up with three things: its instructions, the tools it may call, and handoffs, which let it pass work to another agent.

There are a few common wirings. A manager can treat other agents like tools. It gets one piece back from a specialist, then writes the final reply itself. A handoff works the other way round. The specialist takes over and replies for the rest of that turn. Or a plain program can set the order. It can feed one agent’s output into the next. It can also run several side by side on tasks that do not depend on each other. Setting the order in code makes run time and cost easier to predict.

Research papers have tried other shapes. In a CAMEL demo, one chat agent plays a stock trader and keeps giving instructions to another that plays a programmer, and together they work on a trading bot. AutoGen split the work between two agents. An assistant writes code. A “user proxy” stands in for the person: it runs that code or asks a real human for input. In multiagent debate, several copies of one model each answer a question. Each reads the others’ answers and revises its own, and they usually end up agreeing. The authors found better maths and strategy reasoning, and fewer invented facts, known as hallucinations. A Berkeley study later sorted the ways these systems fail into three groups. They were system design, agents falling out of step, and checking whether the task was really done.

So when is one agent better? Think of a group project at school. Splitting it up helps when each person can research a different chapter on their own. It hurts when everyone needs the same notes or waits on each other. Agent teams work the same way. Models read and write text in tokens, small chunks of text. A single agent already uses about four times the tokens of a normal chat. A team uses about fifteen times, so the job has to be worth that bill. Agents are also not yet good at handing work to each other on the fly. Teams earn their cost on valuable jobs that split into many parallel parts, need more reading than one agent’s context window (the text it can hold in view at once) can take, or use many complicated tools.

2 · Why it exists

One agent working alone runs into three walls on big, open-ended jobs.

Too much to holdSome research jobs gather more material than fits in one context window, the text a model can read at once.
One step at a timeIn one of Anthropic's tests, a lone agent searching one thing after another failed to find the answer to a broad lookup, while a team of subagents split it up and got it right.
One line of thoughtIn research, what you find early shapes what you look for next. Subagents that each work with their own instructions, tools and search route explore more directions and get stuck on the first lead less often.
3 · How it works

Follow one research question through Anthropic's Research system.

The lead agent writes each subagent a brief, lets them search in parallel, and reads back only their short summaries.
  1. 1 · planA lead agent works out an approach and stores the plan in a separate memory, so it is not lost if the conversation outgrows the context window.
  2. 2 · splitIt starts several subagents at once and writes each a brief: what to find, how to report it, which tools to try and where its part of the job stops.
  3. 3 · searchEach subagent explores its slice with a fresh context window, then returns a short summary of what matters instead of everything it read.
  4. 4 · combineThe lead combines the summaries and, if something is still missing, sends out more subagents. A citation agent then matches each claim to the source it came from.

Tokens are the small units a model reads and writes: whole words, parts of words or single characters. Teams of agents spend far more of them than one agent does.

4 · Where it's used
WhoWhat they askWhat it works with
Research assistant“Who sits on the boards of these companies?”Parallel subagents, each taking a slice of the list
Blog writer“Can you research, outline and draft this post, then critique it?”A chain of agents, each taking the last one's output as its input
Developer“Can you plot this stock data and fix the script if it fails?”An assistant agent that writes code and a proxy agent that runs it
Maths homework check“What is the right answer to this word problem?”Several model copies that debate the answer over rounds
5 · What it solves, and what it doesn't
solves
  • On Anthropic's internal research test, a Claude Opus 4 lead with Claude Sonnet 4 subagents scored 90.2% higher than a single Opus 4 agent working alone.
  • It suits questions that spread out into many separate threads you can chase at the same time.
  • Running 3 to 5 subagents at once, each calling 3 or more tools in parallel, shortened complex research runs by as much as 90%.
  • A front-desk agent, called a triage agent, can pass each chat to the matching specialist, so every specialist keeps a narrow, focused set of instructions.
doesn't solve
  • It is expensive. Anthropic measured multi-agent runs spending about 15 times as many tokens as ordinary chats.
  • Jobs where every agent needs the same context, or where agents depend on each other, fit poorly. Most coding tasks have fewer truly parallel parts than research.
  • More agents do not guarantee success. A study of 7 open-source multi-agent systems found failure rates from 41% to 86.7%.
  • Coordination is fragile. Early versions of Anthropic's system sometimes started 50 subagents to answer an easy question, and a small tweak to the lead agent could shift what the subagents did in ways nobody predicted.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. officialHow we built our multi-agent research system, Anthropic · read 28 Sept 2026
  2. docsAgent orchestration, OpenAI Agents SDK · read 28 Sept 2026
  3. paperAutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation, Wu et al., Microsoft Research · read 28 Sept 2026
  4. paperCAMEL: Communicative Agents for "Mind" Exploration of Large Language Model Society, Li et al., KAUST · read 28 Sept 2026
  5. paperImproving Factuality and Reasoning in Language Models through Multiagent Debate, Du et al., MIT and Google Brain · read 28 Sept 2026
  6. paperWhy Do Multi-Agent LLM Systems Fail?, Cemri et al., UC Berkeley, NeurIPS 2025 · read 28 Sept 2026
  7. docsGlossary, Anthropic · read 28 Sept 2026