Prompt chaining
In prompt chaining, you give a model one small job at a time in a set order, and each answer becomes the starting material for the next job.
Prompt chaining breaks one task into a set sequence of model calls. Each prompt covers a single step, and the reply to one prompt becomes part of the next. Whatever the last prompt returns is the answer you keep. Google’s cloud prompting guide says smaller prompts can make results easier to control, debug and get right. Google’s Gemini API guide sets chaining apart from aggregation, which runs several prompts side by side on separate slices of the data and then merges what they return.
The approach was studied in a 2022 paper at the CHI Conference on Human Factors in Computing Systems. Its worked example rewrites critical feedback about Alex’s presentation into something kinder and more useful. One prompt on the same model produced a mostly impersonal rewrite. That rewrite did not give concrete suggestions for all three problems. The chained version, shown in the diagram below, gave the paragraph-writing call a ready list of problems and fixes.
The paper then tested chaining on people. Twenty people each did tasks with both a chained tool and an ordinary single-prompt tool, and the same language model sat behind both. Because every person tried both tools, each one served as their own comparison. Two separate judges then read the results: one preferred the chained version 85% of the time, the other 80%. Chaining took slightly longer on average, about 14.6 minutes against 12.4. The gap was too small to rule out chance.
Anthropic’s current guide says that newer Claude models, helped by adaptive thinking and subagents, handle most multi-step reasoning internally. With adaptive thinking, the model decides for itself when and how much to think. Subagents are specialised helpers the model can hand parts of the work to. The guide still recommends a hand-built chain when you need to see the results in the middle, or when the steps must run in a fixed order. Its most common example is self-correction. One call writes a draft, a second grades that draft against a checklist, and a third improves it using the grades. OpenAI’s Agents SDK is a lightweight code package for building agent apps. The SDK describes the same move for agents, reshaping what one agent returns so the next agent can start from it. Its docs add that when your code, not a model, decides the order, speed and cost are easier to predict.
One prompt that asks for everything at once can fail in ways that are hard to see or fix.
Follow one piece of feedback through three model calls.
- 1 · splitThe first call reads the original feedback and pulls out each problem as a separate item.
- 2 · checkYour code receives that list and can run a check on it, or let a person edit it, before the next call.
- 3 · ideateCode drops each problem into a fixed prompt (Problem: {item}, Suggestion:) and the second call brainstorms fixes for it.
- 4 · composeThe third call gets both lists, the problems and the suggestions, and writes them up as one friendly paragraph, the final output.
Each call gets an easier task, and fixed code decides what flows into the next prompt.
| Who | What they ask | What it works with |
|---|---|---|
| Marketing team | “Write launch copy for this product, then translate it into Spanish” | Call 1 writes the copy, call 2 translates that copy |
| Telecom support analysts | “What are customers complaining about, and how could we fix each kind of issue?” | Pull out issues, sort them into categories, then suggest fixes for each category |
| Report writer | “Draft the quarterly report from our notes” | Write an outline, check it against the criteria, then write from the outline |
| Content team | “Write a blog post about our new API” | Research, outline, draft, critique, then improve |
- Each call gets an easier task. Anthropic describes the trade-off as waiting longer for an answer (more latency) in return for getting it right more often.
- Your code sends each step to the model as its own request, so it can record that step, score it, or pick a different next step depending on the answer.
- In the AI Chains study, people debugged surprising outputs by testing one step of a chain on its own.
- Every extra model call costs its own share of computing, and the chain waits for each one in turn.
- An early mistake can spread. In the paper's example, if the first call pulls out the wrong problems, the final paragraph suffers too.
- It suits tasks that split cleanly into fixed steps. When the steps cannot be known in advance, Anthropic points to other designs, such as orchestrator-workers, where a lead model splits up the work as it goes and hands pieces to helper models.
- Handing outputs on is real work. The PromptChainer study found that people building chains had to plan carefully how each step's output gets reshaped for the next one.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperAI Chains: Transparent and Controllable Human-AI Interaction by Chaining Large Language Model Prompts, Wu, Terry and Cai, CHI 2022 · read 28 Sept 2026
- paperPromptChainer: Chaining Large Language Model Prompts through Visual Programming, Wu et al., CHI 2022 Late-Breaking Work · read 28 Sept 2026
- officialBuilding effective agents, Anthropic · read 28 Sept 2026
- docsPrompting best practices, Anthropic · read 28 Sept 2026
- docsBreak down complex tasks into simpler prompts, Google Cloud · read 28 Sept 2026
- docsPrompt design strategies, Google AI for Developers · read 28 Sept 2026
- docsAgent orchestration, OpenAI Agents SDK · read 28 Sept 2026
- docsOpenAI Agents SDK, OpenAI · read 28 Sept 2026