ReActBuilding with AI

ReAct (Reason + Act)

5 min readintermediateUpdated 28 Sept 2026
1 · In one line

ReAct is a prompting method in which a language model alternates written thoughts with tool actions, reading each result before it picks the next step.

1 · What it is

ReAct is a way of prompting a language model so it takes turns. A prompt is the text you give the model to start it off. With ReAct, the model thinks, does one thing, looks at what came back, then thinks again. Researchers from Princeton University and Google’s Brain team presented it at ICLR 2023, a major AI conference.

Some steps are actions that touch the outside world, like a search. Others are thoughts: sentences of reasoning that touch nothing, like notes to itself.

The paper used a big model, PaLM-540B, left “frozen” (its inner settings unchanged). The prompt showed it a few worked examples: six for HotpotQA, a question set, and three for FEVER, a fact-checking set.

The model had three moves: search, lookup (like Ctrl+F on a page) and finish.

In FEVER, the model labels a claim as supported, refuted, or not enough information. Two claims can differ by one small detail, so correct facts matter.

In the figure, the model never guesses the years. It reads 1844 and 1989 from search results, then compares them.

Why not just act, or just think?

The paper tested ReAct against simpler set-ups. “Act” is ReAct with the thoughts taken out, so the model only issues actions and reads results. “Chain-of-thought” (CoT) is the opposite: thoughts only, with no actions and no results from outside.

ReAct beat Act on both HotpotQA and FEVER. Thinking helps the model choose better actions. It helps most with the final answer.

It also beat Act on ALFWorld, a text game, and WebShop, a shopping site.

Against chain-of-thought, the picture is mixed. ReAct won on FEVER. On HotpotQA it came out slightly behind. Chain-of-thought lays out its reasoning better, but it often makes up facts. ReAct sticks to facts it looked up. But its fixed rhythm makes it less flexible.

Mixing the two

Each method is strong where the other is weak, so the authors combined them. They used self-consistency: ask chain-of-thought 21 times and keep the most common answer. Then simple rules switch methods:

  • ReAct first, then CoT. If ReAct has not answered within a set number of steps (7 for HotpotQA, 5 for FEVER), fall back to chain-of-thought.
  • CoT first, then ReAct. If the most common chain-of-thought answer shows up in fewer than half the samples, the model’s own knowledge may be shaky, so fall back to ReAct and search.

These mixes gave the best prompting results. ReAct-then-CoT was best on HotpotQA, and CoT-then-ReAct was best on FEVER. Both matched self-consistency’s 21 samples using only 3 to 5.

Where you see it now

The same loop now sits inside agent libraries. Hugging Face’s smolagents docs call ReAct the main way people build agents today. Every smolagents agent is built on one shared class that packages the ReAct loop.

OpenAI’s Agents SDK (a toolkit for building agents) runs a similar loop. It runs the tools the model asks for. It adds the results. Then it calls the model again. It stops at a limit called max_turns. LangChain’s docs put it simply: an agent keeps using tools, round after round, until its job is finished.

Anthropic stresses getting real results, such as tool outputs, at every step. It also notes a common stopping rule to stay in control: cap how many times the loop may run.

2 · Why it exists

A model that only reasons can state wrong facts, and one that only acts can fail to finish the job.

No outside checkChain-of-thought prompting has the model write out several reasoning steps before it answers. But it uses only what the model already knows. It cannot look anything up or update what it knows.
Invented factsOn HotpotQA, a question set that needs two or more Wikipedia passages, the authors checked 50 sampled chain-of-thought failures by hand. Hallucination, where the reasoning or its facts are made up, caused 56% of them, more than any other cause.
Acting without a planAn agent that only issues actions can make the same searches as ReAct and still fail to put the final answer together.
3 · How it works

Follow one question through three thought, action and observation turns.

Each observation becomes part of what the model reads next, so the next thought can use it to choose the next action.
  1. 1 · promptThe prompt shows a few worked examples, each a human-written trail of thoughts, actions and observations.
  2. 2 · thinkThe model writes a thought, which changes nothing outside and only adds useful notes to its context.
  3. 3 · actIt then issues an action such as search[entity], and the tool's reply comes back as an observation.
  4. 4 · adjustThe next thought reads that observation and uses it to choose what to retrieve next or to change the plan.
  5. 5 · finishThe loop ends when the model issues finish[answer].

Reason to act, act to reason: thoughts pick the next tool call, and tool results feed new information back into the thoughts.

4 · Where it's used
WhoWhat they askWhat it works with
Research assistant“Which of these two magazines was started first?”One search per magazine, then a comparison
Shopping agent (WebShop test)“Find a nightstand with drawers and a nickel finish under 140 dollars”A search for nightstand drawers, then product option buttons
Household game agent (ALFWorld test)“Examine the paper under the desk lamp”Text moves such as go to coffeetable 1 and take paper 2
5 · What it solves, and what it doesn't
solves
  • ReAct beat chain-of-thought on FEVER fact checking, 60.9 to 56.3 accuracy.
  • Right answers reached through made-up facts were rarer, 6% of ReAct's sampled successes against 14% for chain-of-thought.
  • On the ALFWorld text game, it beat an agent trained on many expert example runs by 34 points of success rate. On the WebShop shopping site, it beat agents trained on human example runs by 10 points. It needed only one or two examples in its prompt.
  • People can tell which facts came from a tool, and they can correct the agent by editing a thought.
doesn't solve
  • On HotpotQA it scored slightly below chain-of-thought, 27.4 to 29.4.
  • A useless search result derails it. Such searches caused 23% of its sampled errors.
  • It can get stuck repeating its previous thoughts and actions.
  • Harder tasks need more examples. Those can overflow the limit on how much text the prompt can hold.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. paperReAct: Synergizing Reasoning and Acting in Language Models, Yao et al., ICLR 2023 · read 28 Sept 2026
  2. officialReAct: Synergizing Reasoning and Acting in Language Models (project page), ReAct authors · read 28 Sept 2026
  3. officialReAct: Synergizing Reasoning and Acting in Language Models, Google Research · read 28 Sept 2026
  4. repoysymyth/ReAct, ReAct authors · read 28 Sept 2026
  5. docsHow do multi-step agents work?, Hugging Face · read 28 Sept 2026
  6. docsRunning agents, OpenAI · read 28 Sept 2026
  7. officialBuilding effective agents, Anthropic · read 28 Sept 2026
  8. docsAgents, LangChain · read 28 Sept 2026