Building with AI

Tool use

5 min readintermediateUpdated 28 Sept 2026
1 · In one line

Tool use is a loop where a model proposes a call, a program outside the model runs it, and the result is written back.

1 · What it is

Language models can solve new tasks from just a few examples or from written instructions. They still struggle with basic work such as arithmetic or looking up a fact. Thoughts that stay inside the model use its internal representations. That reasoning is not grounded in the outside world.

The model proposes a call to an outside tool. A runtime that is not the model executes the call. The result is written back into the text, and generation continues. Toolformer learns which APIs to call, when to call them, and what arguments to pass. It also learns how to use the results in later tokens. The published model has 6.7 billion parameters. Its tools include a calculator. They also include a search engine, a translation system, and a calendar. One calculator input in the paper is 27 + 4 * 2. The same table lists 35 as the output.

While the model is writing, it continues until it expects an API response. The system then interrupts decoding, calls the API, and inserts the response before it goes on. The call might be another neural network, a Python script, or a retrieval system. On a fact-completion task, the model asked the question-answering tool in 98.1% of cases. ReAct makes the same kind of loop visible as text. It generates reasoning traces and actions in an interleaved way. Those traces help it update a plan and handle exceptions. For questions, the action is a simple Wikipedia API. Each cycle is a thought, an action, and an observation. A thought does not affect the external environment. An action can gather more information from a knowledge base or an environment. When the model plans to invoke a tool, the API response can include a tool use block. Anthropic names that piece a tool use block. The application executes the structured call. A second request carries the tool output back to the model. The model uses the result to answer the original question. It may also return more tool calls. At each step the useful signal is a tool result or the output of running code. Working agents are often models using tools from environmental feedback in a loop. ReAct set a cap of 7 steps on one question task and 5 steps on a fact-check task. Builders are told to run examples and look at the mistakes the model makes. Changing the arguments can make those mistakes harder.

2 · Why it exists

On their own, language models struggle with arithmetic and with looking up facts.

No fresh factsThey cannot access up-to-date information on recent events. They also tend to hallucinate facts.
Weak arithmeticThey lack the mathematical skill for precise calculations.
Actions stay outsideIssuing a refund or updating a ticket is work a program performs. The model returns a structured call that the application executes.
3 · How it works

Follow one calculator question through the loop.

The key step is the call the model writes for 27 + 4 * 2. The example output is 35. That text is written back before decoding continues.
  1. 1 · askThe walk-through uses the calculator input 27 + 4 * 2.
  2. 2 · proposeThe model decides which API to call, when to call it, and what arguments to pass.
  3. 3 · runA program outside the model executes the calculator call.
  4. 4 · returnThe system inserts the API response, and the model keeps generating from that text.
  5. 5 · repeatA maximum number of iterations can stop the loop.

The model writes the call. The runtime executes it and inserts the result before generation continues.

4 · Where it's used
WhoWhat they askWhat it works with
Student“What is 27 plus 4 times 2?”A calculator call whose input is that sum
Fact checker“What does the Wikipedia page say about this person?”A search action and the sentences that come back
Support desk“Refund this lost order and update the ticket”A refund action the program actually performs
Coding assistant“Why did this test fail after the folder changed?”The text a shell command returned
5 · What it solves, and what it doesn't
solves
  • A model can reach data outside its training data by deciding that a tool is needed.
  • Inserting the API response lets the next words depend on text the model did not write by itself.
  • A thought does not affect the external environment.
  • An action can gather information from a knowledge base or an environment.
doesn't solve
  • On a fact task, the model used a different tool in 0.7% of examples and no tool in 1.2%. One failure mode is a wrong reasoning trace, including a failure to break out of repeated steps.
  • If a required detail is missing, the model may infer a value. That behavior is not guaranteed. Relative file paths also made a coding agent fail after it left the root directory.
  • A refund or a ticket update is performed by a program. A tool can also come back with an error instead of an answer.
  • Search can return nothing useful. That miss was 23% of studied ReAct errors. The model can repeat the previous thought and action. The number of steps can be difficult or impossible to predict. Mistakes can compound. One setup allowed at most one API call per input.