Building with AI

Code interpreters

4 min readintermediateUpdated 28 Sept 2026
1 · In one line

A code interpreter lets an AI model write a small program, run it in a sealed-off sandbox, and use the result in its answer.

1 · What it is

A language model can plan a maths problem step by step. But it often gets the sums wrong. A code interpreter gives the model a place to run code instead of computing in its head. The model writes a program, the program runs, and the output comes back to the model.

Google’s documentation gives an example question: add up the first fifty primes. The model writes Python code that finds the primes and adds them up, runs it, and reports the result. The code runs in a sandbox, a sealed-off, secure container.

OpenAI calls its version Code Interpreter, but to the model the tool is called python. Anthropic’s code execution tool can also run shell commands (typed instructions to the computer) and edit files. Google’s tool runs only Python.

One research paper, Program of Thoughts, reported an average gain of about 12 percent over chain-of-thought prompting, where the model does the reasoning and the sums in plain text.

2 · Why it exists

Language models can plan a solution, yet still slip when carrying it out.

Slips in the mathsModels often make logic and arithmetic mistakes when solving, even after breaking the problem into steps correctly.
Files need processingWorking with data files, such as spreadsheets, means actually processing them.
No way to testRunning the code lets the model see whether its program works and learn from the result.
3 · How it works

Follow one question through the sandbox.

The model plans and writes the code; the sandbox does the computing and reports back.
  1. 1 · askA person asks a question that needs calculation or a file to be processed.
  2. 2 · writeThe model writes a short Python program for the job.
  3. 3 · runThe program runs inside a sandbox, a sealed-off machine, and prints its output.
  4. 4 · retryIf the code fails, the model can fix it and run it again.
  5. 5 · answerThe model reads the output and writes its reply, sometimes with a chart or file attached.

The model decides what to compute; the sandbox does the computing.

4 · Where it's used
WhoWhat they askWhat it works with
Student“Solve 3x + 11 = 14 for x.”A short Python calculation
Analyst“Plot monthly sales from this CSV file.”An uploaded CSV turned into a graph image
Researcher“What are the average and spread (standard deviation) of these ten numbers?”A list of numbers
Designer“Crop and zoom into the corner of this photo.”An image processed with code
5 · What it solves, and what it doesn't
solves
  • Arithmetic is done by a Python interpreter instead of the model's guesswork.
  • It can read uploaded files and produce new ones, such as charts.
  • A model can keep rewriting failed code until it runs.
  • In one study, handing the computing to Python beat a chain-of-thought approach on the GSM8K maths word-problem test.
doesn't solve
  • The model still has to break the problem down correctly; the sandbox only runs what it is given.
  • Anthropic's sandbox has no internet, so only pre-installed libraries can be used.
  • Gemini stops code after 30 seconds, and an OpenAI sandbox expires after a period of disuse.
  • Turning it on can make other kinds of output, such as story writing, a little worse.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. docsCode Interpreter, OpenAI · read 28 Sept 2026
  2. docsCode execution tool, Anthropic · read 28 Sept 2026
  3. docsCode execution, Google AI for Developers · read 28 Sept 2026
  4. paperPAL: Program-aided Language Models, Gao et al. · read 28 Sept 2026
  5. paperProgram of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks, Chen et al., TMLR 2023 · read 28 Sept 2026