Code interpreters
A code interpreter lets an AI model write a small program, run it in a sealed-off sandbox, and use the result in its answer.
A language model can plan a maths problem step by step. But it often gets the sums wrong. A code interpreter gives the model a place to run code instead of computing in its head. The model writes a program, the program runs, and the output comes back to the model.
Google’s documentation gives an example question: add up the first fifty primes. The model writes Python code that finds the primes and adds them up, runs it, and reports the result. The code runs in a sandbox, a sealed-off, secure container.
OpenAI calls its version Code Interpreter, but to the model the tool is called python. Anthropic’s code execution tool can also run shell commands (typed instructions to the computer) and edit files. Google’s tool runs only Python.
One research paper, Program of Thoughts, reported an average gain of about 12 percent over chain-of-thought prompting, where the model does the reasoning and the sums in plain text.
Language models can plan a solution, yet still slip when carrying it out.
Follow one question through the sandbox.
- 1 · askA person asks a question that needs calculation or a file to be processed.
- 2 · writeThe model writes a short Python program for the job.
- 3 · runThe program runs inside a sandbox, a sealed-off machine, and prints its output.
- 4 · retryIf the code fails, the model can fix it and run it again.
- 5 · answerThe model reads the output and writes its reply, sometimes with a chart or file attached.
The model decides what to compute; the sandbox does the computing.
| Who | What they ask | What it works with |
|---|---|---|
| Student | “Solve 3x + 11 = 14 for x.” | A short Python calculation |
| Analyst | “Plot monthly sales from this CSV file.” | An uploaded CSV turned into a graph image |
| Researcher | “What are the average and spread (standard deviation) of these ten numbers?” | A list of numbers |
| Designer | “Crop and zoom into the corner of this photo.” | An image processed with code |
- Arithmetic is done by a Python interpreter instead of the model's guesswork.
- It can read uploaded files and produce new ones, such as charts.
- A model can keep rewriting failed code until it runs.
- In one study, handing the computing to Python beat a chain-of-thought approach on the GSM8K maths word-problem test.
- The model still has to break the problem down correctly; the sandbox only runs what it is given.
- Anthropic's sandbox has no internet, so only pre-installed libraries can be used.
- Gemini stops code after 30 seconds, and an OpenAI sandbox expires after a period of disuse.
- Turning it on can make other kinds of output, such as story writing, a little worse.
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsCode Interpreter, OpenAI · read 28 Sept 2026
- docsCode execution tool, Anthropic · read 28 Sept 2026
- docsCode execution, Google AI for Developers · read 28 Sept 2026
- paperPAL: Program-aided Language Models, Gao et al. · read 28 Sept 2026
- paperProgram of Thoughts Prompting: Disentangling Computation from Reasoning for Numerical Reasoning Tasks, Chen et al., TMLR 2023 · read 28 Sept 2026