Sandboxing AI agents
Sandboxing an agent means running its commands inside a fence, set up by the operating system, that limits which files and websites they can reach.
Coding agents such as Claude Code and Codex edit files and run commands on your computer to check their own work. Sandboxing wraps those commands in two layers. Filesystem isolation decides which paths a command can read and write. Network isolation decides which domains it can reach. In Claude Code, no domain is allowed until you approve it, and writes are limited to the working folder and a temporary folder unless you add more. Codex also sandboxes local commands, with no network by default and writes limited to the active workspace.
Both layers matter. Without the network layer, a hijacked agent could send out private files such as SSH keys. Without the filesystem layer, it could escape the sandbox and reach the network anyway. So the sandbox refuses to let commands change the agent’s own settings files, even inside the project folder. Editing them could give the agent new permissions.
Some sandboxes isolate much more than one command. Codex cloud runs tasks in isolated containers managed by OpenAI, which go offline after setup unless you turn internet access on. A container is a sealed box for running a program. gVisor catches a container’s system calls, its requests to the operating system, and answers them with its own kernel (the core of an operating system) written in Go. The goal is to lower the risk of a container breaking out to the host computer. Firecracker, built at Amazon Web Services for AWS Lambda, starts small virtual machines called microVMs in under 125 milliseconds. Each workload can then sit behind its own virtual machine barrier.
A coding agent runs real commands on your computer, and asking you about every one can make you less careful.
Follow four commands from one agent through the same boundary.
- 1 · setBefore work starts, you choose which folders commands may write, which paths they may not read and which domains they may reach.
- 2 · wrapEach shell command the agent runs starts inside the sandbox, and any program that command launches inherits the same limits.
- 3 · enforceThe operating system blocks file access outside the rules, and a proxy, a relay program outside the sandbox, lets through only allowed domains.
- 4 · reportAllowed work runs without a prompt, and a blocked command gets an error naming the path or host that was denied.
- 5 · escalateTo go past the boundary, the command must be retried outside the sandbox, which sends it through the normal permission prompt.
Permission rules judge a command before it runs; the sandbox limits what it can touch while it runs.
| Who | What they ask | What it works with |
|---|---|---|
| Developer using Claude Code | “Run the test suite and fix what fails.” | Test commands that may write only inside the project folder |
| Team using Codex cloud | “Update the dependencies in this repository.” | An isolated container that works offline after setup |
| Developer running a local MCP server | “Let the filesystem server write only to this folder.” | The server launched through the sandbox runtime with a write allowlist |
| Platform team running untrusted code | “Keep each customer's workload apart on shared machines.” | One lightweight virtual machine per workload |
- Anthropic reported that sandboxing cut permission prompts by 84% in its own internal use.
- With both layers on, a hijacked agent is held back from changing files outside its folders and from sending data to hosts not on the list.
- The limits cover every script, program and subprocess a command starts, not just the command itself.
- In Codex cloud, secrets set for an environment are removed before the agent starts working.
- Claude Code's sandbox lets commands read files such as ~/.ssh by default; they stay readable unless you add them to a deny list.
- Allowing a broad domain such as github.com can still open a path for leaking data.
- It covers shell commands only. Claude Code's built-in Read, Edit and Write tools use permission rules instead.
- Anthropic's own documentation says it reduces risk but is not a complete isolation boundary.
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsConfigure the sandboxed Bash tool, Anthropic · read 28 Sept 2026
- officialBeyond permission prompts: making Claude Code more secure and autonomous, Anthropic · read 28 Sept 2026
- repoanthropics/sandbox-runtime: A lightweight sandboxing tool for enforcing filesystem and network restrictions on arbitrary processes at the OS level, Anthropic · read 28 Sept 2026
- docsSandbox, OpenAI · read 28 Sept 2026
- docsAgent approvals & security, OpenAI · read 28 Sept 2026
- docsWhat is gVisor?, The gVisor Authors · read 28 Sept 2026
- officialFirecracker: secure and fast microVMs for serverless computing, Firecracker project, Amazon Web Services · read 28 Sept 2026