sandboxing code execution for agents safely
How to run agent-generated code in throwaway containers with no network, read-only filesystems, resource caps, and non-root users. Use whenever an agent writes and runs code, installs packages, or executes anything influenced by untrusted input. Triggers: 'sandbox agent code', 'run untrusted agent code safely', 'docker sandbox AI agents'. Not for: agents that only suggest code for humans to run, or already-isolated serverless execution.
sandboxing code execution for agents safely
TL;DR
Run every line of agent-generated code in a throwaway container with no network, a read-only filesystem, and hard resource caps. The sandbox assumes the code is hostile until it proves otherwise. If the code misbehaves, you delete the container and move on.
sandboxing code execution for agents safelyUse this when
- An agent writes and runs code: scripts, notebooks, migrations, one-off fixes
- You let agents install packages or fetch dependencies
- Untrusted input (web pages, user uploads) flows into code the agent runs
- You are setting up a code-execution tool for an agent for the first time
Not for this skill when
- The agent only suggests code for a human to run (the human is the sandbox)
- Code runs in a fully managed serverless platform that already isolates it
- You need to sandbox the agent's file reads rather than code execution
Steps
1. Run code in throwaway containers, never on the host.
Each execution gets a fresh container that is destroyed afterward. Nothing persists between runs except what you explicitly copy out. Host execution is how a bad script becomes a bad day.
docker run --rm --network none --read-only --tmpfs /tmp:rw,noexec --memory 512m --cpus 1 --user 1000:1000 sandbox-image python /workspace/script.pyExpected: the script runs, prints its output, and the container is gone when it exits.
2. Default to no network.
Most agent scripts do not need the internet. Start with networking disabled and only enable it for specific runs that need package installs or API calls, ideally through an allowlisted proxy.
Expected: a script that tries to open a socket fails immediately with a network error.
3. Mount the filesystem read-only except one scratch directory.
The agent's code can read inputs and write to a single scratch dir. It cannot modify the image, the tooling, or anything outside its workspace. Mount inputs read-only too.
Expected: attempts to write outside the scratch dir fail with a read-only filesystem error.
4. Cap CPU, memory, and wall-clock time.
A runaway loop or a memory-hungry script should die from the cap, not from taking down the host. Set a timeout that kills the container: minutes, not hours.
docker run --rm --network none --memory 512m --cpus 1 --stop-timeout 30 sandbox-image python /workspace/script.pyExpected: a script that loops forever is killed at the timeout, and the host stays healthy.
5. Drop privileges: never run as root.
Run the container as an unprivileged user and drop Linux capabilities you do not need. Agent code does not need to mount disks or change the clock.
Expected: whoami inside the container prints a non-root user, and privileged syscalls fail.
6. Log the code, the inputs, and the result.
Store what ran, with what inputs, and what it printed, tied to the agent session. When a run does something odd next month, this is your replay.
Expected: every execution has a retrievable record: script hash, input hash, exit code, output.
7. Destroy, do not reuse.
Even with all the above, treat each sandbox as single-use. Reused sandboxes accumulate state, and state is where persistence hides.
Expected: no container outlives its run; a fresh image starts every execution.
Variant: run untrusted agent code safely
Same recipe. "Untrusted" just means you skip straight to the strictest settings: no network, read-only, short timeout.
Variant: docker sandbox for AI agents
Docker with the flags above is the common starting point. For stronger isolation, look at gVisor or Firecracker-based sandboxes, which add a second boundary between the container and the host.
Variant: gVisor vs docker for agents
gVisor intercepts syscalls in userspace, which shrinks the kernel attack surface a lot. Use it when agents run code from untrusted sources; plain Docker with tight flags is fine for code your own agent wrote from trusted inputs.
Variant: limit what agent-generated code can do
That is this whole skill: no network, read-only fs, resource caps, non-root, single-use. The limits are the feature.
Why this happens
Agent-written code is exactly as trustworthy as the agent's judgment in that moment, which includes moments of confusion, prompt injection, and plain bugs. rm -rf with the wrong variable is a typo until it runs on your host. The sandbox converts "the agent made a mistake" from an incident into a log line.
Edge cases and pitfalls
- The script genuinely needs the network: allow it per-run through an egress proxy with an allowlist, and log the destinations. Do not flip the default to network-on.
- Package installs need to write somewhere: give installs their own layer or a writable cache dir inside the scratch space, rebuilt from a lockfile each run.
- Secrets the script needs: inject them as environment variables at run time from your secret store, never bake them into the image or the script.
- The agent needs to debug interactively: give it a sandbox shell with the same restrictions, not host access. Debugging convenience never justifies host shells.
- Performance overhead worries: container startup is milliseconds to seconds. If that is too slow, you are running code far more often than you should be; batch it.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_pPvLi3r8irAE6tq0vP9WAQ
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.