how to sandbox an agent with tool access
A step-by-step skill for sandboxing agents that can call tools: inventorying capabilities, allowlisting tools and paths, read-only filesystems, egress control, and audit logging. Use when an agent deployment needs containment, a team grants an agent production tool access, or reviewers ask 'what can this agent touch'. Triggers: 'sandbox an agent', 'agent tool permissions', 'contain AI agent', 'agent safety boundaries'. Not for: model safety tuning, prompt design, or general container hardening.
how to sandbox an agent with tool access
TL;DR
An agent with tool access is a privileged user that acts fast and never sleeps, so sandbox it like one: list every tool it can call, allowlist only what the task needs, mount filesystems read-only by default, block network egress except to named destinations, and log every action. The sandbox is defined by what is denied, not by what is allowed.
how to sandbox an agent with tool accessUse this when
- You are deploying an agent with shell, file, or API tool access
- A team wants an agent to work in production or near-production data
- Reviewers ask for the blast radius of an agent deployment
- An agent needs third-party tools you do not fully trust
- You are writing the runbook for agent operations
Not for this skill when
- The question is about model behavior or prompt design
- You are hardening containers in general with no agent involved
- The agent only chats and calls no tools
- You need formal compliance certification (separate process)
Steps
1. Inventory every tool the agent can reach
List the full tool surface: shell commands, file operations, network calls, API clients, code execution. Include indirect reach, like a shell tool that can curl anywhere or a file tool that can read secrets.
grep -rn "tools\s*=\|available_tools\|tool_choice" [HOME]/... | head -20Expected: a written inventory with no "and whatever else it finds". Unknown tools are the ones that cause incidents.
2. Allowlist tools per task, default deny
Grant the minimum tool set each task needs and nothing more. A research task gets search and read; it does not get shell write or deploy. Review the grants whenever the task changes.
Expected: a grant matrix mapping tasks to tools, stored with the agent config. If you cannot explain why a tool is granted, remove it.
3. Make the filesystem read-only except for scratch
Mount the working tree read-only and give the agent one writable scratch directory. Secrets, configs, and credentials live outside both.
Expected: a write attempt outside scratch fails loudly in testing. Verify by having the agent attempt it in a dry run before granting real access.
4. Control network egress
Default-deny outbound network access and allowlist only the destinations the task needs: the API it queries, the repo host, nothing else. Log every connection attempt, allowed or denied.
Expected: an egress test shows only allowlisted hosts reachable, and the denied attempts appear in the log. An agent that can reach the open internet can exfiltrate.
5. Bound execution time and side effects
Set timeouts on tool calls and on the whole run. Require explicit confirmation for irreversible actions: deletes, publishes, payments, messages sent externally. The agent proposes; a human or a policy approves.
Expected: a runaway loop terminates on its own, and the confirmation gate has a test proving it triggers on a destructive action.
6. Log every tool call with its arguments
Record what was called, with what arguments, when, and what came back. This log is your audit trail and your incident timeline; without it you cannot answer "what did the agent do".
ls -la [HOME]/... | tail -5Expected: complete, timestamped, tamper-evident logs. If a tool call can happen without appearing here, the sandbox has a hole.
Variant: coding agents with repo write access
Scope writes to specific branches or directories, require PRs instead of direct pushes, and run the test suite on agent-produced changes before merge. The repo's own review process is part of the sandbox.
Variant: agents with customer data access
Add data minimization: the agent sees only the records the task needs, PII is masked in tool output, and every data access is logged against a ticket or task ID. Assume the logs will be audited.
Variant: agents calling other agents
Delegation multiplies the tool surface. Each subagent gets its own narrower grant, and the parent cannot hand down tools it was never granted. Log the delegation chain so responsibility is traceable.
Why this happens
Tool access turns language output into real-world effects, but teams grant it with the casualness of a chat permission because the agent "seemed safe" in demos. Demos do not probe boundaries; production does. Sandboxing works because it replaces trust in the agent's judgment with enforcement at the boundary, which holds even when the judgment fails.
Edge cases and pitfalls
- Tool descriptions are attacker-influenced when tools come from third parties; treat tool metadata as untrusted content.
- Time-of-check gaps: validating an action and executing it later lets state change in between; validate and execute atomically where it matters.
- Log volume from chatty agents can hide the important calls; log at the tool boundary with structured fields, not free text.
- Scratch directories accumulate state between runs; clean them or scope them per run to avoid cross-task contamination.
- Human approvers rubber-stamp; rotate approvers and sample approved actions for real review.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_CV8tRkR3hA3mKLsHSqRobw