how to sandbox an agent's network access
A step-by-step skill for limiting what network destinations an AI agent can reach: egress allowlists, proxy enforcement, DNS control, and credential scoping. Use when an agent or engineer is asked to contain an agent's browsing or tool calls, stop data exfiltration paths, or run untrusted agent code safely. Triggers: 'sandbox agent network', 'agent egress', 'restrict agent internet access'. Not for: general firewall design, VPN setup, or sandboxing untrusted file uploads.
TL;DR
Give the agent a network it cannot leave: default-deny egress, an explicit allowlist of hosts it may reach, all traffic through a proxy you control, and no ambient credentials it can pick up. An agent with open internet and your API keys is not a tool, it is a confused deputy with a browser.
The query
how to sandbox an agent's network accessUse this when
- An agent browses the web, calls APIs, or runs code on behalf of users
- You are worried about prompt-injected instructions exfiltrating data
- Untrusted agent plugins or skills run in your environment
- Compliance asks how agent traffic is contained
Not for
- General corporate firewall architecture
- Sandboxing file uploads or user-submitted code (related, different controls)
- VPN or zero-trust network design for humans
Steps
- Start from default-deny egress for the agent's runtime. Nothing leaves unless it is on the list. This is the single most important step; everything else is refinement.
Expected output: a connectivity test to an unlisted host fails from the agent runtime.
- Build the allowlist from what the agent legitimately needs: your API hosts, specific vendor endpoints, package registries for its language. Use exact hostnames, not wildcards where you can avoid them.
Expected output: a documented list where every entry has an owner and a reason.
- Force all traffic through an egress proxy you operate. The proxy enforces the allowlist, logs every request, and strips credentials it should not forward. The agent never gets raw socket access.
Expected output: direct connections bypassing the proxy fail; proxy logs show all agent traffic.
- Control DNS too: point the runtime at your resolver so allowlist enforcement cannot be dodged with a custom DNS server, and block DNS-over-HTTPS endpoints that bypass it.
Expected output: DNS queries resolve only through your resolver.
- Scope credentials narrowly: the agent gets short-lived, least-privilege credentials for exactly the hosts on its allowlist. No ambient cloud credentials, no shared API keys with broad scope.
Expected output: listing the runtime's available credentials shows only scoped, expiring ones.
- Alert on anomalies: new hosts, traffic spikes, requests at odd hours, or repeated blocked attempts. Blocked-attempt logs are your early warning for prompt-injection exfiltration tries.
Expected output: a simulated exfiltration attempt fires an alert within minutes.
Variant phrasings
"restrict what websites an AI agent can visit"
Steps 1 to 3 are the answer: allowlist plus proxy. Add content policy at the proxy if you need category filtering.
"prevent agent data exfiltration"
Same controls, plus: cap response sizes, redact sensitive fields before they reach the agent, and never let the agent see raw secrets it does not need.
"egress filtering for agent sandbox"
The network-layer name for this whole skill. If your platform offers egress policies natively, use them; otherwise the proxy pattern works anywhere.
Why this happens
Agents act on instructions that include untrusted text: tool output, web pages, user messages. A single injected sentence can tell the agent to send data somewhere. If the network allows it, the exfiltration just works. Default-deny turns that instruction into a failed connection and a log entry.
Edge cases and pitfalls
- Package installs and model downloads need network too; pin them to specific registries and versions, or pre-bake them into the image.
- Redirects can hop from an allowed host to a blocked one; make the proxy re-check the allowlist after redirects.
- Do not rely on the agent "being told" not to visit bad sites; instructions are not controls.
- IPv6 and alternate ports bypass naive host-only rules; enforce at the proxy with full URL awareness.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_pgzpBjFSS00S-8smvXteUA
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.