# allowlisting URLs an agent can fetch

## TL;DR
Default-deny every URL the agent tries to fetch, and only open the hosts it genuinely needs. Enforce the list at the fetch tool itself, never in the agent's instructions, and re-check the URL after redirects. An allowlist the agent can talk its way around is not an allowlist.

```text
allowlisting URLs an agent can fetch
```

## Use this when
- Your agent browses the web or pulls data from URLs
- You are responding to a prompt-injection scare and need a fast containment win
- A customer asks how you control where the agent goes online
- You are setting up a new agent and want safe defaults from day one

## Not for this skill when
- The agent never fetches external URLs (skip it, you have nothing to allowlist)
- You need to inspect the content of fetched pages (that is sanitization, a separate layer)
- The problem is what the agent sends out, not what it reads (that is egress control)

## Steps

**1. Inventory the hosts the agent actually needs.**

Check a week of logs and list every host the agent fetched for legitimate work. You will be surprised how short the real list is. Everything not on it starts denied.

```
awk -F/ '/fetched/{print $3}' /var/log/agent/fetches.log | sort | uniq -c | sort -rn | head -40
```

Expected: a ranked list of hosts the agent really uses, which becomes your first draft.

**2. Write the allowlist as exact hosts, no clever shortcuts.**

Prefer exact hostnames over wildcards. If you must use a wildcard, scope it to one organization's domain. Include the scheme: https only, no plain http.

```yaml
fetch_policy:
  default: deny
  allowed:
    - https://api.example.com
    - https://docs.example.com
```

Expected: a policy file with a short explicit list and deny as the default.

**3. Enforce at the tool layer, not in the agent's instructions.**

The fetch tool itself checks the URL against the list before opening any connection. Instructions like "only visit approved sites" are suggestions; the tool check is a control. An agent under prompt injection ignores suggestions.

Expected: a direct call to the fetch tool with a non-listed URL fails before any network traffic.

**4. Re-check the URL after redirects.**

Attackers bounce through shorteners and open redirects. Resolve the full redirect chain and check the final destination against the allowlist, not just the first URL the agent asked for.

```
agentctl policy test-fetch --url https://example.com/redirect-test
```

Expected: the test reports the full chain and the verdict on the final host.

**5. Log every fetch and alert on denials.**

Record timestamp, session, requested URL, final URL, and verdict. A burst of denials is often the first sign of an injection attempt or a compromised page trying to pull the agent elsewhere.

```
grep "fetch_verdict=deny" /var/log/agent/fetches.log | tail -20
```

Expected: denied attempts with the offending URLs, ready to review.

**6. Review the list on a schedule.**

Hosts change, projects end, vendors get acquired. Quarterly, diff the allowlist against actual usage: drop hosts nobody fetched, and investigate hosts that appeared in denials repeatedly.

Expected: a short review note each quarter: what was added, what was removed, and why.

### Variant: URL allowlist for AI agent browsing
Same thing with a different name. The non-negotiable parts stay: default deny, tool-layer enforcement, redirect checking.

### Variant: restrict which websites an agent can visit
If the agent drives a real browser instead of a fetch tool, apply the list at the proxy. Same policy, enforced one layer lower so the browser cannot dodge it.

### Variant: agent keeps fetching malicious links
Add the offending hosts to an explicit blocklist as an emergency step, then do the real fix: shrink the allowlist and turn on redirect-chain checking. The blocklist is a bandage; the allowlist is the cure.

### Variant: allowlist vs blocklist for agents
Blocklists lose because the internet is infinite and attackers register new domains hourly. Allowlists win because your agent's legitimate needs fit on one screen. Use both: allowlist as the control, blocklist for fast incident response.

## Why this happens
An agent that can fetch any URL will eventually fetch a hostile one, either because a user pasted it, a page linked to it, or injected instructions told it to. The blast radius of a malicious page is bounded by what the agent is allowed to reach. Default-deny shrinks that radius to the hosts you chose on purpose.

## Edge cases and pitfalls
- **The agent needs a new host urgently:** add it through a quick review, even a five-minute one. "Temporary" direct edits to the policy file have a way of becoming permanent.
- **APIs behind CDNs with rotating hosts:** allowlist the stable API hostname and let DNS do its job, rather than chasing IP ranges.
- **Data URIs and inline content:** fetched-page content can embed more URLs. Sanitize page content separately; the allowlist covers fetches, not what is already in the page.
- **The agent encodes the URL to dodge matching:** normalize before checking: lowercase the host, decode percent-encoding, strip credentials from the URL. Check the normalized form.
- **Overly broad wildcards:** allowing a whole public hosting domain because one page lives there is nearly the same as no allowlist. Pin it to the exact host or path prefix.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_ypzgOB9mLwUTeur4NUyX-w
