# how to detect a compromised agent

## TL;DR
Watch what the agent does, not what it says. A compromised agent chats normally while its tool calls go sideways: new hosts, odd hours, bulk data reads, credential use you never authorized. Baseline normal behavior first, then alert on deviations and isolate on the first strong signal.

```text
how to detect a compromised agent
```

## Use this when
- An agent starts behaving oddly: strange tool calls, unexpected destinations
- You suspect prompt injection succeeded and want to confirm the blast radius
- You are building monitoring for a fleet of agents
- A vendor or customer asks how you would know if an agent got hijacked

## Not for this skill when
- The agent is just bad at its task (incompetence is not compromise)
- You already confirmed compromise and need cleanup (that is incident response)
- You are trying to prevent compromise in the first place (harden first, detect second)

## Steps

**1. Baseline what normal looks like.**

For each agent, record its usual hosts, tools, data sources, run hours, and action volume over a couple of weeks. You cannot spot abnormal until you have defined normal. Store the baseline where alerts can reference it.

Expected: a short profile per agent: typical hosts, typical tools, typical hours, typical daily action count.

**2. Alert on new outbound hosts.**

The first thing a hijacked agent usually does is phone home somewhere new. Any fetch or connection to a host the agent has never used before deserves an alert, especially in the hour after the agent read external content.

```
grep "fetch_verdict=deny" /var/log/agent/fetches.log | tail -20
```

Expected: a list of blocked or first-seen hosts with timestamps, ready to triage.

**3. Alert on bulk data reads and unusual data access.**

An agent that normally reads ten records pulling ten thousand is a signal. So is an agent touching a restricted datastore it has never touched. Set thresholds per data source and alert when they break.

Expected: alerts fire on volume spikes and first-time access to sensitive sources, with the session ID attached.

**4. Alert on credential and secrets access outside the task.**

Watch for the agent reading secrets, using credentials at odd hours, or using a credential for a system outside its current task scope. Credentials are the crown jewels; their use should be boring and predictable.

Expected: every credential use is logged with task context, and out-of-context use pages someone.

**5. Alert on changes to the agent itself.**

If the agent's instructions get edited, a new scheduled job appears, or its tool permissions widen, treat it as a compromise signal until proven otherwise. Persistence is the attacker's goal.

Expected: any modification to agent configuration generates an alert with a before and after diff.

**6. Isolate on the first strong signal, then investigate.**

Do not wait for three corroborating alerts. Halt the session, revoke its credentials, and then dig through the transcript. You can always resume a falsely accused agent; you cannot un-send exfiltrated data.

```
agent sessions halt --session [SESSION_ID] --reason "anomaly: first-seen exfil host"
```

Expected: the session stops within seconds and the investigation starts from a frozen state.

### Variant: signs an AI agent is hijacked
New outbound hosts, bulk reads, odd-hour activity, tool calls that do not match the task, and chat responses that dodge questions about what it just did. Any two together mean isolate first.

### Variant: agent behaving strangely how to check
Pull the session transcript with full tool calls and read it forward from the last external input the agent consumed. Strange behavior almost always starts right after a fetch, an upload, or a pasted blob of text.

### Variant: prompt injection detection in logs
Look for the pattern: external content in, then behavior change within minutes. Correlate fetch logs with tool-call logs on the session ID and the timeline tells the story.

### Variant: audit agent actions after a suspected breach
Export the full tool-call history, list every irreversible action, and assume the log is incomplete. Rotate everything the session could reach regardless of what the log shows.

## Why this happens
Compromise rarely announces itself in chat. The model keeps producing fluent, plausible responses while following injected instructions in its tool use. Teams that only review chat output miss it entirely; teams that monitor tool calls catch it in minutes. The chat is the mask, the tool log is the face.

## Edge cases and pitfalls
- **The baseline is noisy because the agent does varied work:** baseline per task type, not per agent. A research agent and a deploy agent have nothing in common.
- **Alert fatigue from first-seen hosts:** tune with an allowlist of known-good hosts so only truly new ones alert, and auto-resolve when the host gets approved.
- **The attacker deletes log entries:** ship logs off the agent host in real time to append-only storage. Local logs are evidence; remote logs are proof.
- **Slow, low-volume exfiltration under thresholds:** pair threshold alerts with impossible-travel style rules: credential use from a new network, or data access at 3am for a 9-to-5 agent.
- **You have no logging at all:** start with tool-call logging today. Detection without logs is guessing, and this whole skill assumes the logs exist.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst__WWFeSRKppcnoyAFAB-B1w
