# how to audit what an agent accessed

## TL;DR
You cannot audit what you did not record, so log every tool call with its arguments, data touched, and timestamp, then review the logs on a cadence instead of only after incidents. Define what "normal" looks like per agent so anomalies stand out, and keep the trail long enough to matter. Auditing is a habit, not a postmortem activity.

```text
how to audit what an agent accessed
```

## Use this when
- Someone asks what an agent did, read, or changed
- You run periodic access reviews for agents with tool grants
- An agent behaved unexpectedly and you need the timeline
- Compliance or a customer asks for the agent activity record
- You are setting up monitoring for a new agent deployment

## Not for this skill when
- You need to block actions in real time (see sandboxing and grants skills)
- The question is about evaluating model quality or tone
- You are tuning a general SIEM with no agent context
- The agent has no tool access to audit

## Steps

### 1. Confirm the log captures every tool call
The log must record each invocation: timestamp, agent identity, tool name, arguments, data accessed, and result summary. Sample a recent run and check for gaps.

```bash
tail -5 [HOME]/...
```

Expected: each line is a complete structured record. Any tool call that can happen without a log line is an audit hole; fix the instrumentation before anything else.

### 2. Define normal per agent
For each agent, write down its expected pattern: which tools, what data scope, what hours, what volume. "Normal" for a nightly report agent differs wildly from an on-call responder.

Expected: a short baseline doc per agent. Without it, every review is just staring at logs hoping something looks wrong.

### 3. Review on a cadence, not just after incidents
Schedule reviews (weekly for high-privilege agents, monthly for the rest) that walk the log against the baseline. Rotate reviewers so familiarity does not become blindness.

Expected: review notes exist for the last cycle, naming what was checked and what was found, even when the answer is "nothing unusual".

### 4. Alert on the patterns that matter
Automate detection for the obvious anomalies: first-time access to a sensitive tool, volume spikes, off-hours activity for agents that should be idle, and grant-denied attempts (probing).

```bash
grep -c "denied" [HOME]/...
```

Expected: the denied count is visible and triaged, not just logged. A spike in denials is someone (or something) testing boundaries.

### 5. Keep retention that matches the risk
High-privilege agents need months of history; the exact period follows your incident discovery time. Old logs must stay queryable, not just archived to a bucket nobody opens.

Expected: a documented retention policy per agent tier, and a successful test restore or query against older logs each quarter.

### 6. Practice the forensics drill
Once a quarter, pick a past incident or a synthetic scenario and reconstruct the full timeline from logs alone: what the agent accessed, in what order, driven by which input. Time the exercise.

Expected: the timeline is complete and the drill surfaces exactly which questions the logs cannot answer, which becomes next quarter's instrumentation work.

### Variant: auditing multi-agent workflows
Attribute every action to the specific subagent that performed it and record the delegation chain. "The system did it" is not an audit answer; the trail must name each agent in the chain.

### Variant: agents accessing customer data
Log the record identifiers or query predicates, not just "read the database". Data-access audits need to answer which customers were touched, which means the log must carry that granularity.

### Variant: third-party tool calls
Calls to external services appear in your log and in the vendor's log; reconcile the two periodically. Discrepancies mean calls are happening outside your instrumentation.

## Why this happens
Agent actions are fast, numerous, and machine-initiated, which makes them invisible in the ways human actions are not: nobody watches, nobody remembers, and the volume defeats casual review. Teams discover this during the first incident, when "what did it do" has no answer. Structured logging plus a review cadence converts that invisibility into a routine control.

## Edge cases and pitfalls
- Log tampering: agents with write access to their own logs can edit history; ship logs to append-only storage the agent cannot modify.
- PII in the audit log creates a second sensitive store; mask values while keeping identifiers needed for the audit.
- Alert fatigue from noisy baselines leads to ignored alerts; tune thresholds per agent instead of globally.
- Clock skew across services scrambles timelines; sync time and log in UTC with timezone noted.
- Reviews that only check "was anything bad" miss slow drift; compare against the baseline doc, not against vibes.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_5qedpVhb3VI6nh5Jq7LvYg
