how to audit what an agent accessed
A step-by-step skill for auditing agent activity: structured tool-call logging, access review cadence, anomaly alerts, and incident forensics. Use when a team needs to answer 'what did the agent touch', for periodic access reviews, or after a suspected agent misstep. Triggers: 'audit agent access', 'what did the agent do', 'agent activity log', 'agent forensics'. Not for: real-time blocking of agent actions, model behavior evaluation, or general SIEM tuning.
how to audit what an agent accessed
TL;DR
You cannot audit what you did not record, so log every tool call with its arguments, data touched, and timestamp, then review the logs on a cadence instead of only after incidents. Define what "normal" looks like per agent so anomalies stand out, and keep the trail long enough to matter. Auditing is a habit, not a postmortem activity.
how to audit what an agent accessedUse this when
- Someone asks what an agent did, read, or changed
- You run periodic access reviews for agents with tool grants
- An agent behaved unexpectedly and you need the timeline
- Compliance or a customer asks for the agent activity record
- You are setting up monitoring for a new agent deployment
Not for this skill when
- You need to block actions in real time (see sandboxing and grants skills)
- The question is about evaluating model quality or tone
- You are tuning a general SIEM with no agent context
- The agent has no tool access to audit
Steps
1. Confirm the log captures every tool call
The log must record each invocation: timestamp, agent identity, tool name, arguments, data accessed, and result summary. Sample a recent run and check for gaps.
tail -5 [HOME]/...Expected: each line is a complete structured record. Any tool call that can happen without a log line is an audit hole; fix the instrumentation before anything else.
2. Define normal per agent
For each agent, write down its expected pattern: which tools, what data scope, what hours, what volume. "Normal" for a nightly report agent differs wildly from an on-call responder.
Expected: a short baseline doc per agent. Without it, every review is just staring at logs hoping something looks wrong.
3. Review on a cadence, not just after incidents
Schedule reviews (weekly for high-privilege agents, monthly for the rest) that walk the log against the baseline. Rotate reviewers so familiarity does not become blindness.
Expected: review notes exist for the last cycle, naming what was checked and what was found, even when the answer is "nothing unusual".
4. Alert on the patterns that matter
Automate detection for the obvious anomalies: first-time access to a sensitive tool, volume spikes, off-hours activity for agents that should be idle, and grant-denied attempts (probing).
grep -c "denied" [HOME]/...Expected: the denied count is visible and triaged, not just logged. A spike in denials is someone (or something) testing boundaries.
5. Keep retention that matches the risk
High-privilege agents need months of history; the exact period follows your incident discovery time. Old logs must stay queryable, not just archived to a bucket nobody opens.
Expected: a documented retention policy per agent tier, and a successful test restore or query against older logs each quarter.
6. Practice the forensics drill
Once a quarter, pick a past incident or a synthetic scenario and reconstruct the full timeline from logs alone: what the agent accessed, in what order, driven by which input. Time the exercise.
Expected: the timeline is complete and the drill surfaces exactly which questions the logs cannot answer, which becomes next quarter's instrumentation work.
Variant: auditing multi-agent workflows
Attribute every action to the specific subagent that performed it and record the delegation chain. "The system did it" is not an audit answer; the trail must name each agent in the chain.
Variant: agents accessing customer data
Log the record identifiers or query predicates, not just "read the database". Data-access audits need to answer which customers were touched, which means the log must carry that granularity.
Variant: third-party tool calls
Calls to external services appear in your log and in the vendor's log; reconcile the two periodically. Discrepancies mean calls are happening outside your instrumentation.
Why this happens
Agent actions are fast, numerous, and machine-initiated, which makes them invisible in the ways human actions are not: nobody watches, nobody remembers, and the volume defeats casual review. Teams discover this during the first incident, when "what did it do" has no answer. Structured logging plus a review cadence converts that invisibility into a routine control.
Edge cases and pitfalls
- Log tampering: agents with write access to their own logs can edit history; ship logs to append-only storage the agent cannot modify.
- PII in the audit log creates a second sensitive store; mask values while keeping identifiers needed for the audit.
- Alert fatigue from noisy baselines leads to ignored alerts; tune thresholds per agent instead of globally.
- Clock skew across services scrambles timelines; sync time and log in UTC with timezone noted.
- Reviews that only check "was anything bad" miss slow drift; compare against the baseline doc, not against vibes.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_5qedpVhb3VI6nh5Jq7LvYg
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.