how to audit what an agent changed in production
Audits agent changes in production by attributing actions to the agent's identity and collecting evidence from Kubernetes audit logs, cloud trail events, GitOps history, and Terraform state. Use to scope blast radius after a scare, answer auditor questions, or build the accountability story for a safety review. Not for agents without write access.
TL;DR
Attribute every agent change to the agent's identity, then collect the evidence from the layers that already log everything: Kubernetes audit logs, cloud trail events, GitOps commit history, and Terraform state. If you cannot answer "what did the agent touch in the last hour" in under five minutes, your audit setup is incomplete.
Error / query
how to audit what an agent changed in productionUse this skill when
- something broke and an agent was active around the same time
- auditors ask for a record of automated production changes
- you are building the agent's accountability story for a safety review
- you need to roll back an agent's change and want the full blast radius
Not for this skill when
- the agent never had write access (there is nothing to audit)
- you are setting up the guardrails themselves (different skill)
- the question is about prompt or conversation logs rather than infrastructure changes
Steps
Step 1: Give every agent action a stable identity to search for
kubectl get events -n prod --sort-by=.lastTimestamp | grep -i "sre-agent" | tail -20Expected: events attributed to the agent's ServiceAccount. All auditing starts here: if agent actions are not attributable to a distinct identity, stop and fix the identity setup first.
Step 2: Query the Kubernetes audit log for the agent's writes
grep "system:serviceaccount:ops:sre-agent" /var/log/kubernetes/audit/audit.log | grep -E "\"verb\":\"(create|update|patch|delete)\"" | tail -20Expected: one line per mutating API call with the object, namespace, and timestamp. This is the ground truth for what the agent did inside the cluster; events alone can miss direct API calls.
Step 3: Check the cloud trail for out-of-cluster changes
aws cloudtrail lookup-events --lookup-attributes AttributeKey=Username,AttributeValue=sre-agent --max-results 20Expected: the agent's assumed-role sessions and API calls. Agents that touch both Kubernetes and cloud APIs leave trails in both places; check both or you only see half the story.
Step 4: Diff the GitOps repo and Terraform state around the incident window
git -C /opt/gitops log --since="3 hours ago" --author="sre-agent" --oneline
terraform show -json | jq ".values.root_module.resources | length"Expected: the commit list shows exactly which manifests the agent changed; the state resource count flags unexpected additions. Git history gives you intent and review trail; state gives you actuality.
Step 5: Build the timeline and confirm completeness
kubectl -n prod rollout history deployment/[APP] | tail -5Expected: revision history lining up with the audit timestamps. Walk the timeline: agent action, observed effect, human response. If there is a gap where something changed with no attributed actor, your logging has a hole; close it before the next incident.
Variant phrasings
"which changes did the ai agent make last night"
Steps 1 through 3 scoped to the overnight window. Start with the audit log, not with asking the agent; its summary of its own actions is not evidence.
"prove the agent did not cause the outage"
Same evidence, used defensively: an empty audit result for the agent's identity across the window is the proof, plus the timeline showing the real cause.
"agent change history for compliance"
Export steps 2-4 on a schedule into immutable storage; auditors want the raw records, not a dashboard screenshot.
Why it happens
Agent incidents are hard to audit because the action, the reasoning, and the identity are split across systems: the agent's log has the reasoning, the cluster has the action, and only the ServiceAccount ties them together. Teams that skip the identity step end up with "some automation did something" and no way to scope the blast radius, which turns a ten-minute rollback into a two-hour archaeology dig.
Edge cases and pitfalls
- Audit log volume is huge; make sure the agent's identity is indexed or filtered at collection time, or queries time out when you need them most.
- Impersonation can make a human's action look like the agent's; log the impersonation chain, not just the final identity.
- Deleted namespaces take their events with them; the cluster audit log and cloud trail are the durable copies.
- Agents that shell out to scripts can hide actions inside one exec; require the agent to log intended commands before running them.
- Retention windows matter: if audit logs age out after 30 days, a slow-burn agent mistake discovered on day 45 is unauditable.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_jhfpuBCDDVf6tZLK4L3gIA
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.