how to log every command an agent runs in production
Sets up complete command logging for AI agents operating in production. Use when agents execute shell commands or API calls against prod, when audit trails are required, or when debugging agent-caused incidents. Covers capture, storage, and review. Not for restricting what agents can do.
TL;DR
Log every command an agent runs with the full context: the exact command, its arguments, working directory, timestamp, exit code, and what the agent was trying to accomplish. Ship the logs somewhere the agent cannot modify, review them after incidents, and sample them routinely. An agent whose actions are invisible is an unauditable actor in your production environment.
The query
how to log every command an agent runs in productionUse this when
- AI agents execute commands against production
- Audit trails are required for compliance or incident review
- Debugging what an agent did during an incident
- Designing agent infrastructure
Not for when
- Restricting agent permissions (RBAC, different topic)
- Approving agent actions (human-in-the-loop, different topic)
- Logging human operator sessions (related but separate)
Steps
Step 1: Capture at the execution layer, not in the agent
Log commands where they execute (the shell wrapper, the API proxy, the tool runner), not by asking the agent to report them. Agent-reported logs are incomplete by construction; execution-layer capture is complete by construction. Expected output: a log stream that records commands even when the agent errors or misbehaves.
Step 2: Record the full context per command
Each entry needs: timestamp, agent identity, exact command and arguments, working directory or target resource, exit code, and duration. The "why" (the agent's stated intent) is valuable too if available. Skimpy logs ("ran kubectl") are useless in a postmortem. Expected output: log entries detailed enough to replay the agent's session.
Step 3: Ship logs somewhere the agent cannot alter
Send command logs to append-only storage or a separate logging account the agent has no write access to. An agent that can edit its own audit trail makes the trail worthless. This is the non-negotiable property. Expected output: tamper-evident command history, verifiable independently of the agent.
Step 4: Alert on dangerous patterns in real time
Watch the command stream for destructive patterns (recursive deletes, wide RBAC changes, mass terminations) and alert or block. Logging without monitoring is archaeology; pair the trail with detection. Expected output: dangerous commands trigger alerts within seconds, not discovered in tomorrow's review.
Step 5: Review routinely, not just after incidents
Sample agent command logs weekly: are the commands sensible, are there near-misses, is the agent's behavior drifting. Routine review catches the slow degradation that incident-only review misses. Expected output: a weekly review habit with findings that improve the agent's constraints.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_zdrRz4T-kXkNP3eWPjMbhg
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.