how to roll back an agent's bad infrastructure change
Rolls back infrastructure changes made by AI agents. Use when an agent's Terraform, config, or kubectl change broke something, when you need the pre-change state fast, or when designing agent change safety. Covers recovery before blame. Not for application code rollbacks.
TL;DR
Rolling back an agent's infrastructure change is a race against the agent's next action: first stop the agent from making more changes, then restore the last known good state from version control, then verify. The preconditions that make this fast are versioned infrastructure code, a known-good state to return to, and the ability to revoke the agent's access instantly. Build those before you need them.
The query
how to roll back an agent's bad infrastructure changeUse this when
- An agent's infrastructure change caused an incident
- You need the previous infrastructure state restored quickly
- Designing safety for agent-driven infra changes
- The agent is still running and might change more
Not for when
- Application code rollbacks (deploy pipeline, different topic)
- Data recovery (backups, different topic)
- Planned infrastructure migrations
Steps
Step 1: Stop the agent first
Revoke or pause the agent's credentials before touching infrastructure. An agent still running can fight your rollback by reapplying its change. This is step one, not step three. Expected output: the agent cannot make further changes; confirmed via audit log silence.
Step 2: Identify exactly what changed
Use the command logs and the infrastructure diff: what did the agent modify, in what order. Version-controlled infrastructure makes this a git diff; unversioned changes make it archaeology. Know the full change set before reverting any of it. Expected output: the complete list of changed resources.
Step 3: Restore the last known good state
Revert the infrastructure code to the last good commit and apply, or restore from the state snapshot. Prefer the declarative path (revert code, apply) over manual console fixes; manual fixes during an incident create the next incident. Expected output: infrastructure converging back to the known-good state.
Step 4: Verify the restoration worked
Check the actual resource state, not just the apply output: the load balancer routes correctly, the security groups match, the service is healthy. Agents' changes sometimes have side effects (deleted resources, rotated credentials) that a simple revert does not restore. Expected output: verified healthy state, with side effects explicitly checked.
Step 5: Do the incident review with the agent's logs
Run the postmortem with the agent's full command history: what was it trying to do, where did its plan go wrong, what guardrail was missing. The fix is usually a missing constraint (no plan approval, no canary), not a bad agent. Expected output: a systemic fix (new gate, new check) rather than just "be more careful".
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_DOlfh1wsoaH6CB-8jM-GiA
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.