agent's cleanup script terminated the wrong EC2 instances - tag filter "env: dev-" matched staging and prod
A playbook for making destructive cost-agent filters safe: ban prefix and substring tag matching on destructive actions, require exact tag matches plus an explicit exclusion list, always preview with a dry run, and gate deletions behind human approval. Use when a cleanup agent's tag filter matched staging and prod and terminated the wrong EC2 instances. Not for rightsizing math errors, stale billing data, or anomaly false positives.
TL;DR
Never let a destructive filter use prefix or substring matching. Require exact tag matches, add an explicit never-touch exclusion list (prod, staging, canary), preview every action with a dry run, and require a human to approve the target list before anything terminates. The wrong instances died because "dev-" is a prefix of nothing but a substring of everything the agent should not touch.
The query
agent's cleanup script terminated the wrong EC2 instances - tag filter "env: dev-" matched staging and prodUse this when
- A cleanup or rightsizing agent terminated, stopped, or deleted the wrong resources
- The filter used prefix, substring, or wildcard matching on tags or names
- Staging or prod resources share a naming pattern with dev resources
- The agent acted without a preview or approval step
Not for
- Rightsizing recommendations that are just wrong (math errors, bad baselines)
- Anomaly alerts firing on normal traffic
- The agent reading stale billing data
- API throttling during a scan (no resources harmed)
Steps
1. Freeze the automation and inventory the damage
Disable the cleanup job immediately. List every instance the script touched: instance IDs, tags, and whether they are recoverable (stopped vs terminated).
Expected output: a complete list of affected instances and the job disabled.
2. Replace fuzzy matching with exact matches
Change the filter from substring/prefix matching to exact tag-value equality. env equals dev is safe; "contains dev-" is not. If the tagging scheme is inconsistent, fix the tags first, then re-enable the job.
# wrong: matches anything containing the string
aws ec2 describe-instances --filters "Name=tag:env,Values=*dev-*"
# right: exact value match only
aws ec2 describe-instances --filters "Name=tag:env,Values=dev"Expected output: the filter returns only resources tagged exactly dev.
3. Add an explicit never-touch list
Maintain an exclusion list the filter can never override: any instance tagged env=prod or env=staging, instances with a do-not-delete tag, and named canary or DR resources. The destructive action must check this list after the filter, as a separate step.
Expected output: a second guard that blocks the action even if the filter misfires.
4. Require a dry-run preview
Every run must first output the exact target list (instance IDs + tags) without acting. Log the preview. The agent may proceed only after the preview step completes and the list is reviewed.
Expected output: a logged target list showing instance IDs and their tags, with zero actions taken.
5. Gate destructive actions behind human approval
Terminations and deletions require an explicit approval: the preview list goes to a human (or an approval queue) and the agent acts only on approval. Read-only scans can stay autonomous; anything that destroys cannot.
Expected output: no destructive run executes without a recorded approval.
Variant phrasings
cost agent stopped the wrong RDS instances by substring tag match
Same fix. Substring matching on DB tags is the same bug with a different resource type; the never-touch list in step 3 should cover prod replicas explicitly.
agent's shutdown schedule stopped the canary instances
Name-pattern matching ("test-") is the same fuzzy-match bug. Canaries belong on the never-touch list by name pattern AND tag.
agent deleted snapshots it thought were orphaned
Same family: the "orphan" heuristic was wrong (a silent tagging failure). The fix is the same shape: exact criteria, preview, approval.
Why it happens
The agent is optimizing for coverage: find everything that looks idle or dev-like. Prefix and substring matches feel thorough, and on a clean tagging scheme they work. But tagging schemes are never clean: staging borrows dev's prefix, prod canaries carry test names, and one inconsistent tag turns a cleanup into an outage. The agent has no model of "resources that must never die," so the filter is the only guard, and a fuzzy filter is no guard at all.
Edge cases
- Tagging is inconsistent across the fleet: fix tags before re-enabling automation. A cleanup agent running on messy tags will always be dangerous.
- Cross-account resources: the never-touch list must be evaluated per account. A filter that is safe in dev-account can be lethal if it ever runs against the prod account.
- Terminated instances cannot be recovered: if the damage is done, the playbook shifts to restore from AMI/snapshot and post-mortem. The prevention steps above are for the next run.
- Dry-run flags that lie: verify the dry run actually skips the API call (some agents' "dry run" only skips the final confirmation). Test it against a sandbox first.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_j2eyql-xEpYYZHyuWcODEQ
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.