## TL;DR
Never let a destructive filter use prefix or substring matching. Require exact tag matches, add an explicit never-touch exclusion list (prod, staging, canary), preview every action with a dry run, and require a human to approve the target list before anything terminates. The wrong instances died because "dev-" is a prefix of nothing but a substring of everything the agent should not touch.

## The query

```text
agent's cleanup script terminated the wrong EC2 instances - tag filter "env: dev-" matched staging and prod
```

## Use this when

- A cleanup or rightsizing agent terminated, stopped, or deleted the wrong resources
- The filter used prefix, substring, or wildcard matching on tags or names
- Staging or prod resources share a naming pattern with dev resources
- The agent acted without a preview or approval step

## Not for

- Rightsizing recommendations that are just wrong (math errors, bad baselines)
- Anomaly alerts firing on normal traffic
- The agent reading stale billing data
- API throttling during a scan (no resources harmed)

## Steps

### 1. Freeze the automation and inventory the damage

Disable the cleanup job immediately. List every instance the script touched: instance IDs, tags, and whether they are recoverable (stopped vs terminated).

Expected output: a complete list of affected instances and the job disabled.

### 2. Replace fuzzy matching with exact matches

Change the filter from substring/prefix matching to exact tag-value equality. `env` equals `dev` is safe; "contains dev-" is not. If the tagging scheme is inconsistent, fix the tags first, then re-enable the job.

```bash
# wrong: matches anything containing the string
aws ec2 describe-instances --filters "Name=tag:env,Values=*dev-*"
# right: exact value match only
aws ec2 describe-instances --filters "Name=tag:env,Values=dev"
```

Expected output: the filter returns only resources tagged exactly `dev`.

### 3. Add an explicit never-touch list

Maintain an exclusion list the filter can never override: any instance tagged `env=prod` or `env=staging`, instances with a `do-not-delete` tag, and named canary or DR resources. The destructive action must check this list after the filter, as a separate step.

Expected output: a second guard that blocks the action even if the filter misfires.

### 4. Require a dry-run preview

Every run must first output the exact target list (instance IDs + tags) without acting. Log the preview. The agent may proceed only after the preview step completes and the list is reviewed.

Expected output: a logged target list showing instance IDs and their tags, with zero actions taken.

### 5. Gate destructive actions behind human approval

Terminations and deletions require an explicit approval: the preview list goes to a human (or an approval queue) and the agent acts only on approval. Read-only scans can stay autonomous; anything that destroys cannot.

Expected output: no destructive run executes without a recorded approval.

## Variant phrasings

### cost agent stopped the wrong RDS instances by substring tag match

Same fix. Substring matching on DB tags is the same bug with a different resource type; the never-touch list in step 3 should cover prod replicas explicitly.

### agent's shutdown schedule stopped the canary instances

Name-pattern matching ("test-") is the same fuzzy-match bug. Canaries belong on the never-touch list by name pattern AND tag.

### agent deleted snapshots it thought were orphaned

Same family: the "orphan" heuristic was wrong (a silent tagging failure). The fix is the same shape: exact criteria, preview, approval.

## Why it happens

The agent is optimizing for coverage: find everything that looks idle or dev-like. Prefix and substring matches feel thorough, and on a clean tagging scheme they work. But tagging schemes are never clean: staging borrows dev's prefix, prod canaries carry test names, and one inconsistent tag turns a cleanup into an outage. The agent has no model of "resources that must never die," so the filter is the only guard, and a fuzzy filter is no guard at all.

## Edge cases

- Tagging is inconsistent across the fleet: fix tags before re-enabling automation. A cleanup agent running on messy tags will always be dangerous.
- Cross-account resources: the never-touch list must be evaluated per account. A filter that is safe in dev-account can be lethal if it ever runs against the prod account.
- Terminated instances cannot be recovered: if the damage is done, the playbook shifts to restore from AMI/snapshot and post-mortem. The prevention steps above are for the next run.
- Dry-run flags that lie: verify the dry run actually skips the API call (some agents' "dry run" only skips the final confirmation). Test it against a sandbox first.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_j2eyql-xEpYYZHyuWcODEQ
