## TL;DR
Databases are the worst thing to stop by accident, and substring tag matching is how it happens. Match DB tags by exact value, describe every matched instance and confirm its role (primary vs replica, prod vs dev) before acting, preview the stop list with a dry run, and require human approval for any database stop. A stopped prod replica is an outage with a billing excuse.

## The query

```text
cost agent stopped the wrong RDS instances - it matched DB tags by substring and killed the prod replica
```

## Use this when

- A cost agent stopped or deleted an RDS instance it should not have touched
- The match used substring, prefix, or wildcard logic on DB tags or identifiers
- Prod and non-prod databases share naming or tagging patterns
- The agent acted on databases without a preview or approval

## Not for

- EC2 instance cleanup with bad filters (same shape, different playbook)
- Rightsizing recommendations with wrong math
- RDS performance problems or replica lag
- Billing data being stale or wrong

## Steps

### 1. Check what is stopped and restart what should be running

List the affected DB instances, their roles (primary, read replica), and their current status. Restart anything prod that should be up. Confirm replication caught up before calling it recovered.

```bash
aws rds describe-db-instances --query "DBInstances[].[DBInstanceIdentifier,DBInstanceStatus,ReadReplicaSourceDBInstanceIdentifier]" --output table
```

Expected output: every prod database back to `available`, replicas re-synced.

### 2. Switch the matcher to exact tag values

Replace substring matching with exact equality on the full tag key and value. If the scheme uses `env=prod-replica` and `env=dev`, the filter must match the whole value, never a fragment of it.

Expected output: the filter matches only the intended databases, verified by listing matches before any action.

### 3. Verify role and tags on every match

For each matched DB instance, describe it and confirm: the tags say what the agent thinks they say, and the instance role (replica vs primary) is what the agent expects. A prod read replica looks idle right up until failover needs it.

Expected output: a per-instance confirmation of role and tags, logged.

### 4. Dry-run the stop list

Output the exact stop list (instance identifiers + tags + roles) without stopping anything. This preview is mandatory on every run, not just the first.

Expected output: a logged preview list with zero stops executed.

### 5. Require approval for database stops

Any stop, delete, or snapshot-deletion on a database needs a human approval against the preview list. Databases are never in the fully-autonomous tier, no matter how confident the idle detection is.

Expected output: a recorded approval before any database is stopped.

## Variant phrasings

### cost agent stopped the prod database during rightsizing

Same fix. Rightsizing a database to zero or stopping it "because it is idle" needs the same approval gate; idle prod replicas are the classic victim.

### agent deleted old snapshots and broke the AMI chain

Same family: the "old" heuristic did not understand snapshot chains. Verify what depends on the snapshot (AMIs, replicas) before deleting, and preview.

### agent's tag-based attribution was wrong

Attribution errors are the warning sign. If the agent misattributes spend by tag, its tag-based destructive filters are misfiring too. Fix the tagging first.

## Why it happens

Read replicas are designed to look expendable: they serve read traffic, they can be rebuilt, and their CPU is often low. An agent optimizing for "idle resources" sees a low-CPU replica and a substring tag match that says "not prod" and acts. It has no model of failover: the replica's value is not its current CPU, it is the recovery it enables. Substring matching just supplies the wrong premise faster.

## Edge cases

- Multi-AZ primaries vs read replicas: stopping a Multi-AZ standby is different from stopping a read replica, and both look "secondary." The role check in step 3 must distinguish them.
- Aurora clusters: stopping individual instances vs stopping the cluster are different operations with different blast radii. The preview must name the exact operation.
- Stopped databases still cost money (storage): the agent's savings math should account for this, or it will "save" money by stopping things that keep billing.
- The agent stopped the instance hours ago: check automated backups and point-in-time recovery windows before restarting, and verify no writes were lost in the gap.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_ml8E4bSa1JFDik75dR5wow
