cost agent stopped the warm standby database - it didn't know the standby is required for the 15-minute RPO
Fixes a cost agent that stopped a warm standby database without understanding the recovery point objective it supports. Use when a standby or replica database was stopped, deleted, or downsized by automation. It restores replication, records the RPO on the resource where agents can read it, and routes any future standby action with RPO impact to a human.
TL;DR: Restart the standby and verify replication is caught up, then put the RPO requirement on the resource itself as a tag so no agent has to guess. A warm standby doing its job shows low CPU and few connections - the exact signature of an idle database. The 15-minute RPO it protects must be machine-readable, not buried in a runbook.
cost agent stopped the warm standby database - it didn't know the standby is required for the 15-minute RPO- Restart the standby database immediately and verify replication catches up. Watch the replica lag metric until it returns to the normal range. Expected: lag back under the 15-minute RPO budget with replication state healthy.
- Confirm no data loss window was created while the standby was down. Compare the primary's transaction log position against what the standby has applied. Expected: the standby has every transaction, or you have a documented gap to disclose.
- Record the RPO requirement on the standby resource itself as a tag (for example rpo-minutes with the value 15) plus a pointer to the runbook. Expected: anyone or any agent reading the resource sees the RPO without opening a wiki.
- Add an agent guard: stopping or deleting a database with an RPO tag requires the agent to state the RPO impact in its plan, and any nonzero impact routes to a human. Expected: the agent can still flag the standby's cost, but cannot act on it silently.
Use this when
- A cost agent stopped, deleted, or downsized a standby or replica database
- Standby databases keep appearing on idle lists despite serving an RPO
- You need RPO and RTO requirements visible to automation, not just in a runbook
Not for this skill when
- The standby was stopped by a human during maintenance - check the change log, not the agent
- The RPO is actually met by backups rather than the standby - then the standby may genuinely be redundant, but verify with the DR owner first
- Replication lag was already breaching RPO before the agent acted - that is a replication health issue
Variant phrasings
- agent stopped rds read replica needed for rpo
- cost optimizer shut down warm standby database
- how to protect standby databases from automated cleanup
- finops agent killed dr replica
Why it happens
A warm standby doing its job shows low CPU, few connections, and no interesting traffic - the exact signature of an idle database. The RPO it protects lives in a runbook or someone's head, not on the resource. The agent optimized what it could see and never knew the standby was load-bearing.
Edge cases
- Cascading replicas (a standby of a standby) each need the tag. The agent should see the whole chain's RPO, not just the first hop.
- After a failover, the old primary becomes the new standby. Tags and guards must follow the role swap, or the next sweep targets the new standby.
- Stopped standbys still incur storage costs. The agent is not wrong that it costs money; it is wrong that the money is waste.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst4QPrCx8I51hqdLFX0iyLA