agent flagged the warm DR region as waste - the failover stack it called idle is the company's only disaster recovery...
Stops a cost agent from flagging disaster recovery infrastructure as waste. Use when the agent targets a DR region, failover stack, or warm standby for deletion or shutdown. It tags DR resources explicitly, excludes them from idle sweeps by policy while still reporting their cost, and verifies recovery objectives after any nearby change.
TL;DR: Tag every DR resource so the agent can recognize it, then exclude tagged resources from idle sweeps by policy. DR infrastructure is idle by design - a warm standby that is busy is a DR plan that already failed. The agent may report what DR costs, but it must never recommend removing it.
agent flagged the warm DR region as waste - the failover stack it called idle is the company's only disaster recovery copy- Confirm what the agent flagged and whether it acted or only recommended. Check the DR region's resources against the agent's action log. Expected: you know exactly which DR resources were touched, if any.
- Tag every DR resource with a clear marker (for example a tag named dr-role with values like primary, standby, or replica) and document what each value means. Expected: a tag query returns the full DR footprint in one pass.
- Add a policy exclusion: resources carrying the DR tag are never idle candidates and never auto-stopped. The agent may report their cost, but may not recommend removing them. Expected: the next sweep lists DR cost as informational only, with no action attached.
- After any change near DR, verify the recovery objectives still hold. Run a failover drill or review the last drill report against the RTO and RPO targets. Expected: a dated drill result showing the DR copy still meets its objectives.
Use this when
- A cost agent flagged DR, failover, or backup-region infrastructure as idle or wasteful
- The agent proposed deleting or stopping resources in a secondary region
- You need the agent to report DR cost without acting on it
Not for this skill when
- The DR setup is genuinely oversized (for example a full prod clone for a 24-hour RTO) - that is a DR architecture review, not an agent bug
- There are two DR copies and one is redundant - a human should make that call explicitly
- The flagged resources are backups, not live DR - backup lifecycle policy is the right tool
Variant phrasings
- cost agent wants to delete disaster recovery region
- finops tool flagged warm standby region as waste
- how to exclude dr infrastructure from idle resource cleanup
- agent recommended removing failover stack to save cost
Why it happens
DR infrastructure is idle by design - a warm standby that is busy is a DR plan that already failed. Utilization-based agents cannot distinguish idle because useless from idle because waiting for a disaster. Without an explicit marker, the most important idle resources in the account look exactly like the waste the agent was built to find.
Edge cases
- Pilot-light DR (minimal resources that scale up on failover) has a smaller footprint but the same rule applies. Tag it and exclude it.
- DR drills temporarily make the region look busy. The agent should not learn from drill windows that DR is in use either; the tag is the source of truth.
- Cross-region replication traffic shows up as data transfer cost. Do not let the agent optimize it away without understanding the RPO it supports.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_p7VQWG1HdozApGjYrqt6Dg
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.