cost agent recommended downsizing a prod database - it didn't know the instance handles the Black Friday traffic spike
Stops a cost agent from rightsizing away the headroom a production database needs for peak events like Black Friday. Use when a downsizing recommendation looks right on average load but ignores known spikes. It feeds the agent a peak-events calendar, gates downsizes on peak-window utilization, and requires the peak window cited in every recommendation.
TL;DR: Block the downsizing recommendation and require the agent to check peak-event utilization before proposing any database shrink. A 30-day lookback in October sees nothing of last November's Black Friday. Feed the agent a calendar of known traffic spikes and gate every recommendation on the p99 load during those windows.
cost agent recommended downsizing a prod database - it didn't know the instance handles the Black Friday traffic spike- Pull the database's utilization over the last 12 months, not the last 30 days. Look at p99 CPU and connection count during known peak events. Expected: spikes 3 to 5 times the recent average around the event dates.
- Give the agent a peak-events calendar: sale dates, product launches, seasonal windows, stored where the agent reads its inputs. Expected: the agent can name which upcoming events affect each database.
- Add a guard rule: never recommend downsizing a database whose peak-window p99 utilization exceeds 60 percent of the proposed smaller instance's capacity. Expected: the bad recommendation disappears while genuinely idle databases still get flagged.
- Require the agent to cite the peak window it checked in every rightsizing recommendation. Expected: each recommendation names the event or states no known peaks, so a human reviewer can sanity-check it.
Use this when
- A cost agent recommended shrinking a database or instance that handles seasonal or event traffic
- Rightsizing output looks right on average load but ignores known spikes
- You need the agent to respect a traffic calendar before recommending downsizes
Not for this skill when
- The workload is genuinely flat year-round with no peaks - standard rightsizing is fine
- The database is already provisioned for peak and you want to cut the baseline - that is an architecture decision, not an agent bug
- The recommendation came from AWS Compute Optimizer rather than your agent - check its lookback settings instead
Variant phrasings
- rightsizing tool wants to downsize database before peak season
- cost optimizer ignored black friday traffic when recommending smaller rds
- agent recommended smaller database instance but traffic spikes quarterly
- how to stop cost agent downsizing prod db
Why it happens
Agents default to short lookback windows (7 to 30 days) because that is what the metrics APIs return cheaply. A 30-day window in October sees nothing of last November's Black Friday. The agent is not wrong about the data it saw; it is wrong about which data matters. Peak capacity planning needs the full annual cycle plus a calendar of future events.
Edge cases
- New services with no history have no peaks to learn. Default to conservative and require human sign-off on the first rightsizing pass.
- One-off events like a viral launch create peaks that never repeat. Mark in the calendar whether each event recurs.
- Read replicas scale differently from primaries. Peak load on the primary does not always mean the replicas need the same headroom.
Provenance
Resolved from the public thread: https://vectle.com/posts/pstxdMbju5CzfaIU4NkVD6NQ
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.