cost agent's baseline included the month we ran the load test - every normal month since looks like an anomaly
Fixes a cost anomaly baseline contaminated by a load-test month, which makes every normal month since look anomalous. Use when the detector flags every period after a one-off event. Key trigger: the training window included a load test, migration, or promo with 2-3x normal spend.
TL;DR: Drop the load-test month from the training data and rebuild the baseline. Every month since has looked anomalous because the model learned 'normal' from artificially inflated traffic. After the rebuild, normal months score clean and the threshold means something again.
ANOMALY: monthly spend $44,000 vs baseline $71,000 (-38%) - flagged every month since March- Identify the contaminated window: find the load-test dates inside the training range. Expected: one month with 2-3x normal traffic and spend, matching the test calendar.
- Exclude that window and retrain on the remaining months. Expected: the baseline drops to true normal and the band tightens.
- Backtest the last 3 normal months against the new baseline. Expected: no anomalies flagged on any of them.
- Add a permanent rule: load-test and chaos-test windows are tagged in the ops calendar and auto-excluded from training data. Expected: the next load test cannot poison the model.
- Sanity-check the restored sensitivity: with clean data the band should be tight enough that a real 30% swing alerts. Expected: the threshold is meaningful again, not just decorative.
Use this when
- Every period looks anomalous after a one-off event
- The training window included a load test, migration, or promo
- The detector's baseline is visibly higher than current normal spend
- The team has started ignoring the anomaly channel entirely
Not for this skill when
- Only some months flag (different pattern, check seasonality instead)
- The load test is still running (exclude it going forward, do not retrain mid-test)
- Spend genuinely dropped because you rightsized (re-baseline deliberately as a decision, not as a fix)
- The anomaly is positive (spend up), not negative (different investigation)
Variant phrasings
- baseline contaminated by load test
- every month flagged as anomaly after test
- how to exclude load test from cost baseline
- detector baseline too high after one-time event
Why it happens
The training window is treated as ground truth, and one artificial month shifts the mean up and widens the band. From then on, real normal looks like a permanent negative anomaly. Teams learn to ignore the detector, which is worse than having no detector: the one time a real anomaly fires, nobody is watching.
Edge cases
- Multiple contaminated months (quarterly load tests): exclude all of them, and make sure enough clean history remains for a stable baseline
- If the load test revealed you need that capacity permanently, the 'contamination' is the new normal: re-baseline on purpose and document why
- Negative anomalies (spend dropping) are still worth a glance: they can mean broken telemetry or a dead data pipeline, not savings
- Backfill jobs and migrations contaminate the same way: the exclusion rule should cover any scheduled artificial-traffic window, not just load tests
- If clean history is too short after exclusions, widen the threshold temporarily rather than shipping a twitchy detector
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_bR-fPj0sW1U-CuOzB6svPQ
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.