cost agent missed a real 3x spend spike - its z-score threshold was trained on a noisy quarter full of promo credits
Fixes a cost anomaly detector that missed a real 3x spend spike because its z-score threshold was trained on a quarter full of promo credits. Use when the detector is suspiciously quiet or missed an obvious spike. Key trigger: the training period included credits, refunds, or load tests.
TL;DR: Rebuild the baseline on clean data. Strip promo credits, refunds, and one-time events from the training window, and use a robust statistic (median and MAD) instead of mean and standard deviation. The noisy quarter had inflated the 'normal' band so wide that a real 3x spike fit inside it.
No anomaly: spend $96,000 vs baseline mean $58,000 (stddev $31,000) - within 2 sigma, no alert- Inspect the training window: list the top 10 daily spend values and their line items. Expected: promo-credit days and one-off events showing as huge swings in both directions.
- Clean the data URIs remove days with promo credits, refunds, and known one-time events from the training set, and use gross spend (before credits) as the metric. Expected: the artificial variance is gone from the training data.
- Switch to robust stats: baseline equals the median, spread equals the MAD (median absolute deviation, scaled by about 1.4826 to approximate a standard deviation). Expected: leftover outliers in the training data no longer inflate the band.
- Backtest: run the new detector over the quarter that contained the missed spike. Expected: the 3x spike now fires, and normal days stay quiet.
- Lock the training pipeline: credits and refunds are excluded by line-item type on every retrain, not just this once. Expected: the next promo quarter cannot poison the model again.
Use this when
- The detector missed a spike that is obvious in hindsight
- The training period included promo credits, refunds, or load tests
- Z-score thresholds feel too lax and nothing ever alerts
- You are setting up anomaly detection for the first time on messy historical data
Not for this skill when
- The spike fell inside a known event window (that is a baseline-shape problem, not a training-data problem)
- You have fewer than about 30 clean days of history (robust stats need enough data)
- The 'spike' was legitimate growth (relabel the expectation, do not retune the detector)
- The detector is too noisy rather than too quiet (opposite problem, tighten instead)
Variant phrasings
- anomaly detection missed spend spike
- z-score threshold too high cloud cost
- promo credits breaking anomaly baseline
- detector never alerts on real spikes
Why it happens
Mean and standard deviation are not robust: a few wild days (promo credits landing, a load test, a refund) inflate the standard deviation enormously, so the 2-sigma band becomes wide enough to hide a genuine 3x spike. Median and MAD ignore the wild days, which is exactly what you want from a baseline.
Edge cases
- After cleaning, the band gets tighter and borderline days may start alerting: expect a short tuning period, not a broken detector
- Recurring promos (annual credits) need permanent exclusion rules, not one-off cleaning
- If credits are material to the business, track net and gross spend as two separate series instead of picking one
- Document the MAD scaling constant you use so the next person can reproduce the threshold
- Retrain on a schedule: a baseline trained once and never refreshed slowly goes stale as the business changes
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_URgg1cLbwabESfccpTAisg
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.