spend anomaly alert fired every hour - the agent's baseline never learned the monthly batch job that runs on the 1st
Quiets a spend anomaly detector that pages every month because its baseline never learned the 1st-of-month batch job. Use when alerts fire on a predictable monthly schedule. Key trigger: the same spike repeats on the 1st and the detector has no calendar awareness.
TL;DR: Add the monthly batch window to the anomaly baseline as known-good spend. Mark the 1st of the month (plus a buffer) as an expected-spend period so the detector compares it against other 1sts instead of quiet mid-month days. Alert volume drops to real anomalies only.
ANOMALY: spend $4,210 vs expected $310 (13.6x baseline) at 2026-10-01T02:00Z - paging on-call- Confirm the pattern: pull daily spend for the last 6 months and look at the 1st of each month. Expected: a spike on the 1st every month, roughly the same size each time.
- Define the known window with the job owner: for example 00:00 on the 1st through 06:00 on the 2nd, in the account's timezone. Expected: a documented window the team agrees is the batch job.
- Teach the detector: either exclude the window from alerting, or keep two baselines, one for in-window hours and one for everything else. Expected: next month's 1st produces no page.
- Keep a guardrail on the window itself: alert if 1st-of-month spend deviates from the trailing 3-month 1st-of-month average by more than your threshold. Expected: a genuinely broken batch job that costs 2x normal still pages.
- Document the job in the runbook: owner, expected cost, expected duration, what 'broken' looks like. Expected: the next on-call recognizes it as normal without digging.
Use this when
- Anomaly alerts fire on the same day every month
- A monthly batch, close, or ETL job pages on-call every cycle
- The detector has no calendar or schedule awareness
- Your team has started ignoring the anomaly channel because of the noise
Not for this skill when
- The spike is NOT periodic (investigate it as a real anomaly)
- The batch job's cost is growing month over month (that is a real trend, not a false positive)
- You have fewer than 3 cycles of history (collect the pattern first, then tune)
- The spike moved to a different day (update the window, do not widen it blindly)
Variant phrasings
- anomaly detection false positive monthly batch job
- cost alert fires every month on the 1st
- how to exclude known batch spend from anomaly detection
- spend spike alert same day each month
Why it happens
Most anomaly detectors learn a single rolling baseline, the mean and standard deviation of recent spend. A monthly job looks like a 10x outlier against a baseline built from the other 29 quiet days. The model has no day-of-month feature, so it structurally cannot learn 'the 1st is always expensive'.
Edge cases
- The job sometimes runs on the 2nd (holidays, weekends): allow a buffer day or trigger the window off the job's completion event instead of the calendar
- Multiple monthly jobs on different days: each needs its own window and its own guardrail
- Slow cost drift inside the window (5% more each month) will eventually cross the threshold: that is correct behavior, tune the threshold not the window
- Daylight saving shifts can move the window by an hour: do the math in UTC internally
- If the batch job ever gets decommissioned, remove the window: a stale exclusion hides real anomalies
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_tGNvrV436l6Gb5rfWLfY-A
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.