agent cried wolf on a data-transfer spike - the spike was a one-time S3 cross-region replication for a migration
Stops false data-transfer alerts caused by one-time S3 cross-region replication during a migration. Use when a transfer spike pages on-call but maps to a scheduled data movement. Key trigger: the spike window matches a migration or replication job.
TL;DR: Annotate known one-time data movements before the detector sees them. Add an ops-calendar check to the agent: when a transfer spike correlates with a scheduled migration or replication job, it logs the spike as explained instead of paging. Keep a guardrail so a runaway replication that blows past the migration estimate still alerts.
ANOMALY: data transfer spend $8,900 vs expected $300 (29x baseline) - paging on-call- Confirm the cause: check S3 replication metrics and the migration runbook for the same window. Expected: bytes transferred match the migration plan within a reasonable margin.
- Build the annotation habit: migrations, replication backfills, and DR tests get a dated entry in the ops calendar the detector can read. Expected: the detector can join a spike to a calendar event automatically.
- Add the correlation step to the agent: on a transfer spike, check the calendar and recent replication job history BEFORE paging. Expected: 'explained by migration X' in the log instead of a page.
- Keep the guardrail: alert if transfer cost exceeds the migration's estimated transfer budget by your threshold (say 50%). Expected: a looping or mis-scoped replication still pages.
- After the migration, verify the replication rule is disabled or re-scoped. Expected: no recurring monthly surprise from a forgotten rule left running.
Use this when
- One-time data movements trigger transfer-cost alerts
- Migrations or replication backfills page on-call
- The agent has billing data but no ops calendar to consult
- Data-transfer spend is bursty and the detector models it like smooth compute spend
Not for this skill when
- The transfer spike has no associated job (investigate: could be exfiltration or a misconfigured replication)
- Transfer costs are chronically high (that is an architecture problem, keep the data local or use cheaper paths)
- The spike is in NAT gateway data-processing charges, not S3 transfer (different line item, different fix)
- The replication is ongoing by design (model it as recurring, not one-time)
Variant phrasings
- data transfer spike false positive
- S3 cross region replication cost alert
- migration caused cost anomaly
- one-time replication flagged as anomaly
Why it happens
Data transfer is bursty by nature, and detectors usually model it with the same smooth baseline they use for compute. A one-time replication moves in hours what the baseline expects in a month. The agent sees the line item but not the project behind it, because nothing connects the billing data to the migration plan.
Edge cases
- Replication rules left on after migration turn the 'one-time' spike into a monthly charge: the post-migration check in step 5 is the real saving
- Cross-region transfer pricing differs by direction and region pair: the estimate must use the right rate or the guardrail misfires
- S3 Batch Replication and live replication show up differently: check both when reconciling
- A partial migration retry can double the transfer: the estimate should assume at least one retry
- Transfer spikes during an actual incident (failover, restore) are also 'explained': the calendar should include DR tests and real failovers
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_p-k4Fe1IB6swws2P0PvBeg
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.