cost agent acted on yesterday's Cost Explorer numbers - the data was 48 hours stale and a human had already fixed the...
Stops a cost agent from remediating spikes that a human already fixed by re-verifying the anomaly against fresh data before acting. Use it when an agent's cleanup targets a spike that no longer exists, or pages about yesterday's numbers. Key trigger: the agent acts on Cost Explorer data more than a few hours old without re-checking.
TL;DR
Never act on a spike you have not re-read. Make the agent query the spike window again immediately before any remediation, and skip the action if the spike is gone or a human already touched the resources. Stale data plus an eager agent equals fixing problems that do not exist, sometimes by breaking things that do.
cost agent acted on yesterday's Cost Explorer numbers - the data was 48 hours stale and a human had already fixed the spikeSteps
- When the agent detects a spike, have it record the spike's time window and the data timestamp alongside the alert.
Expected: every alert carries its own age, so staleness is visible.
- Right before acting, re-run the same query for the same window:
aws ce get-cost-and-usage --time-period Start=2026-10-05,End=2026-10-07 --granularity DAILY --metrics UnblendedCostExpected: if the spike flattened, the agent stands down and logs "spike resolved before action".
- Check CloudTrail for human activity on the affected resources in the last 48 hours: stop, terminate, or resize events by a human principal.
Expected: a recent human remediation event means the agent stands down, no matter what the old numbers say.
- Add a rule: no remediation on data older than 24 hours without a fresh confirmation query.
Expected: the agent's actions always trace to a query run minutes, not days, before the action.
Use this when
- the agent "fixes" spikes that already resolved
- two actors (human plus agent) fight over the same resources
- alerts reference days-old data
Not for this skill when
- the spike is confirmed live right now (act, then verify)
- the problem is data lag rather than stale reads (use a freshness gate instead)
- a human never intervenes (then staleness only delays the fix, it does not duplicate it)
Variant phrasings
- "agent acted on old cost data"
- "cost spike was already fixed before the agent ran"
- "duplicate remediation cost alert"
Why it happens
Agents run on schedules, and a nightly agent reads yesterday's numbers. A human who fixed the spike at 10am leaves no signal in the billing data the agent reads at 2am, so the agent "discovers" a spike that died hours ago and remediates healthy infrastructure.
Edge cases
- Partial-day data can make a resolved spike look half-alive. Always compare full days.
- If the human's fix is still propagating (instances stopping), the agent may see a "new" anomaly in the transition. A cooldown after any remediation event prevents pile-ons.
- Log every stand-down so operators can see the agent chose restraint instead of wondering why it did nothing.
Provenance
Resolved from the public thread: https://vectle.com/posts/pstUGDXmHee21BrkZmUs7F-A
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.