cost agent read CloudWatch metrics in the wrong timezone - its "business hours" window was off by 8 hours and it...
Fixes a cost agent that kills the wrong instances because its "business hours" window is computed in the wrong timezone. Use it when shutdown schedules hit night-shift or follow-the-sun workers, or when the agent's quiet hours are off by exactly the UTC offset. Key trigger: the damage is offset by a round number of hours.
TL;DR
CloudWatch timestamps are UTC, always. Convert the business-hours window to UTC once, explicitly, and filter metrics on the UTC window; never let local time leak into the query. An 8-hour offset in the window is an 8-hour offset in who gets shut down.
cost agent read CloudWatch metrics in the wrong timezone - its "business hours" window was off by 8 hours and it killed the night-shift workersSteps
- Confirm the damage pattern: list the stopped instances' actual busy hours in UTC and compare them to the window the agent used.
Expected: the two windows are the same shape, shifted by the UTC offset. That shift is the whole bug.
- Rewrite the window as explicit UTC. Business hours 9:00 to 17:00 America/Chicago is 14:00 to 22:00 UTC in standard time. Write the UTC numbers into the config, not the conversion.
Expected: the config contains UTC hours with a comment naming the source zone.
- Handle daylight saving explicitly: the UTC offset for the source zone changes twice a year, so either recompute on schedule or store the window in the local zone and convert at query time with a real timezone library.
Expected: the spring and fall transitions do not silently move the window by an hour.
- Dry-run the schedule against last week's metrics before re-enabling it.
Expected: zero instances that were busy during UTC business hours appear in the stop list.
Use this when
- shutdowns land outside intended hours
- the miss is a round number of hours
- the fleet spans timezones and the agent assumed one
Not for this skill when
- the window is right but the instances are still wrong (that is a tagging or scope bug)
- the fleet is UTC-native already
- you want 24/7 parking rather than business-hours parking (drop the window entirely)
Variant phrasings
- "CloudWatch timezone business hours wrong"
- "shutdown schedule killed night shift instances"
- "agent used local time instead of UTC"
Why it happens
The developer's machine, the ticket, and the runbook all speak local time, but CloudWatch speaks UTC. Somewhere between the ticket ("park the dev fleet 7pm to 7am") and the metric query, the zone gets dropped, and the filter silently slides by the offset. It passes every test run in the same zone as the developer.
Edge cases
- Follow-the-sun fleets need per-region windows, not one global window.
- Cron-based schedulers have their own timezone setting. Check the scheduler's zone separately from the metric query's zone.
- When in doubt, log both the local and the UTC timestamp on every action so the next investigation takes minutes.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_M5xC62gFTrAMNGs1FGVY0w
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.