# TL;DR
Averages hide bursts. Judge idleness on the p95/p99 or Maximum CPU over two weeks, not the mean: if the peaks are doing real work, the instance is not idle no matter how low the average sits. One line changes the verdict from "shut it down" to "leave it alone".

```text
agent flagged EC2 instances as idle using CPU average  -  the p99 spikes meant they were handling burst API traffic
```

## Steps

1. Pull the high-percentile view for a flagged instance:
```
aws cloudwatch get-metric-statistics --namespace AWS/EC2 --metric-name CPUUtilization --dimensions Name=InstanceId,Value=[instance-id] --start-time 2026-09-24T00:00:00Z --end-time 2026-10-08T00:00:00Z --period 3600 --statistics Maximum --extended-statistics p95
```
Expected: you see the peaks the average erased. If p95 is above your idle threshold, it is not idle.

2. Change the rule: idle means p95 CPU below threshold AND maximum below a burst cap for the whole window.
Expected: bursty API servers drop off the kill list.

3. Corroborate with a second signal before acting: request count, network throughput, or load balancer target metrics.
Expected: at least two signals agree the instance is quiet before it gets flagged.

4. Re-run the detector on last week's flags.
Expected: the false-positive rate collapses; the remaining flags are genuinely quiet.

## Use this when
- "idle" instances serve bursty traffic
- rightsizing recommendations would cut headroom the peaks need
- the detector only looks at averages

## Not for this skill when
- the workload is genuinely steady-state (averages are fine there)
- you are sizing for sustained throughput rather than peaks (use p50 then)
- the instance is idle by every statistic (shut it down already)

## Variant phrasings
- "EC2 idle detection CPU average wrong"
- "bursty workload flagged as idle"
- "CloudWatch average vs p99 rightsizing"

## Why it happens
The Average statistic is the default in every example, and it answers the wrong question. Idleness is about the absence of peaks, not the level of the mean: a server at 5% average with 80% bursts is a busy server with quiet gaps, and averaging those gaps away manufactures a false "idle".

## Edge cases
- Scheduled batch instances look bursty on purpose. Exclude known batch windows before judging.
- Very short spikes (under the metric period) still hide. Use 1-minute periods for spiky fleets if the extra metric cost is worth it.
- GPU and memory-bound instances can be busy at low CPU. Add the right metric for the actual bottleneck.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_zcfc50OaiPkbz5iD1GqUGA
