CloudWatch "Rate exceeded" on GetMetricData: how to batch requests
Batches CloudWatch GetMetricData calls to avoid rate limits. Use when monitoring hits rate exceeded errors, when metric math is chatty, or when designing metric collection. Not for CloudWatch Logs or other APIs.
TL;DR
GetMetricData allows 500 metrics per call; rate exceeded means you are calling too often with too few metrics per call. Batch aggressively: pack hundreds of metric queries into each call, cache results for the metric period, and stop polling faster than the metric resolution. Most rate problems come from one-metric-per-call patterns that a single batched call replaces.
The query
CloudWatch "Rate exceeded" on GetMetricData: how to batch requestsUse this when
- GetMetricData returns rate exceeded errors
- Metric collection is chatty
- Designing CloudWatch-based monitoring
- After adding many new metrics
Not for when
- Other CloudWatch API rate limits (logs, etc.)
- Metric math expression errors
- Missing metrics (different issue)
Steps
Step 1: Find the chatty callers
Identify which code calls GetMetricData and with how many metrics per call. One metric per call in a loop is the classic pattern; it burns the quota hundreds of times faster than batched calls. Expected output: the callers and their batch sizes mapped.
Step 2: Batch to hundreds of metrics per call
Restructure calls to pack up to 500 metric queries per GetMetricData request. This single change usually eliminates rate errors entirely. Expected output: call count dropping by orders of magnitude.
Step 3: Cache at the metric period
Cache results for at least the metric's period (usually 60s or 300s). Polling faster than the data refreshes wastes quota for identical numbers. Expected output: no duplicate queries within a period.
Step 4: Add backoff with jitter anyway
Even batched callers hit limits during bursts. Implement exponential backoff with jitter so the calls you do make degrade gracefully under throttling. Expected output: throttling absorbed without errors surfacing.
Step 5: Consider metric streams for high volume
For very high metric volumes, CloudWatch metric streams (Kinesis Firehose delivery) replace polling entirely. Polling does not scale to thousands of metrics; streams do. Expected output: the collection architecture matched to the metric volume.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_VEiIUkMU66dvdCPHHub8hg