how to reduce Prometheus cardinality before it falls over
Reduces Prometheus metric cardinality before it causes outages. Use when Prometheus is slow or OOMing, when active series counts are climbing, or when planning label strategy. Covers finding the offenders and fixing them. Not for remote-write or long-term storage setup.
TL;DR
Cardinality kills Prometheus: every unique label combination is a series, and unbounded labels (user IDs, request paths, pod names in the wrong place) multiply series into the millions. Find the worst offenders with the TSDB head stats, then drop or relabel the bad labels at scrape time. Do this before Prometheus falls over, because after is much harder.
The query
how to reduce Prometheus cardinality before it falls overUse this when
- Prometheus is slow, OOMing, or refusing queries
- Active series counts keep climbing
- You are designing label conventions for new services
- Planning capacity for the monitoring stack
Not for when
- Setting up remote write or long-term storage
- Alert rule design
- Non-Prometheus metrics systems
Steps
Step 1: Find the worst offending metrics
Check the TSDB head statistics for series count by metric name. The top few metrics usually account for most of the cardinality. Then check which labels on those metrics have the most values. Expected output: a ranked list of metrics by series count, with the high-cardinality labels named.
Step 2: Identify unbounded labels
Look for labels with ever-growing value sets: user IDs, email addresses, full URL paths, timestamps, pod names on metrics that outlive pods. These are the cardinality bombs. A label with thousands of values on a frequently scraped metric is the problem. Expected output: the specific label names that must go, per offending metric.
Step 3: Drop or rewrite labels at scrape time
Use relabeling to drop the bad labels before ingestion, or aggregate them (full path becomes route template). Do this in the scrape config so the series never enter the TSDB. Dropping labels loses dimensions; make sure nobody's dashboard depends on them first. Expected output: series count drops sharply after the next scrape cycle; the offending labels gone.
Step 4: Fix the instrumentation at the source
Relabeling treats the symptom. The real fix is changing the client library code or exporter config so the bad labels are never emitted. File the issue with the service owner; relabeling buys time. Expected output: new services ship with bounded label sets; the relabeling rules eventually become unnecessary.
Step 5: Set cardinality guardrails
Add alerts on series count growth and per-metric series limits. Review label conventions in code review for new metrics. Cardinality is a budget; spend it deliberately. Expected output: early warning before the next cardinality crisis, not during it.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_-VNbG9EskhdygVrMYDdyoQ
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.