kafka consumer lag monitoring setup
Sets up Kafka consumer lag monitoring so you catch slow consumers before they fall over. Use when you need visibility into consumer group lag, when lag alerts do not exist yet, or when choosing between burrow, kafka-lag-exporter, and built-in tooling. Not for fixing max message size errors, for rebalance storms, or for idempotent producer setup.
TL;DR
Monitor per-partition lag (end offset minus consumer offset) per consumer group, alert on sustained growth rather than absolute values, and use kafka-lag-exporter or Burrow to get the numbers into Prometheus. Lag is a rate problem, not a level.
kafka consumer lag monitoring setupUse this when
- You have no visibility into consumer lag today
- Lag alerts fire too late or too noisily
- You are choosing lag monitoring tooling
Not for this skill when
- Producers get message-too-large errors
- Consumers rebalance constantly
- You need exactly-once producer semantics
Steps
- Get the baseline numbers with the built-in tool first:
kafka-consumer-groups.sh --bootstrap-server YOUR_HOST:9092 \
--group orders-consumer --describeExpected output: per-partition current offset, end offset, and LAG columns. This is the ground truth every exporter ultimately reports.
- Deploy kafka-lag-exporter (or Burrow) to turn those numbers into metrics:
# kafka-lag-exporter: one exporter per cluster, polls consumer groups
# and exposes kafka_consumergroup_group_lag to PrometheusExpected output: kafka_consumergroup_group_lag and kafka_consumergroup_group_lag_seconds metrics in Prometheus, labeled by group, topic, and partition. Burrow is the alternative if you want lag evaluated with its own status logic.
- Alert on lag growth rate, not absolute lag:
# alert when p99 lag grows over 15 minutes, not when it merely exists
- alert: KafkaConsumerLagGrowing
expr: deriv(kafka_consumergroup_group_lag[15m]) > 0Expected output: alerts only when the consumer is actually falling behind, not during normal bursty traffic. Absolute thresholds page you every deploy and every traffic spike.
- Also alert on the consumer being dead, which looks like zero lag growth with no commits:
# a group with no committed offsets for a while is not healthy, it is absent
kafka-consumer-groups.sh --bootstrap-server YOUR_HOST:9092 --listExpected output: the group list. A consumer that crashed shows stale offsets and flat lag, which a growth-rate alert misses. Alert on missing heartbeats or stale kafka_consumergroup_group_max_lag timestamps too.
- Build the dashboard around the questions you will ask at 3am:
- lag per group/topic right now
- lag trend over the last 6 hours
- which partitions are lagging (one hot partition = skew, all = slow consumer)Expected output: a dashboard that distinguishes skew from slowness at a glance. Per-partition lag is the key panel; aggregate lag hides hot partitions.
Variant phrasings
kafka lag exporter vs burrow
kafka-lag-exporter is simpler and Prometheus-native; Burrow adds its own consumer-status evaluation (OK/WARN/ERR/STOPPED). Pick the exporter for metrics pipelines, Burrow for status semantics.
kafka consumer lag alert best practice
Alert on sustained positive derivative of lag plus staleness detection (steps 3-4). Absolute thresholds are either noisy or late.
monitor kafka consumer group lag prometheus
kafka-lag-exporter plus the recording rules above is the standard stack. The built-in JMX metrics cover brokers, not consumer lag.
Why it happens
Kafka exposes consumer offsets and log end offsets, and lag is just their difference, but nothing aggregates or alerts on it out of the box. Teams discover lag when users complain because the raw numbers live behind a CLI nobody watches. An exporter plus rate-based alerts closes that gap.
Edge cases
- Compacted topics: lag in messages is misleading when old messages vanish, watch lag in time (seconds) too.
- Consumer group rebalances reset per-partition assignments, expect brief lag spikes, not alerts.
__consumer_offsetsretention can expire offsets for idle groups, making them look new rather than stale.- Exactly-once consumers commit transactionally, lag semantics stay the same but offset commits lag behind processing slightly more.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_BpuFSXPWWPhUDgi09k-MdA
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.