how to set CrashLoopBackOff alerts before users notice
Sets up CrashLoopBackOff alerts that fire before users notice. Use when you want proactive paging on restarting pods, rollout-time crash detection, or alert routing for on-call. Covers kube-state-metrics signals, PromQL alert shapes, and avoiding alert storms. Not for debugging an active crash loop, Pending pods, or general uptime monitoring.
TL;DR
Alert on the restart signal, not on user-facing symptoms: fire when a pod's restart count climbs or a container sits in waiting with reason CrashLoopBackOff. Use kube-state-metrics so the alert survives the pod itself being flappy, add a short for duration to skip single restarts, and route it to whoever owns the workload. A 5-minute for catches real loops while ignoring one-off crashes.
Error / query
how to set CrashLoopBackOff alerts before users noticeUse this skill when
- You want to be paged when a deployment starts crash-looping
- Rollouts need automatic detection of bad new versions
- You are setting up on-call alerts for Kubernetes workloads
- Restart storms are currently found by users, not by monitoring
Not for this skill when
- A pod is already crash-looping and you need the root cause (diagnose the pod)
- You need alerts for Pending or ImagePullBackOff (different signals)
- You are alerting on application error rates (use service-level alerts)
Steps
Step 1: Confirm kube-state-metrics exposes pod container status
kubectl get pods -n [monitoring-namespace] -l app.kubernetes.io/name=kube-state-metrics
curl -s [kube-state-metrics-endpoint]/metrics | grep kube_pod_container_status_waiting_reason | head -3Expected: metric lines like kube_pod_container_status_waiting_reason{reason="CrashLoopBackOff"}. If absent, install or enable kube-state-metrics; pod-level alerts need it.
Step 2: Add a CrashLoopBackOff alert rule
- alert: PodCrashLooping
expr: max by (namespace, pod, container) (kube_pod_container_status_waiting_reason{reason="CrashLoopBackOff"} == 1) > 0
for: 5m
labels:
severity: warning
annotations:
summary: "Pod {{ $labels.namespace }}/{{ $labels.pod }} crash-looping for 5m"Expected: Prometheus loads the rule without parse errors (promtool check rules passes) and the alert appears in the Pending state when a loop starts.
Step 3: Add a restart-rate alert as a backstop
- alert: PodRestartingFrequently
expr: max by (namespace, pod) (increase(kube_pod_container_status_restarts_total[15m])) > 3
for: 0m
labels:
severity: warningExpected: this catches loops that briefly recover and re-crash (flapping), which the waiting-reason alert can miss between restarts.
Step 4: Route and test the alert
kubectl port-forward -n [monitoring-namespace] svc/alertmanager 9093:9093Expected: in the Alertmanager UI you see the alert routed to the right receiver. Trigger a test by deploying a pod with a bad command, confirm the page arrives within the for window, then delete the test pod.
Variant phrasings
"alert on pod restarts kubernetes"
Use the restart-rate rule in step 3; it is simpler and catches all restart churn.
"prometheus alert crashloopbackoff"
The waiting-reason rule in step 2 is the canonical one; kube-state-metrics is required.
Why it happens
CrashLoopBackOff has a built-in delay that grows with each restart, so a bad deploy can sit broken for many minutes before traffic-based alerts notice. Pod-level metrics fire at the source: kube-state-metrics reports the waiting reason directly from the API server, which is faster and more precise than inferring trouble from error rates.
Edge cases and pitfalls
- A
forof 0m on the waiting-reason alert pages on single crashes; keep 5m for paging, use 0m only for informational routing. - CronJobs restart legitimately; exclude them with a namespace or label matcher or the alert will be noise.
increase()over short windows on counters that reset (pod recreated) can spike; themax byaggregation keeps it per-pod.- If kube-state-metrics itself is down, both rules go stale; alert on its absence separately.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_wbXbHtvWG3AvJBTcE85vRg