## TL;DR
Alert on the restart signal, not on user-facing symptoms: fire when a pod's restart count climbs or a container sits in waiting with reason CrashLoopBackOff. Use kube-state-metrics so the alert survives the pod itself being flappy, add a short `for` duration to skip single restarts, and route it to whoever owns the workload. A 5-minute `for` catches real loops while ignoring one-off crashes.

## Error / query
```text
how to set CrashLoopBackOff alerts before users notice
```

## Use this skill when
- You want to be paged when a deployment starts crash-looping
- Rollouts need automatic detection of bad new versions
- You are setting up on-call alerts for Kubernetes workloads
- Restart storms are currently found by users, not by monitoring

## Not for this skill when
- A pod is already crash-looping and you need the root cause (diagnose the pod)
- You need alerts for Pending or ImagePullBackOff (different signals)
- You are alerting on application error rates (use service-level alerts)

## Steps

### Step 1: Confirm kube-state-metrics exposes pod container status
```bash
kubectl get pods -n [monitoring-namespace] -l app.kubernetes.io/name=kube-state-metrics
curl -s [kube-state-metrics-endpoint]/metrics | grep kube_pod_container_status_waiting_reason | head -3
```
Expected: metric lines like `kube_pod_container_status_waiting_reason{reason="CrashLoopBackOff"}`. If absent, install or enable kube-state-metrics; pod-level alerts need it.

### Step 2: Add a CrashLoopBackOff alert rule
```yaml
- alert: PodCrashLooping
  expr: max by (namespace, pod, container) (kube_pod_container_status_waiting_reason{reason="CrashLoopBackOff"} == 1) > 0
  for: 5m
  labels:
    severity: warning
  annotations:
    summary: "Pod {{ $labels.namespace }}/{{ $labels.pod }} crash-looping for 5m"
```
Expected: Prometheus loads the rule without parse errors (`promtool check rules` passes) and the alert appears in the Pending state when a loop starts.

### Step 3: Add a restart-rate alert as a backstop
```yaml
- alert: PodRestartingFrequently
  expr: max by (namespace, pod) (increase(kube_pod_container_status_restarts_total[15m])) > 3
  for: 0m
  labels:
    severity: warning
```
Expected: this catches loops that briefly recover and re-crash (flapping), which the waiting-reason alert can miss between restarts.

### Step 4: Route and test the alert
```bash
kubectl port-forward -n [monitoring-namespace] svc/alertmanager 9093:9093
```
Expected: in the Alertmanager UI you see the alert routed to the right receiver. Trigger a test by deploying a pod with a bad command, confirm the page arrives within the `for` window, then delete the test pod.

## Variant phrasings

### "alert on pod restarts kubernetes"
Use the restart-rate rule in step 3; it is simpler and catches all restart churn.

### "prometheus alert crashloopbackoff"
The waiting-reason rule in step 2 is the canonical one; kube-state-metrics is required.

## Why it happens
CrashLoopBackOff has a built-in delay that grows with each restart, so a bad deploy can sit broken for many minutes before traffic-based alerts notice. Pod-level metrics fire at the source: kube-state-metrics reports the waiting reason directly from the API server, which is faster and more precise than inferring trouble from error rates.

## Edge cases and pitfalls
- A `for` of 0m on the waiting-reason alert pages on single crashes; keep 5m for paging, use 0m only for informational routing.
- CronJobs restart legitimately; exclude them with a namespace or label matcher or the alert will be noise.
- `increase()` over short windows on counters that reset (pod recreated) can spike; the `max by` aggregation keeps it per-pod.
- If kube-state-metrics itself is down, both rules go stale; alert on its absence separately.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_wbXbHtvWG3AvJBTcE85vRg
