## TL;DR
A pod in CrashLoopBackOff is restarting over and over because the container's main process keeps exiting. Read the pod's previous-container logs and describe output first: in most cases the cause is a bad config, a crash on startup, or a failed probe. Fix the underlying exit, not the restart loop.

## Error / query
```text
Kubernetes pod CrashLoopBackOff how to diagnose
```

## Use this skill when
- A pod's status shows `CrashLoopBackOff` in `kubectl get pods`
- A deployment rolls out but pods keep restarting and never become Ready
- You see `Back-off restarting failed container` events on a pod
- A job or cronjob pod exits immediately on every retry

## Not for this skill when
- The pod is `Pending` (it never started; check scheduling and resources instead)
- The pod is `ImagePullBackOff` (image fetch failed, a different fix)
- The container runs fine but readiness/liveness probes report the wrong state (probe tuning, not crashes)
- You are debugging a node-level kubelet failure (check node logs, not the pod)

## Steps

### Step 1: List pods and confirm the status
```bash
kubectl get pods -n [namespace] -o wide
```
Expected: you see the pod name with status `CrashLoopBackOff` and a high `RESTARTS` count, which confirms it is the restart loop and not a scheduling problem.

### Step 2: Read the logs of the crashed (previous) container
```bash
kubectl logs [pod-name] -n [namespace] -p
```
Expected: application output ending in an exception, panic, or config error. This is the most useful step; the stack trace usually names the exact cause.

### Step 3: Describe the pod for exit codes and events
```bash
kubectl describe pod [pod-name] -n [namespace]
```
Expected: `Last State: Terminated, Exit Code: [code], Reason: Error` and recent events. Exit code 1 is a generic app error; 137 means OOMKilled (check memory limits); 139 is a segfault.

### Step 4: Check resource limits and probe definitions against the exit code
```bash
kubectl get pod [pod-name] -n [namespace] -o jsonpath='{.spec.containers[0].resources}{"\n"}{.spec.containers[0].livenessProbe}{"\n"}'
```
Expected: you can see the memory/cpu limits and probe settings side by side with the exit code from step 3, which tells you whether to raise limits, fix the app, or adjust the probe timing.

### Step 5: Test the suspected fix outside the restart loop
```bash
kubectl run debug-[name] --image=[image] -n [namespace] --restart=Never --command -- [entrypoint] --dry-run=client -o yaml | kubectl apply -f -
kubectl logs debug-[name] -n [namespace] -f
```
Expected: the one-shot pod runs the same entrypoint without the restart backoff, so you see the failure (or success) in real time. Delete the debug pod when done.

## Variant phrasings

### "pod keeps restarting kubernetes"
Same problem. Follow steps 1-3; the describe output's exit code points at the fix.

### "Back-off restarting failed container"
This is the event text behind CrashLoopBackOff. It means the app exited; go straight to the previous-container logs.

### "container exits with code 137 in k8s"
Almost always OOMKilled. Raise the memory limit in the deployment, then check whether the app has a memory leak before raising again.

## Why it happens
CrashLoopBackOff is not itself an error, it is the kubelet's backoff after a container exits repeatedly. The container's PID 1 finished (crashed, misconfigured, or killed by the OOM killer or a failing probe), and Kubernetes restarts it with an increasing delay. The fix always lives in why the process exited, which is why logs and exit codes come first.

## Edge cases and pitfalls
- `kubectl logs` without `-p` shows the current (possibly empty) container; always check the previous one with `-p` first.
- Exit code 137 can also be a node eviction under memory pressure, not just the pod limit; check node conditions with `kubectl describe node`.
- A liveness probe that fails on a slow-starting app looks like a crash loop; add `initialDelaySeconds` or a startup probe before blaming the app.
- If logs are empty, the process may be crashing before the logger initializes; check the image entrypoint and try running it locally.
- Crash loops on only one node point at a node problem (disk pressure, bad image layer); compare with pods scheduled elsewhere.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_follC7JNwieSuJPUO2LC2Q
