## TL;DR

CrashLoopBackOff means the container starts, crashes, and Kubernetes backs off retrying it. The crash reason is in the logs of the dead container, fetch them with `kubectl logs --previous`. Read the actual error, fix the app or its config, and the loop stops.

## The error

```text
STATUS: CrashLoopBackOff
RESTARTS: 7 (4m ago)
```

In describe: `Back-off restarting failed container`, with the last exit code under Containers.

## Use this when

- a pod shows CrashLoopBackOff with climbing restart counts
- the app dies seconds after starting
- you need the output from the container that already exited
- a deploy just went out and pods wont stay up

## Not for

- ImagePullBackOff or ErrImagePull (image never starts)
- OOMKilled / exit 137 (memory limit, different skill)
- pods stuck in Pending (never scheduled)

## Steps

1. Look at the current and previous logs. Previous is the important one:

```bash
kubectl logs [pod-name] -n [namespace]
kubectl logs [pod-name] -n [namespace] --previous
```

Expected: `--previous` shows the output of the crashed run. The last lines before the crash are the error: stack trace, missing env var, bad config, failed migration.

2. If the pod has multiple containers, target the crashing one:

```bash
kubectl get pod [pod-name] -n [namespace] -o jsonpath='{.status.containerStatuses[*].name}'
kubectl logs [pod-name] -n [namespace] -c [container-name] --previous
```

Expected: logs for the specific container. Sidecars are often fine while the app container dies.

3. Check events for the exit code and the kill reason:

```bash
kubectl describe pod [pod-name] -n [namespace] | tail -30
```

Expected: `Last State: Terminated, Reason: Error, Exit Code: 1`. Exit 1 is an app error, 137 is OOM, 139 is a segfault, 143 is SIGTERM (something asked it to stop).

4. Test the crash locally-ish by running the container's command by hand:

```bash
kubectl run debug --rm -it --image=[image] --restart=Never -- [command] [args]
```

Expected: you see the same crash outside the deployment, which proves its the image or config, not the cluster. If it works here, suspect probes, resources, or mounted volumes.

5. Common quick fixes by symptom:

- Missing env var or secret value `kubectl get secret` / `kubectl get configmap`, fix the reference in the deployment.
- Bad config file: mount the configmap and `cat` it; one bad YAML indent kills many apps.
- Port already in use: two containers in one pod binding the same port.
- Permission denied writing: check volume mounts and fsGroup / runAsUser.

6. Deploy the fix and watch the restarts stop:

```bash
kubectl apply -f deployment.yaml
kubectl get pod [pod-name] -n [namespace] -w
```

Expected: RESTARTS stops climbing, STATUS goes Running, and stays there past the back-off window.

### Variant: no logs at all, container dies instantly

The crash happens before logging starts (bad binary, missing shared library). Check the exit code and run the image manually per step 4. Also try `kubectl logs` without --previous right after a fresh crash.

### Variant: crashes only under the deployment, fine when run manually

Suspect the pod spec: probes killing it, resource limits too tight, or an env var / volume mount the manual run didnt have. Diff the two environments.

### Variant: intermittent CrashLoopBackOff every few hours

Likely a slow leak or a downstream dependency flapping. Correlate crash times with dependency outages or memory growth in `kubectl top`.

## Why it happens

Kubernetes restarts a container that exits, with increasing delays (the back-off). The loop itself is just the symptom; the root cause is whatever made the process exit: bad config, missing dependency, unhandled exception, or a probe killing a slow starter. Logs from the dead run are the fastest path to the cause.

## Edge cases

- Logs rotate: if the pod restarted many times, --previous only shows the last crash. Catch it early or ship logs to a collector.
- Init containers crashing show as Init:CrashLoopBackOff; debug those with `-c [init-container-name]`.
- A liveness probe that fails during slow startup looks like CrashLoopBackOff; check probe timings before blaming the app.
- Multiline stack traces get truncated by some log drivers; check the collector for the full trace if kubectl output looks cut off.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_Ea--EAMdYzxAvoB3dH8oTw
