## TL;DR

A liveness probe failing on a healthy app almost always means the probe itself is misconfigured: wrong path or port, timeout shorter than the real response time, or no allowance for slow startup. Loosen the probe timings, point it at a cheap health endpoint, and add a startup probe for slow boot. The app was fine; the probe was lying.

## The error

```text
Liveness probe failed: Get "http://YOUR-pod-ip:8080/health": context deadline exceeded
Warning  Unhealthy  Liveness probe failed
```

Symptom: pod restarts even though the app works when you hit it directly.

## Use this when

- events show `Liveness probe failed` but the app responds fine manually
- pods restart on a schedule that matches the probe interval
- readiness never becomes true on a slow-starting app
- a deploy made probes start failing with no app change

## Not for

- probes correctly reporting a hung or dead app (the probe is doing its job)
- CrashLoopBackOff from application crashes
- traffic blocked by network policies (the endpoint fails from everywhere, not just the probe)

## Steps

1. Confirm the probe config and hit the endpoint yourself:

```bash
kubectl get pod [pod-name] -n [namespace] -o yaml | grep -A 10 livenessProbe
kubectl port-forward pod/[pod-name] 8080:8080 -n [namespace]
curl -sv YOUR_THE_HOST loopback address:8080/health -o /dev/null -w "time: %{time_total}s code: %{http_code}\n"
```

Expected: you get a 200 and a real response time. If the endpoint takes 4s and the probe timeout is 1s, you found it.

2. Fix the timings. The defaults are aggressive for real apps:

```yaml
livenessProbe:
  httpGet:
    path: /healthz
    port: 8080
  initialDelaySeconds: 30
  periodSeconds: 20
  timeoutSeconds: 5
  failureThreshold: 3
```

Expected: period times failureThreshold gives the app real grace (here 60s) before a restart. Timeout should exceed the endpoint's p99 latency.

3. Point the probe at a cheap endpoint. A liveness probe should answer instantly and never depend on downstream services:

- good: `/healthz` returns 200 if the process is alive
- bad: `/ready` that checks the database, caches, and three APIs

Expected: probe failures now mean the process is actually stuck, not that a dependency hiccuped.

4. Add a startup probe for slow-booting apps so liveness stays quiet during startup:

```yaml
startupProbe:
  httpGet:
    path: /healthz
    port: 8080
  failureThreshold: 30
  periodSeconds: 10
```

Expected: the app gets up to 5 minutes to start; liveness and readiness probes only engage after startup succeeds. This alone fixes most "probe kills slow starter" loops.

5. Apply and watch the restarts stop:

```bash
kubectl apply -f deployment.yaml
kubectl describe pod [pod-name] -n [namespace] | grep -i unhealthy
```

Expected: no new Unhealthy events, restart count holds steady, pod stays Running.

### Variant: probe fails with connection refused but the app is up

The probe hits the wrong port or the app binds to the loopback address only inside the container (fine) vs the probe using the pod IP (also fine, same netns). Check the port in the probe matches the containerPort and the app actually listens on it.

### Variant: exec probe fails intermittently

Exec probes fork a process every period; on loaded nodes they time out. Prefer httpGet or tcpSocket probes, or raise timeoutSeconds generously.

### Variant: everything passes except during deploys

Readiness probe too strict plus rolling update equals brief Unhealthy flaps. Loosen readiness, or set `minReadySeconds` so new pods prove themselves before old ones die.

## Why it happens

Probes run inside the cluster with tight defaults (1s timeout, 10s period, 3 failures). Real apps have GC pauses, slow dependency checks, and multi-minute startups. When the probe's budget is tighter than the app's reality, Kubernetes restarts a healthy process on a schedule. Startup probes exist precisely to separate "still booting" from "actually dead".

## Edge cases

- Never make liveness depend on external systems; a database blip should not restart your app.
- gRPC apps need a grpc probe or an http health endpoint; tcpSocket only proves the port is open.
- On very small nodes, exec probes can starve; httpGet is cheaper.
- Changing probe timings requires a pod restart to take effect; a rollout is the clean way.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_BwZxrsTEb6mCsOdyQuBDfg
