Kubernetes liveness probe failing but app is fine
Fixes Kubernetes liveness probes that fail while the app is healthy. Use when probes restart a working app, readiness never turns green, or probe timings are too aggressive for real startup. Triggers: Liveness probe failed, Unhealthy, probe timeout, restart loops on a healthy app. Not for: probes correctly catching a dead app, CrashLoopBackOff from app crashes, network policies blocking traffic.
TL;DR
A liveness probe failing on a healthy app almost always means the probe itself is misconfigured: wrong path or port, timeout shorter than the real response time, or no allowance for slow startup. Loosen the probe timings, point it at a cheap health endpoint, and add a startup probe for slow boot. The app was fine; the probe was lying.
The error
Liveness probe failed: Get "http://YOUR-pod-ip:8080/health": context deadline exceeded
Warning Unhealthy Liveness probe failedSymptom: pod restarts even though the app works when you hit it directly.
Use this when
- events show
Liveness probe failedbut the app responds fine manually - pods restart on a schedule that matches the probe interval
- readiness never becomes true on a slow-starting app
- a deploy made probes start failing with no app change
Not for
- probes correctly reporting a hung or dead app (the probe is doing its job)
- CrashLoopBackOff from application crashes
- traffic blocked by network policies (the endpoint fails from everywhere, not just the probe)
Steps
- Confirm the probe config and hit the endpoint yourself:
kubectl get pod [pod-name] -n [namespace] -o yaml | grep -A 10 livenessProbe
kubectl port-forward pod/[pod-name] 8080:8080 -n [namespace]
curl -sv YOUR_THE_HOST loopback address:8080/health -o /dev/null -w "time: %{time_total}s code: %{http_code}\n"Expected: you get a 200 and a real response time. If the endpoint takes 4s and the probe timeout is 1s, you found it.
- Fix the timings. The defaults are aggressive for real apps:
livenessProbe:
httpGet:
path: /healthz
port: 8080
initialDelaySeconds: 30
periodSeconds: 20
timeoutSeconds: 5
failureThreshold: 3Expected: period times failureThreshold gives the app real grace (here 60s) before a restart. Timeout should exceed the endpoint's p99 latency.
- Point the probe at a cheap endpoint. A liveness probe should answer instantly and never depend on downstream services:
- good:
/healthzreturns 200 if the process is alive - bad:
/readythat checks the database, caches, and three APIs
Expected: probe failures now mean the process is actually stuck, not that a dependency hiccuped.
- Add a startup probe for slow-booting apps so liveness stays quiet during startup:
startupProbe:
httpGet:
path: /healthz
port: 8080
failureThreshold: 30
periodSeconds: 10Expected: the app gets up to 5 minutes to start; liveness and readiness probes only engage after startup succeeds. This alone fixes most "probe kills slow starter" loops.
- Apply and watch the restarts stop:
kubectl apply -f deployment.yaml
kubectl describe pod [pod-name] -n [namespace] | grep -i unhealthyExpected: no new Unhealthy events, restart count holds steady, pod stays Running.
Variant: probe fails with connection refused but the app is up
The probe hits the wrong port or the app binds to the loopback address only inside the container (fine) vs the probe using the pod IP (also fine, same netns). Check the port in the probe matches the containerPort and the app actually listens on it.
Variant: exec probe fails intermittently
Exec probes fork a process every period; on loaded nodes they time out. Prefer httpGet or tcpSocket probes, or raise timeoutSeconds generously.
Variant: everything passes except during deploys
Readiness probe too strict plus rolling update equals brief Unhealthy flaps. Loosen readiness, or set minReadySeconds so new pods prove themselves before old ones die.
Why it happens
Probes run inside the cluster with tight defaults (1s timeout, 10s period, 3 failures). Real apps have GC pauses, slow dependency checks, and multi-minute startups. When the probe's budget is tighter than the app's reality, Kubernetes restarts a healthy process on a schedule. Startup probes exist precisely to separate "still booting" from "actually dead".
Edge cases
- Never make liveness depend on external systems; a database blip should not restart your app.
- gRPC apps need a grpc probe or an http health endpoint; tcpSocket only proves the port is open.
- On very small nodes, exec probes can starve; httpGet is cheaper.
- Changing probe timings requires a pod restart to take effect; a rollout is the clean way.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_BwZxrsTEb6mCsOdyQuBDfg
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.