readiness probe failing during deploys: how to tune
Tunes readiness probes that fail during deploys and flap pods out of service. Use when new pods never become Ready, rolling updates stall on readiness, or probes pass manually but fail under kubelet timing. Covers periodSeconds, failureThreshold, initialDelaySeconds, and startup probes. Not for liveness crashes, Pending pods, or application bugs.
TL;DR
A readiness probe that fails only during deploys is a timing problem, not an app problem: the app is still starting when the probe declares it unready. Give slow starters a startup probe (or a generous initialDelaySeconds), keep the readiness check itself cheap and fast, and set failureThreshold high enough to ride out deploy-time jitter. Tune the probe to the app's real startup curve, not to the happy path.
Error / query
readiness probe failing during deploys: how to tuneUse this skill when
- New pods never reach Ready during a rollout
- Rollouts stall with
waiting for deployment spec update to be observed - The readiness endpoint answers fine when you curl it manually
- Pods flap Ready/NotReady under load or at startup
Not for this skill when
- The container is crash-looping (fix the crash, not the probe)
- Liveness is killing a healthy app (tune liveness separately)
- The app genuinely returns errors (fix the app)
Steps
Step 1: Watch what the probe actually sees during a deploy
kubectl get events -n [namespace] --field-selector reason=Unhealthy --sort-by=.lastTimestamp | tail -10Expected: Readiness probe failed: messages with the concrete error (connection refused, 500, timeout). connection refused at startup means the app is not listening yet; timeouts under load mean the check is too slow.
Step 2: Give slow starters a startup probe
startupProbe:
httpGet:
path: /healthz
port: [port]
failureThreshold: 30
periodSeconds: 10Expected: the startup probe gates readiness and liveness for up to 5 minutes (30 x 10s). Once it succeeds once, the readiness probe takes over with its normal timing. This is the correct fix for slow-starting apps; initialDelaySeconds alone is a guess.
Step 3: Make the readiness check cheap and honest
readinessProbe:
httpGet:
path: /ready
port: [port]
periodSeconds: 10
timeoutSeconds: 2
failureThreshold: 3Expected: a fast check (under 2s) against an endpoint that reflects real readiness (dependencies up, caches warm), not just 200 OK from the framework. Cheap checks do not flap under load.
Step 4: Verify with a rollout and watch pod state transitions
kubectl rollout restart deployment/[deployment] -n [namespace]
kubectl get pods -n [namespace] -wExpected: new pods go Pending to Running to Ready without long NotReady stretches, and the rollout completes. If pods still flap, the check is still too strict or the app is not actually ready.
Variant phrasings
"readiness probe fails kubernetes deployment"
Same tuning flow. Startup probe first, then check cost and thresholds.
"pod not ready after deploy"
Readiness is the gate. Steps 1 and 2 cover the deploy-time case.
Why it happens
At deploy time everything is worst-case: cold caches, JIT warmup, migrations running, and thundering-herd traffic. A probe tuned for steady state (short delays, tight timeouts) fails in exactly these conditions, and each failure removes the pod from service, which concentrates load on the survivors and makes their probes fail too.
Edge cases and pitfalls
- Using the liveness endpoint for readiness is common but wrong if liveness is lax; keep the semantics separate.
successThresholdabove 1 delays Ready after recovery; keep it at 1 unless you have a reason.- Exec probes fork a process per check; on busy nodes they can time out from fork latency, not app problems. Prefer httpGet or tcpSocket.
- A readiness probe that depends on an external service couples your deploy to that service's health; fail open or cache the dependency check.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_Tt6xagZTRzugZlP8GNw47g
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.