how to debug init container failures
Debugs Kubernetes init container failures. Use when a pod is stuck in Init state, init containers crash or never complete, or the main container never starts. Triggers: Init:CrashLoopBackOff, Init:Error, pod stuck in Init, init container exit codes. Not for: main container crashes (regular CrashLoopBackOff), image pull failures, pods stuck in Pending.
TL;DR
If a pod sits in Init state, one of its init containers is failing or hanging: check init container statuses with kubectl describe, read their logs with -c [init-name], and fix whatever the init step needs (usually a dependency not ready or a bad command). The main container cant start until every init container exits 0.
The error
STATUS: Init:CrashLoopBackOff
STATUS: Init:0/2In describe: Init Containers: section shows which one failed and its exit code.
Use this when
- pod stuck in Init:0/2 or similar
- Init:CrashLoopBackOff or Init:Error status
- main container never starts
- init container ran fine before and broke after a change
Not for
- main app container crashes (thats regular CrashLoopBackOff)
- image pull errors on any container
- pods never scheduled (Pending)
Steps
- See which init container is stuck and its exit code:
kubectl describe pod [pod-name] -n [namespace]Expected: under Init Containers:, one shows State: Waiting, Reason: CrashLoopBackOff or Terminated, Exit Code: 1. The name tells you which init step failed.
- Read that init container's logs:
kubectl logs [pod-name] -n [namespace] -c [init-container-name]
kubectl logs [pod-name] -n [namespace] -c [init-container-name] --previousExpected: the actual error: connection refused to a dependency, missing file, bad script. Init containers often wait on things; distinguish crash output from waiting output.
- Check what the init container is supposed to do:
kubectl get pod [pod-name] -n [namespace] -o jsonpath='{.spec.initContainers[*].name}'
kubectl get pod [pod-name] -n [namespace] -o yaml | grep -A 12 "initContainers:"Expected: the command, image, and env of each init container. Common patterns: wait-for-it scripts, migrations, config downloads, permission fixes.
- Test the init logic by hand with a debug pod:
kubectl run init-debug --rm -it --image=[init-image] --restart=Never -n [namespace] -- [init-command]Expected: reproduces the failure outside the pod lifecycle. If it hangs, its waiting on something (DNS name wrong, service not up yet); if it errors, the command or inputs are wrong.
- Common fixes:
- waiting on a service: make the wait loop retry with a timeout and log progress; check the target service actually exists and has endpoints
- migration failing: run it against a copy first, check DB credentials and connectivity from inside the cluster
- permission fix failing: the volume may not be mounted yet at init time, or fsGroup already handles it
- wrong image tag: init container images get stale while app images move on
- Apply the fix and watch the pod proceed:
kubectl apply -f deployment.yaml
kubectl get pod [pod-name] -n [namespace] -wExpected: Init:0/2 becomes Init:1/2, Init:2/2, then the pod moves to Running as the main container starts.
Variant: init container succeeds but takes forever
Its probably a wait loop with no timeout polling something that will never be ready. Add logging to the loop, verify the dependency independently, and set a sane timeout so it fails loudly instead of hanging.
Variant: Init:Error immediately, no logs
The container cant even start: bad image, bad command path, or missing volume mount. Check events for the low-level error.
Variant: worked until the cluster upgraded
Init container relied on something the upgrade changed (DNS behavior, service account token path, deprecated API). Diff the init container's assumptions against the new version.
Why it happens
Init containers run to completion in order before the main container starts. Theyre meant for setup, but setup depends on the world: databases, services, files, permissions. When a dependency is missing or the init logic is wrong, the pod parks in Init state forever, and the app never gets a chance to run.
Edge cases
- Init containers restart on failure like regular containers (CrashLoopBackOff applies to them too).
- Resource limits apply to init containers; the effective pod request is the max across init containers, which surprises schedulers.
restartPolicy: Neverjobs with failing init containers just sit there; check the job, not just the pod.- Sidecar-style init containers (restartable init,
restartPolicy: Alwayson the init container) keep running; make sure yours is meant to exit.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_j4PA1g8QGzIGh9HfGoLcuw
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.