Kubernetes "FailedCreatePodSandBox": CNI sandbox creation debugging
Diagnoses the Kubernetes FailedCreatePodSandBox error: reading the pod event, checking kubelet logs, and validating CNI config, binaries, and the pause image. Use it when pods never start and events blame sandbox creation. Not for ImagePullBackOff or CrashLoopBackOff, which fail later in the pod lifecycle.
Fix Kubernetes FailedCreatePodSandBox errors
TL;DR
Describe the pod and read the FailedCreatePodSandBox event, then check kubelet logs for the underlying CNI error. The cause is almost always a broken CNI config, missing CNI binaries, or an unpullable pause image. Fix that layer and pods start. Every pod needs its pause-container network sandbox before any app container can run, so when sandbox creation fails nothing else gets a chance.
Kubernetes "FailedCreatePodSandBox": CNI sandbox creation debuggingSteps
- Read the actual error. Run
kubectl describe pod [pod-name] -n [namespace]and find the FailedCreatePodSandBox event.
Expected: the event message names the underlying failure (CNI plugin error, bad config, image problem).
- Check kubelet logs on the node. Run something like
journalctl -u kubelet --since "30 min ago"and search for CNI mentions, or check your container runtime logs.
Expected: log lines naming the exact CNI call that failed.
- Verify the CNI config. On the node, list the CNI config directory and confirm there is at least one valid network config file in it.
Expected: one default config file present, parseable, and pointing at an installed plugin.
- Verify the CNI binaries. List the CNI bin directory on the node and confirm the binaries match the plugin names in the config.
Expected: every plugin referenced by the config has a binary present.
- Confirm the pause image is pullable. Check that the node's runtime can fetch the configured pause image (network, registry auth, image name).
Expected: the pause image is present locally or pulls cleanly.
- Restart the networking layer after fixing. Restart the CNI DaemonSet pods or the kubelet, then watch a new pod get created.
Expected: new sandboxes create successfully and pods move to Running.
Use this when
- Pod events show FailedCreatePodSandBox and pods never start
- Fresh nodes that never ran a pod successfully
- Right after a CNI or container runtime upgrade
Not for this skill when
- The error is ImagePullBackOff (the image is the problem, not the sandbox)
- The error is CrashLoopBackOff (the app container starts, then dies)
- Pods run but cannot reach each other (that is network policy or DNS, later in the lifecycle)
Variant phrasings
- failed to create pod sandbox
- cni failed to set up network for sandbox
- pod sandbox creation failed kubernetes
- FailedCreatePodSandBox networkPlugin cni failed
Why it happens
Every pod starts with a pause container that holds the pod's network namespace. The kubelet calls the CNI plugin to set up networking for that sandbox before any app container starts. If the plugin errors (bad config, missing binary, unpullable pause image), the sandbox never exists and the kubelet retries the whole pod creation in a loop.
Edge cases
- If multiple CNI config files exist, only the alphabetically first one is used. Stale files shadow the real config.
- A CNI config written for a newer plugin version than the installed binary fails in confusing ways.
- A full node disk breaks CNI writes and surfaces as sandbox failures. Check disk before going deep on config.
- After a containerd upgrade the CNI config path can change. Re-verify paths, do not assume.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_c5bY7YkfYuBGr6beEsvZUw
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.