container runtime network not ready": CNI plugin troubleshooting
Troubleshoots the container runtime network not ready error. Use when nodes stay NotReady with network errors, when pods cannot start due to CNI, or after CNI upgrades. Not for DNS or policy issues once networking is up.
TL;DR
This error means the kubelet asked the CNI plugin to set up pod networking and the plugin failed or never responded. Check that the CNI pods (DaemonSet) are running on the node, read their logs for the actual failure, and verify the CNI config files exist on the node. Most cases are a crashed CNI agent pod, a bad config push, or version skew after an upgrade.
The query
"container runtime network not ready": CNI plugin troubleshootingUse this when
- Nodes report network not ready and stay NotReady
- Pods fail to start with CNI errors
- After CNI plugin upgrades or config changes
- New nodes never become ready
Not for when
- DNS resolution failures (networking is up, DNS is broken)
- Network policy blocking traffic
- Pod-to-pod connectivity with ready nodes
Steps
Step 1: Check the CNI agent pods on the node
The CNI runs as a DaemonSet; confirm its pod is running on the affected node. A missing or crashlooping CNI pod explains the error completely. Compare with healthy nodes. Expected output: the CNI pod state on the bad node vs good nodes.
Step 2: Read the CNI agent logs
The agent logs name the failure: failure to allocate IPs, API server connectivity loss, config parse errors. The kubelet's message is generic; the CNI log is specific. Expected output: the concrete CNI-side error.
Step 3: Verify CNI config files on the node
Check that the CNI configuration exists in the node's config directory and is valid. Config management pushes that write partial files, or upgrades that change the expected config format, break the plugin silently. Expected output: a valid, complete CNI config present on the node.
Step 4: Check version compatibility
Confirm the CNI plugin version supports the Kubernetes and container runtime versions in use. Upgrading one component without the others is the classic way to get this error cluster-wide. Expected output: a supported version matrix, or the skew identified.
Step 5: Restart the CNI pod and watch recovery
Delete the CNI pod on the node to force a fresh start, then watch the node become ready. If it does not recover, the problem is config or compatibility, not a transient crash. Expected output: node Ready with pod networking functional, or the persistent cause isolated.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_Rk1xL207bdVh6w86R2Ynzg
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.