kubectl exec gives i/o timeout: how to debug
Debugs i/o timeout errors from kubectl exec. Use when exec into a pod fails with a dial tcp i/o timeout while other kubectl commands still work. Covers the API-server-to-kubelet path, node health, kubelet port reachability, and CNI problems. Not for exec permission errors, exec into crashed containers, cluster-wide API outages, or slow commands that time out for their own reasons.
TL;DR
kubectl exec failing with i/o timeout while kubectl get and kubectl logs work means the API server cannot reach the kubelet on that node in time; the pod itself is usually fine. Check node health, confirm TCP port 10250 is reachable from the control plane, and look at the CNI pods. The fix is almost always network between control plane and node, not the workload.
Error / query
kubectl exec gives i/o timeout: how to debugSteps
Step 1: Confirm the API server is healthy and scope the failure
kubectl get pods -n [namespace]
kubectl logs [pod-name] -n [namespace] --tail=20Expected: both work. That proves your credentials and the API server are fine and isolates the problem to the exec streaming path.
Step 2: Check which node the pod runs on and whether it is healthy
kubectl get pod [pod-name] -n [namespace] -o wide
kubectl describe node [node-name]Expected: the pod's node is listed, and the node shows Ready under Conditions. A NotReady node explains the timeout immediately.
Step 3: Test whether the kubelet port is reachable
From a machine with network access to the nodes, run:
nc -zv [node-ip] 10250Expected: succeeded or open. The exec path is client to API server to kubelet on TCP 10250; a refused or timed-out connection here is the timeout you are seeing. On managed clusters this usually means a security group or firewall rule between control plane and nodes.
Step 4: Check the CNI and kubelet on the node
kubectl get pods -n kube-systemExpected: the CNI pods (calico, cilium, flannel, or your cluster's equivalent) are all Running. If they are crash-looping, pod networking on that node is broken and exec cannot set up its stream. On the node itself, journalctl -u kubelet --since "30 min ago" shows whether the kubelet is healthy or restarting.
Step 5: Try exec variants to rule out the container side
kubectl exec [pod-name] -n [namespace] -- [command]
kubectl exec [pod-name] -n [namespace] -c [container-name] -- [command]Expected: if the plain form times out but everything above is healthy, try naming the container explicitly in multi-container pods. A container whose main process hangs on startup can also stall exec session setup; check the container is actually Running first.
Step 6: Fix the network path, then re-test
Open TCP 10250 from the control plane security group to the nodes, fix the broken CNI pods, or resolve the node pressure, then re-run the exec. Expected: exec connects within a couple of seconds.
When to use
kubectl execreturnsdial tcp ...: i/o timeout(often mentioning port 10250)- exec fails but
kubectl get,describe, andlogsall work - exec works on some nodes but times out on others
When not to use
- exec fails with
forbiddenorunauthorized; that is RBAC, not networking - The container is crashed or has no shell; exec cant attach to a dead container
- The whole API server is unreachable; fix cluster connectivity first
- A command inside the pod is just slow; that is the workload, not an i/o timeout on dial
Tool compatibility
kubectl 1.27 and newer against any Kubernetes cluster; kubelet TCP port 10250 on every node; managed clusters (EKS, GKE, AKS) where control-plane-to-node traffic crosses security groups or firewall rules; self-hosted clusters with host firewalls like firewalld or ufw; any CNI.
Variant phrasings
"kubectl exec dial tcp 10250 i/o timeout"
The exact signature of this problem. The API server dialed the kubelet and got no answer; start at step 3.
"kubectl exec timeout only on one node"
Node-specific: compare describe node output between a working and a broken node, and check whether 10250 is reachable on each. Usually a per-node firewall rule or a dead kubelet.
"kubectl exec works but is very slow to start"
Same path, marginal network. It has not timed out yet, but the control-plane-to-node link is degraded; investigate before it becomes a full timeout.
Why it happens
kubectl exec is not a direct connection to the pod. Your client asks the API server, which opens a second connection to the kubelet on the pod's node over TCP port 10250, and the kubelet attaches to the container runtime. kubectl logs and get only need the API server, which is why they keep working. Anything breaking the API-server-to-kubelet leg (firewall, dead kubelet, broken CNI, NotReady node) surfaces as an i/o timeout on dial.
Edge cases
- Cloud managed clusters sometimes restrict 10250 to specific control plane security groups; node OS firewalls are a second place the port gets blocked.
- A kubelet under heavy disk or memory pressure can accept the TCP connection but never respond; node pressure looks like a network timeout.
kubectl execthrough a corporate VPN or proxy can add its own timeouts; test from inside the cluster network to separate client-side from server-side causes.- A hung container runtime on the node produces the same symptom; restarting the container runtime on the node (not the kubelet) is the fix there.
Provenance
Resolved from the public thread: https://vectle.com/posts/pstm5x0rTTw8saEvUVfn3sOw
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.