service ClusterIP unreachable from pods: kube-proxy debugging
Debugs ClusterIP services that pods cannot reach. Use when curl from a pod to a service IP times out, when DNS resolves but connections fail, or when kube-proxy might be misconfigured. Not for ingress, NodePort, or external connectivity issues.
TL;DR
A ClusterIP is not a real address; it is a virtual IP that kube-proxy programs into iptables or IPVS rules on every node. When pods cannot reach a service, the usual suspects are kube-proxy not running on the node, stale or missing rules, or network policies blocking the traffic. Debug node by node: is kube-proxy healthy there, do the rules exist, is something else blocking.
The query
service ClusterIP unreachable from pods: kube-proxy debuggingUse this when
- Connections from a pod to a ClusterIP time out or refuse
- DNS resolves the service name but TCP fails
- Some nodes can reach the service but others cannot
- After CNI or kube-proxy configuration changes
Not for when
- Ingress or NodePort access problems
- External traffic into the cluster
- DNS resolution failures (fix DNS first)
Steps
Step 1: Confirm kube-proxy is running on the affected node
Check the kube-proxy pods (usually a DaemonSet) and specifically the instance on the node where the failing pod runs. A dead or crashlooping kube-proxy on one node explains node-specific failures exactly. Expected output: kube-proxy healthy on all nodes, or the broken node identified.
Step 2: Check that the service has endpoints
A ClusterIP with no endpoints blackholes traffic. Verify the service's endpoints are populated and the backing pods are ready. This is the most common cause overall and takes ten seconds to check. Expected output: endpoints present and ready, or confirmation the backend is missing.
Step 3: Inspect the programmed rules on the node
Look at the iptables or IPVS rules for the service's ClusterIP on the affected node. Missing rules mean kube-proxy is not programming them (check its logs); stale rules pointing at dead pods mean endpoint updates are not propagating. Expected output: rules present and pointing at live endpoints, or the specific programming failure identified.
Step 4: Rule out network policies
List network policies in the involved namespaces. A policy that selects the client pods or the service pods can silently drop the traffic, and the symptoms look exactly like a kube-proxy problem. Test by checking policy selectors against the pod labels. Expected output: no policy blocking the path, or the blocking policy found and fixed.
Step 5: Test from multiple pods on multiple nodes
Run the connectivity test from pods on different nodes to map the blast radius: all nodes, one node, or one pod. All-nodes failure points at the service or cluster-wide config; single-node points at that node's kube-proxy or CNI. Expected output: a clear blast-radius map directing the fix at the right layer.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_xkuXNNH33TPsHwIHsge-mQ
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.