CoreDNS "no such host" for in-cluster service names
Fixes CoreDNS failing to resolve in-cluster service names. Use when pods cannot resolve service DNS, when short names fail but FQDN works, or when DNS works from some pods but not others. Not for external DNS resolution.
TL;DR
In-cluster DNS failures come in flavors: CoreDNS pods down or overloaded, the client's ndots and search domains mangling the query, or network policies blocking DNS traffic (UDP/TCP 53 to kube-dns). Test with the FQDN first to split name-construction problems from server problems, then check whether DNS packets even reach CoreDNS.
The query
CoreDNS "no such host" for in-cluster service namesUse this when
- Pods fail to resolve service names
- Short names fail but FQDN works
- DNS works from some pods but not others
- After CoreDNS config changes
Not for when
- External hostname resolution (upstream DNS, different path)
- Service has no endpoints (DNS resolves, connections fail)
- Node-level DNS issues
Steps
Step 1: Test with the FQDN
Query the full service FQDN (service.namespace.svc.cluster.local) from the failing pod. If the FQDN resolves but the short name does not, the problem is search domains or ndots in the pod's resolv.conf, not CoreDNS. Expected output: FQDN works (client config issue) or fails too (server/connectivity issue).
Step 2: Check CoreDNS pod health and load
Verify CoreDNS pods are running and not overloaded: check their CPU, query rates, and logs for errors. Overloaded CoreDNS drops queries, which clients experience as resolution failures. Expected output: CoreDNS healthy with headroom, or overload identified.
Step 3: Verify DNS traffic can reach CoreDNS
Check network policies and firewalls for rules blocking UDP/TCP port 53 to the kube-dns service. DNS uses UDP first and falls back to TCP; blocking either breaks resolution in confusing ways. Expected output: DNS packets flowing to CoreDNS confirmed, or the blocking policy found.
Step 4: Inspect the pod's resolv.conf
Read the pod's DNS configuration: nameserver should point at kube-dns, ndots and search domains should be the cluster defaults unless deliberately changed. Custom DNS configs are a frequent culprit. Expected output: sane DNS config, or the misconfiguration spotted.
Step 5: Query CoreDNS directly to isolate
Run a DNS query from a debug pod straight at a CoreDNS pod IP, bypassing the service. If direct queries work but service queries fail, the kube-dns service or its endpoints are the problem. Expected output: CoreDNS itself vindicated or implicated, precisely.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_gtqLpIZIyfCHooilErcKGA
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.