## TL;DR
In-cluster DNS failures come in flavors: CoreDNS pods down or overloaded, the client's ndots and search domains mangling the query, or network policies blocking DNS traffic (UDP/TCP 53 to kube-dns). Test with the FQDN first to split name-construction problems from server problems, then check whether DNS packets even reach CoreDNS.

## The query
```text
CoreDNS "no such host" for in-cluster service names
```

## Use this when
- Pods fail to resolve service names
- Short names fail but FQDN works
- DNS works from some pods but not others
- After CoreDNS config changes

## Not for when
- External hostname resolution (upstream DNS, different path)
- Service has no endpoints (DNS resolves, connections fail)
- Node-level DNS issues

## Steps

### Step 1: Test with the FQDN
Query the full service FQDN (service.namespace.svc.cluster.local) from the failing pod. If the FQDN resolves but the short name does not, the problem is search domains or ndots in the pod's resolv.conf, not CoreDNS.
Expected output: FQDN works (client config issue) or fails too (server/connectivity issue).

### Step 2: Check CoreDNS pod health and load
Verify CoreDNS pods are running and not overloaded: check their CPU, query rates, and logs for errors. Overloaded CoreDNS drops queries, which clients experience as resolution failures.
Expected output: CoreDNS healthy with headroom, or overload identified.

### Step 3: Verify DNS traffic can reach CoreDNS
Check network policies and firewalls for rules blocking UDP/TCP port 53 to the kube-dns service. DNS uses UDP first and falls back to TCP; blocking either breaks resolution in confusing ways.
Expected output: DNS packets flowing to CoreDNS confirmed, or the blocking policy found.

### Step 4: Inspect the pod's resolv.conf
Read the pod's DNS configuration: nameserver should point at kube-dns, ndots and search domains should be the cluster defaults unless deliberately changed. Custom DNS configs are a frequent culprit.
Expected output: sane DNS config, or the misconfiguration spotted.

### Step 5: Query CoreDNS directly to isolate
Run a DNS query from a debug pod straight at a CoreDNS pod IP, bypassing the service. If direct queries work but service queries fail, the kube-dns service or its endpoints are the problem.
Expected output: CoreDNS itself vindicated or implicated, precisely.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_gtqLpIZIyfCHooilErcKGA
