## TL;DR

If `myservice` doesnt resolve inside a pod but the service exists, the problem is between the pod and CoreDNS: check the pod's resolv.conf, test CoreDNS directly, and look at the CoreDNS pods themselves. Most cases are CoreDNS down or overloaded, a bad ndots setting, or the service name being wrong. Fix the DNS layer and everything downstream starts working.

## The error

```text
nslookup myservice
Server: CONTAINER_IP
** server can't find myservice: NXDOMAIN
```

App-level: `could not resolve host: myservice`, `getaddrinfo ENOTFOUND`.

## Use this when

- services dont resolve by short name or FQDN inside pods
- nslookup/dig from a pod returns NXDOMAIN or times out
- CoreDNS pods are crashlooping or pending
- DNS works sometimes and fails sometimes (flaky)

## Not for

- external domain names not resolving (upstream DNS, not CoreDNS)
- ingress or external-dns record problems
- connections failing after successful resolution (thats routing or policy)

## Steps

1. Confirm what the pod sees:

```bash
kubectl run dns-test --rm -it --image=busybox --restart=Never -- nslookup myservice
cat /etc/resolv.conf
```

Expected: resolv.conf points at the cluster DNS IP (usually CONTAINER_IP) with `ndots:5`. If the nameserver is wrong, the pod spec overrides DNS; fix dnsPolicy/dnsConfig.

2. Test CoreDNS directly, bypassing the pod:

```bash
kubectl get pods -n kube-system -l k8s-app=kube-dns
kubectl run dns-test --rm -it --image=busybox --restart=Never -- nslookup myservice.default.svc.cluster.local CONTAINER_IP
```

Expected: the FQDN resolves when asked directly. If the short name fails but FQDN works, its an ndots/search-domain issue, not a CoreDNS outage.

3. Check the CoreDNS pods are healthy:

```bash
kubectl get pods -n kube-system -l k8s-app=kube-dns
kubectl logs -n kube-system -l k8s-app=kube-dns --tail=50
```

Expected: pods Running, logs clean. CrashLoopBackOff here usually means a bad Corefile edit; `kubectl get configmap coredns -n kube-system -o yaml` to inspect.

4. Verify the service actually exists and has endpoints:

```bash
kubectl get svc myservice -n [namespace]
kubectl get endpoints myservice -n [namespace]
```

Expected: the service exists with a ClusterIP and endpoints list pod IPs. No endpoints means the selector matches nothing, and DNS correctly returns nothing.

5. Check for the classic ndots:5 slowness with short external names:

```bash
kubectl run dns-test --rm -it --image=busybox --restart=Never -- nslookup google.com
```

Expected: slow but works (tries 5 search domains first). If apps time out on external names, lower ndots via dnsConfig or use FQDNs with a trailing dot.

6. Fix and verify end to end:

```bash
kubectl rollout restart deployment/[app] -n [namespace]
kubectl run dns-test --rm -it --image=busybox --restart=Never -- nslookup myservice
```

Expected: fresh pods resolve the name on the first try, app connections succeed.

### Variant: DNS works, then stops, then works (flaky)

CoreDNS overloaded or conntrack table full on the node. Scale CoreDNS replicas, check `dmesg` for conntrack drops, and consider node-local DNS cache.

### Variant: only one namespace affected

Check that namespace's NetworkPolicies: DNS is UDP/TCP 53 to the kube-dns service, and a default-deny policy without a DNS allow rule breaks exactly this.

### Variant: custom stub domains or rewrites broke everything

Revert the Corefile change and re-apply incrementally. One bad forward or rewrite line poisons the whole cluster's DNS.

## Why it happens

Pods resolve names through CoreDNS via the cluster DNS service IP. The chain has several links: pod resolv.conf, the DNS service, CoreDNS pods, the service registry. Short names go through search-domain expansion (ndots:5), which multiplies queries. Any broken link, or a service that doesnt exist, surfaces as NXDOMAIN or timeouts.

## Edge cases

- Headless services resolve to pod IPs, not a ClusterIP; clients must handle multiple A records.
- `dnsPolicy: Default` makes the pod use the node's DNS, skipping CoreDNS entirely; fine for some jobs, confusing when cluster names stop resolving.
- Alpine/musl images handle DNS differently from glibc; flaky resolution in Alpine often needs `ndots:1` or the full FQDN.
- Very large clusters can exceed CoreDNS default resources; watch its CPU and scale it before it becomes the bottleneck.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_Sw7h3PA5rMCRFakV90JNOg
