## TL;DR
Intermittent DNS failures are almost never the DNS records; they are the path: overloaded resolvers, UDP packet loss, search-domain misconfiguration, or app-level caching of failures. Debug by reproducing the failure mode (not just testing when it works), capturing whether it is timeout vs NXDOMAIN vs SERVFAIL, and moving up the chain from the app to the resolver. Each failure type points at a different layer.

## The query
```text
how to debug intermittent DNS resolution failures in production
```

## Use this when
- Services fail to resolve hostnames intermittently
- Restarts temporarily fix resolution
- Only some hosts or pods are affected
- Errors cluster in time (suggesting resolver overload)

## Not for when
- Managing DNS records and zones
- Permanent resolution failures (check the records)
- Service discovery inside orchestrators (different mechanism)

## Steps

### Step 1: Capture the exact failure mode
Next time it happens, record whether it is a timeout, NXDOMAIN, or SERVFAIL, and for which hostname. Timeouts point at network or resolver overload; NXDOMAIN points at search domains or bad names; SERVFAIL points at upstream resolver problems. Guessing without this wastes the whole investigation.
Expected output: the failure classified, which determines the entire debug path.

### Step 2: Test resolution at each layer
From the affected host, query the configured resolver directly, then query upstream resolvers, then test with search domains disabled. This isolates whether the problem is the local resolver config, the upstream, or name construction.
Expected output: the layer where resolution breaks identified.

### Step 3: Check resolver health and load
Look at the resolver's query volume, cache hit rate, and response times during failure windows. Overloaded resolvers drop UDP queries first, which looks exactly like intermittent failure from the client. Corporate DNS forwarders are a common bottleneck.
Expected output: resolver metrics showing overload (or ruling it out).

### Step 4: Look for client-side caching of failures
Many runtimes and libraries cache DNS failures (negative caching) far longer than successes. A 30-second resolver blip becomes a 5-minute outage if the app caches the failure. Check the runtime's DNS cache behavior and tune negative TTLs down.
Expected output: negative caching identified and tamed; blips stop becoming outages.

### Step 5: Add redundancy and monitoring
Run local caching resolvers on hosts, configure multiple upstream resolvers, and monitor resolution latency and failure rate as a first-class metric. DNS is infrastructure; monitor it like infrastructure.
Expected output: resolution failures become rare, and when they happen, the monitoring names the layer immediately.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_8ecl-BPgr7DwjHBSiqIGAw
