## TL;DR

Context deadline exceeded means the API call took longer than the client's timeout, and the cause is usually an overloaded API server or slow etcd, not your command. Check API server latency and etcd health first, then look at whats hammering the API (a runaway controller or a huge list call). Fix the load and the timeouts stop.

## The error

```text
Error from server: context deadline exceeded
Unable to connect to the server: context deadline exceeded
```

Controllers log: `failed to list *v1.Pod: context deadline exceeded`.

## Use this when

- kubectl commands time out intermittently
- controllers or operators log deadline exceeded
- API feels slow but isnt down
- list or watch calls hang

## Not for

- connection refused (API server not listening)
- 401/403 errors (auth, not timeouts)
- timeouts from inside pods to services (different path)

## Steps

1. Check if its you or the server. Time a trivial call:

```bash
time kubectl get --raw /healthz
```

Expected: responds in well under a second. If even /healthz is slow, the API server itself is struggling; if its fast, the problem is specific expensive calls.

2. Look at API server and etcd health:

```bash
kubectl get --raw /healthz/etcd
kubectl top pods -n kube-system -l component=kube-apiserver
```

Expected: etcd healthy, apiserver CPU sane. etcd slowness is the classic root cause: every API write waits on etcd quorum.

3. Find whos hammering the API:

```bash
kubectl get --raw /metrics | grep -i apiserver_request_total | sort -t ' ' -k2 -rn | head -10
```

Expected: the top request sources. A controller listing all pods cluster-wide every 30 seconds, or a broken operator in a hot loop, shows up immediately.

4. Check for expensive list calls without limits:

```bash
kubectl get pods -A --chunk-size=500 --request-timeout=30s | head
```

Expected: works fine with chunking. Code that lists everything at once (no pagination, no field selectors) times out as the cluster grows; fix the client to page and filter.

5. Raise the client timeout as a stopgap, fix the server as the cure:

```bash
kubectl get pods -A --request-timeout=60s
```

Expected: the call completes. But dont stop here: sustained deadline exceeded means the control plane needs help (more apiserver CPU, etcd on faster disks, fewer abusive clients).

6. Confirm recovery:

```bash
time kubectl get pods -A --request-timeout=30s
```

Expected: completes in seconds, no deadline errors. Watch for a full day; timeouts that return point at a recurring load pattern like a cronjob that lists everything.

### Variant: only watch streams stall

Long watches timing out is often a proxy or load balancer idle timeout killing the connection, not the API server. Check LB timeouts in front of the API.

### Variant: deadline exceeded right after a cluster upgrade

New admission webhooks or aggregated API servers added latency to every call. Check webhook timeouts and the health of extension API servers.

### Variant: one controller logs it constantly, everything else fine

That controller's client is misconfigured (too-short timeout, no retries, huge list scope). Fix the controller, not the cluster.

## Why it happens

Every API call carries a deadline. When etcd is slow, the apiserver is CPU-starved, or a client makes enormous unpaginated list calls, the work exceeds the deadline and the client gives up. The error blames the context, but the fault is almost always server-side load or client-side greed.

## Edge cases

- etcd on slow disks is the silent killer; watch etcd disk wal fsync durations, not just API latency.
- Huge CRDs with thousands of objects make list calls expensive; paginate and use label selectors.
- Mutating webhooks that call out to slow services add their latency to every matching API call.
- Cloud load balancers in front of the API server have their own idle timeouts that masquerade as deadline exceeded on watches.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_L4wpKrRcfYW5s8dKND9GlQ
