## TL;DR
Most `context deadline exceeded` errors from kubectl or controllers are client-side timeouts firing before a slow API server answers. Raise the client timeout, add retries with backoff in controllers, and check whether client-side rate limiting is throttling you before blaming the API server. Tune the client to the server's real latency, not to defaults meant for fast local clusters.

## Error / query
```text
how to tune Kubernetes API client timeouts and retries
```

## Use this skill when
- kubectl or controllers report `context deadline exceeded`
- API calls flap or time out under load
- Automation is slow and client-side throttling is suspected
- Controllers log rate-limiter waits

## Not for this skill when
- The API server itself is slow or overloaded (server-side tuning)
- etcd is the bottleneck (different diagnosis)
- Requests are denied by RBAC (authorization, not timeouts)

## Steps

### Step 1: Measure whether the server or the client is slow
```bash
time kubectl get pods -n [namespace] --request-timeout=60s
```
Expected: the wall time tells you the server's real latency. If `--request-timeout=60s` succeeds where the default (30s) failed, the client timeout was the problem, not the server.

### Step 2: Set a sane request timeout for interactive use
```bash
kubectl get pods -n [namespace] --request-timeout=2m
```
Expected: long list/watch calls complete instead of timing out. For scripts, export it per-command rather than relying on the default; large list calls against big clusters legitimately take over 30s.

### Step 3: Tune client-go rate limiting in controllers
```go
config.QPS = 50
config.Burst = 100
```
Expected: in controller code, raising QPS/Burst from the client-go defaults (5/10) stops the client from self-throttling. Watch controller logs for `Waited for` messages from the rate limiter; their disappearance confirms the fix.

### Step 4: Add retries with backoff around API calls
```bash
for i in 1 2 3 4 5; do kubectl get [resource] [name] -n [namespace] && break || sleep $((2 ** i)); done
```
Expected: transient timeouts succeed on retry with exponential backoff. For controllers, use the client-go retry utilities rather than failing the reconcile loop on the first timeout.

## Variant phrasings

### "kubectl context deadline exceeded"
Step 1 first: raise `--request-timeout` and see if the call succeeds. If it does, it was the client.

### "kubernetes client throttling"
Check controller logs for rate-limiter waits, then raise QPS/Burst as in step 3.

## Why it happens
Client defaults assume a fast, nearby API server: 30s request timeouts and aggressive client-side rate limits. Against large, loaded, or distant clusters, legitimate calls exceed those defaults, and the client reports a timeout that looks like a server failure. The fix is matching client patience to server reality.

## Edge cases and pitfalls
- Raising timeouts masks a genuinely sick API server; if p99 latency keeps climbing, tune the server instead of the client.
- Watch calls with no timeout can hang forever on network partitions; always bound watches, even generous ones.
- Very high QPS from many controllers can DDoS a small API server; raise limits gradually and watch API server latency.
- `--request-timeout=0` means no timeout, which is dangerous in scripts; always set an explicit bound.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_jf7Qv9jBV9iWwtT-gg-2fA
