## TL;DR
`overloaded_error` means Anthropic's serving capacity is saturated right now; it is transient and the correct response is retry with exponential backoff, not a code fix. Implement retries with jitter, shed non-critical load, and check the status page to distinguish an incident from normal contention.

```text
anthropic overloaded_error: Overloaded
```

## Use this when
- Claude API calls fail with overloaded_error
- An agent pipeline needs to survive capacity blips
- Errors cluster in time then clear on their own

## Not for this skill when
- The key is invalid (thats auth)
- The model name is wrong (thats not-found)
- Prompts exceed context (thats a 400, not overload)

## Steps

1. Confirm it is transient, not your bug:

```text
Check https://status.anthropic.com for an active incident
```
Expected output: either an incident (wait it out) or all-clear (your retries will succeed). Overload is never fixed by changing your prompt.

2. Add exponential backoff with jitter around the call:

```python
import time, random
for attempt in range(5):
    try:
        return call_claude()
    except OverloadedError:
        time.sleep((2 ** attempt) + random.random())
raise RuntimeError("still overloaded after retries")
```
Expected output: most blips clear within a few retries. Jitter keeps your retries from stampeding with everyone else's.

3. Reduce concurrent load during the incident:

```python
# lower worker count / add a semaphore around API calls
```
Expected output: fewer in-flight requests means fewer overload errors. An agent fleet hammering the API during an incident makes its own problem worse.

4. For pipelines that cant wait, queue and defer:

```python
# push failed items to a retry queue with a delay instead of failing the run
```
Expected output: the pipeline completes late rather than failing. Batch and offline workloads should treat overload as "try later," not "error."

## Variant phrasings

### overloaded only at certain hours
Peak contention. Shift batch work off-peak or accept retries as normal during those windows.

### 529 vs overloaded_error
Same family: 529 is the HTTP status some clients surface for the same condition. Same handling.

## Why it happens
LLM serving has finite GPU capacity; when demand spikes (new model launch, big incident reroutes, your own fleet scaling up), the API sheds load with overloaded_error instead of queueing forever. It is a backpressure signal, and the API contract is "retry later." Treating it as a fatal error is the bug in your code, not theirs.

## Edge cases
- Dont retry instantly in a tight loop; that is a self-inflicted DDoS and prolongs the incident for you.
- Long-running agents should distinguish overload (retry) from auth/billing errors (stop immediately).
- If overload persists for hours, check whether your account or tier changed; sustained overload can also reflect reduced quota.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_dEh0io_imG0G5Ply4Ts0bQ
