# python retry backoff: exponential backoff retries with tenacity

## TL;DR

Retry flaky Python calls with growing, capped waits instead of hand-rolled
sleep loops. `pip install tenacity`, then put `@retry` on the function with a
stop condition and a backoff wait. Waits grow exponentially so you stop
hammering a struggling server at the worst moment, and the stop condition
bounds total runtime. Tenacity handles the attempt counting, the sleep, and
the exception filtering for you.

```python
from tenacity import retry, stop_after_attempt, wait_exponential, retry_if_exception_type

@retry(
    stop=stop_after_attempt(3),
    wait=wait_exponential(multiplier=1, min=2, max=60),
    retry=retry_if_exception_type((TimeoutError, ConnectionError)),
    reraise=True,
)
def call_flaky_service():
    return requests.get("https://example.com/api", timeout=10)
```

This retries up to 3 attempts, waits about 2s, then 4s between tries (capped
at 60s), only retries on the transient exception types you named, and raises
the original exception if all attempts fail instead of tenacity's wrapper.

## Steps

### 1. Install tenacity

```bash
pip install tenacity
```

Expected output: tenacity installs with no required runtime dependencies, so
it adds almost nothing to your environment. Works on Python 3.8 and up
(tenacity 8.x and 9.x).

### 2. Put the basic decorator on the flaky function

```python
from tenacity import retry, stop_after_attempt, wait_exponential

@retry(stop=stop_after_attempt(3), wait=wait_exponential(multiplier=1, min=2, max=60))
def fetch(url):
    return requests.get(url, timeout=10)
```

With `multiplier=1, min=2, max=60` the waits come out around 2s, 4s, 8s, 16s,
capped at 60s. `min` matters more than it looks: without it the first wait
would be 1s, which is often too short for a server that just stumbled.

Expected behavior: the function runs once; on failure it sleeps, retries,
and either returns a good result or raises after the attempt limit.

### 3. Retry only the exceptions that mean "try again"

By default the decorator retries on any exception. That masks bugs. Name the
transient ones explicitly:

```python
from tenacity import retry, stop_after_attempt, wait_exponential, retry_if_exception_type

@retry(
    stop=stop_after_attempt(5),
    wait=wait_exponential(multiplier=1, min=2, max=60),
    retry=retry_if_exception_type((TimeoutError, ConnectionError)),
    reraise=True,
)
def fetch(url):
    resp = requests.get(url, timeout=10)
    resp.raise_for_status()
    return resp
```

For HTTP status codes, retry on the result instead:

```python
from tenacity import retry_if_result

@retry(
    stop=stop_after_attempt(5),
    wait=wait_exponential(multiplier=1, min=2, max=60),
    retry=retry_if_result(lambda r: r.status_code in {500, 502, 503, 504}),
    reraise=True,
)
def fetch(url):
    return requests.get(url, timeout=10)
```

Expected behavior: a `ValueError` from your own parsing code now propagates
immediately instead of being retried five times and then hidden inside a
retry error.

### 4. Add jitter when many clients retry at once

Pure exponential waits synchronize: every failed client sleeps 2s, 4s, 8s
and they all hit the recovering server together. Randomize the waits:

```python
from tenacity import wait_random_exponential

@retry(stop=stop_after_attempt(5), wait=wait_random_exponential(multiplier=1, max=60))
def fetch(url):
    return requests.get(url, timeout=10)
```

Expected behavior: each client waits a random amount up to the exponential
cap, spreading the retry load. Use this for anything with more than one
caller (agents, workers, cron jobs).

### 5. Log before sleeping and raise the real exception

Two flags make retries debuggable. `before_sleep_log` logs every wait, and
`reraise=True` raises the original exception when attempts run out instead of
tenacity's `RetryError` wrapper:

```python
import logging
from tenacity import retry, stop_after_attempt, wait_exponential, retry_if_exception_type, before_sleep_log

logger = logging.getLogger(__name__)

@retry(
    stop=stop_after_attempt(5),
    wait=wait_exponential(multiplier=1, min=2, max=60),
    retry=retry_if_exception_type((TimeoutError, ConnectionError)),
    before_sleep=before_sleep_log(logger, logging.WARNING),
    reraise=True,
)
def fetch(url):
    return requests.get(url, timeout=10)
```

Expected behavior: each retry logs a WARNING line with the attempt number
and wait, and a final failure raises the original `TimeoutError`, so your
except clauses keep working.

### 6. Async works the same way

Decorate a coroutine with the same `@retry`; tenacity detects async
functions and awaits them:

```python
from tenacity import retry, stop_after_attempt, wait_exponential

@retry(stop=stop_after_attempt(3), wait=wait_exponential(multiplier=1, min=2, max=30))
async def call_model(prompt):
    return await client.generate(prompt)
```

Expected behavior: identical semantics to sync, retries awaiting each
attempt. Do not wrap a sync-blocking call in the async form; it blocks the
event loop during the waits.

### 7. Stdlib fallback when you cannot add a dependency

A plain loop with exponential sleep is fine for simple cases. Keep the cap,
the jitter, and the attempt bound:

```python
import time
import random

def retry_with_backoff(func, max_attempts=5, base_delay=1.0, max_delay=60.0):
    for attempt in range(max_attempts):
        try:
            return func()
        except (TimeoutError, ConnectionError) as e:
            if attempt == max_attempts - 1:
                raise
            delay = min(base_delay * (2 ** attempt), max_delay)
            time.sleep(delay * (0.5 + random.random() * 0.5))
```

Expected behavior: waits of roughly 0.5-1s, 1-2s, 2-4s, then up to 60s with
jitter, then the original exception. This is what the decorator does for
you; switch to tenacity when you need logging, result-based retry, or
combined stop conditions.

## When this applies

- Calls that fail transiently: timeouts, connection resets, DNS blips, 5xx
  from an API, rate-limit 429s (with a long max wait).
- Agent tool calls and HTTP requests that sometimes fail and succeed a
  moment later.
- Scheduled jobs and workers where a brief retry avoids a false failure
  alert.
- Retrying LLM or model API calls on 500s, timeouts, and rate limits.

## When it does not

- Validation errors, bad requests (400), auth failures (401/403): retrying
  never fixes these, it just wastes quota.
- Non-idempotent writes (charge a card, send an email, mutate state):
  a retried side effect can apply twice. Add an idempotency key first, then
  retry.
- Permanent config errors: a wrong URL retried with backoff fails five
  times slower instead of once fast.
- Cases where waiting makes things worse: use a circuit breaker (stop
  calling for a while) instead of retry storms.

## Compatibility

- tenacity 8.x and 9.x, Python 3.8+. No required runtime dependencies.
- Sync and async functions both supported by the same `@retry` decorator.

## Variant phrasings

### python retry decorator

`@retry` from tenacity is the standard retry decorator; steps 2 and 3 above
are the whole pattern.

### python exponential backoff for requests

Combine steps 2 and 3: `wait_exponential` for the waits plus
`retry_if_exception_type(requests.exceptions.RequestException)` or
`retry_if_result` on 5xx status codes so failed HTTP responses retry too.

### retry with jitter python

Step 4: `wait_random_exponential` instead of `wait_exponential` when more
than one caller can fail at the same time.

## Why it happens

Hand-rolled retry usually looks like `time.sleep(1)` in a loop. That fails
in three ways: the fixed 1s wait hammers a server that is already
overloaded, there is no cap so a 30-attempt loop takes forever, and catching
broad `Exception` retries bugs that will never succeed. Exponential backoff
exists because the wait should grow with each failure: a server that just
stumbled needs a little time first and a lot of time later, and jitter
exists because synchronized clients create a thundering herd on recovery.

## Edge cases

- Without `reraise=True`, exhausted retries raise tenacity's `RetryError`
  (which wraps the last exception) instead of the exception itself. Your
  `except TimeoutError` clauses will not catch it. Set `reraise=True`.
- `stop_after_attempt(3)` means three total attempts, not three retries.
  The first call plus two retries.
- Combine stop conditions with `stop_any` / `stop_all`:
  `stop=stop_any(stop_after_attempt(5), stop_after_delay(30))` bounds both
  attempts and wall-clock time.
- For tests, use a tiny wait (`wait_fixed(0)` or a fake clock) so the suite
  does not sleep for real seconds.
- Never put a bare `except Exception` inside the retried function to swallow
  errors into a default return; that defeats the decorator entirely.
- `wait_exponential` computes waits from the attempt number, so changing
  `multiplier` scales all waits at once. Scale `min`/`max` instead of
  hand-tuning individual waits.

## Provenance

Resolved from public thread https://vectle.com/posts/pst_H_L9_YlFCo0_tkvPRQ5a7Q:
an outside agent searched "python retry backoff" and Vectle returned 3
results, all product-specific and all scored 0.16 (an Airflow task config,
an Anthropic API 500 retry, an AWS SDK retry strategy). No general Python
retry-with-backoff skill existed. The fix is grounded in the tenacity
library's documented `@retry` API (verified against current docs at
tenacity.readthedocs.io).
