## TL;DR
Airflow's default retry behavior uses a flat `retry_delay`, so every retry waits the same amount of time. That is fine for quick blips but brutal for a dependency that is down for an hour: you just hammer it on a fixed schedule. Turn on `retry_exponential_backoff=True` so the wait doubles (roughly) between attempts, and set `max_retry_delay` so it stops growing past something sane.

## The query
```
airflow task retry exponential backoff config
```

## Use this when
- a task fails on transient errors (API rate limits, brief DB blips, flaky external services)
- you want retries that back off instead of retrying every 5 minutes for 2 hours
- you are writing default_args for a DAG and want a sensible retry policy from day one

## Not for
- fixing a task that fails every single time (that is a code bug, not a retry problem)
- tuning sensors (use `poke_interval` and `timeout` for those instead)
- Airflow 1.x (the exponential flag exists but behaves slightly differently there)

## Steps
1. Set a base `retry_delay` and `retries` in `default_args`. A 5-minute base with 3 retries is a reasonable starting point.
```python
default_args = {
    "retries": 3,
    "retry_delay": timedelta(minutes=5),
    "retry_exponential_backoff": True,
    "max_retry_delay": timedelta(hours=1),
}
```
Expected output: DAG parses with no import errors; the args show up in the task's rendered config.

2. Enable `retry_exponential_backoff=True`. Each retry waits roughly `retry_delay * 2^(try_number - 1)` with some jitter added by the scheduler.
Expected output: first retry after ~5 min, second after ~10 min, third after ~20 min.

3. Cap the growth with `max_retry_delay`. Without it, a task with many retries can end up waiting many hours between attempts.
Expected output: retry intervals never exceed the cap you set, visible in the task instance history.

4. Override per task when needed. A task hitting a rate-limited API might want a longer base delay than the DAG default; pass the same keys in the operator's constructor.
Expected output: the task's try timings differ from the DAG default, everything else follows it.

5. Verify with a deliberately failing task or by checking the logs of a real one. Look at the timestamps of consecutive tries in the task instance details.
Expected output: the gap between try 1 and try 2 is clearly larger than between attempt 0 and try 1, and it plateaus at your cap.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_cvdKwBrGIMcmE1oRlKRc5A
