airflow task retry exponential backoff config
Configures Airflow task retries with exponential backoff. Use it when a data agent or operator needs retries that space out over time instead of hammering a flaky dependency on a fixed interval. Covers retry_exponential_backoff, retry_delay, and max_retry_delay in default_args. Not for fixing the underlying flaky dependency, and not for sensor poke-interval tuning.
TL;DR
Airflow's default retry behavior uses a flat retry_delay, so every retry waits the same amount of time. That is fine for quick blips but brutal for a dependency that is down for an hour: you just hammer it on a fixed schedule. Turn on retry_exponential_backoff=True so the wait doubles (roughly) between attempts, and set max_retry_delay so it stops growing past something sane.
The query
airflow task retry exponential backoff configUse this when
- a task fails on transient errors (API rate limits, brief DB blips, flaky external services)
- you want retries that back off instead of retrying every 5 minutes for 2 hours
- you are writing default_args for a DAG and want a sensible retry policy from day one
Not for
- fixing a task that fails every single time (that is a code bug, not a retry problem)
- tuning sensors (use
poke_intervalandtimeoutfor those instead) - Airflow 1.x (the exponential flag exists but behaves slightly differently there)
Steps
- Set a base
retry_delayandretriesindefault_args. A 5-minute base with 3 retries is a reasonable starting point.
default_args = {
"retries": 3,
"retry_delay": timedelta(minutes=5),
"retry_exponential_backoff": True,
"max_retry_delay": timedelta(hours=1),
}Expected output: DAG parses with no import errors; the args show up in the task's rendered config.
- Enable
retry_exponential_backoff=True. Each retry waits roughlyretry_delay * 2^(try_number - 1)with some jitter added by the scheduler.
Expected output: first retry after ~5 min, second after ~10 min, third after ~20 min.
- Cap the growth with
max_retry_delay. Without it, a task with many retries can end up waiting many hours between attempts.
Expected output: retry intervals never exceed the cap you set, visible in the task instance history.
- Override per task when needed. A task hitting a rate-limited API might want a longer base delay than the DAG default; pass the same keys in the operator's constructor.
Expected output: the task's try timings differ from the DAG default, everything else follows it.
- Verify with a deliberately failing task or by checking the logs of a real one. Look at the timestamps of consecutive tries in the task instance details.
Expected output: the gap between try 1 and try 2 is clearly larger than between attempt 0 and try 1, and it plateaus at your cap.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_cvdKwBrGIMcmE1oRlKRc5A
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.