## TL;DR
Use the backfill command with explicit start and end dates and let Airflow create the historical runs in date order. Cap max_active_runs so the backfill does not stampede your workers, and watch the first few runs before walking away. Backfills are routine operations, the failures come from rushing them.

```text
how to backfill an Airflow DAG for past dates
```

## Use this when
- a new DAG needs historical runs to populate history
- a bug corrupted past outputs and you must recompute them
- catchup did not cover the range you need

## Not for this skill when
- you need to backfill dbt models, that is a different skill
- only one date is wrong, clear and rerun that single task instead
- the DAG has depends_on_past chains you have not thought through

## Steps

1. Check existing runs first so the backfill covers only dates that are actually missing:

```shell
airflow dags list-runs --dag-id my_dag | head -20
```

Expected output: the runs that already exist with their states. Your backfill range should cover dates with no successful runs, not the whole calendar blindly.

2. Run the backfill with explicit start and end dates:

```shell
airflow dags backfill my_dag --start-date 2024-01-01 --end-date 2024-01-31
```

Expected output: Airflow creates one run per schedule interval in the range and starts executing them in date order.

3. Confirm the DAG's max_active_runs is sane before a large backfill, since the DAG setting governs parallelism:

```python
# in your DAG definition:
with DAG(
    dag_id="my_dag",
    max_active_runs=3,
):
    ...
```

Expected output: at most 3 runs execute at once. A backfill with unlimited parallelism will stampede workers and the warehouse, so set this deliberately.

4. Watch the first few runs before letting the rest of the range go:

```shell
airflow dags list-runs --dag-id my_dag | head -10
```

Expected output: runs progressing from queued toward success. If the first runs fail, fix the cause before the backfill burns through the whole range producing failures.

5. Verify the full range completed successfully with one query against the metadata DB:

```sql
SELECT logical_date, state, COUNT(*)
FROM dag_run
WHERE dag_id = 'my_dag' AND logical_date >= '2024-01-01'
GROUP BY 1, 2 ORDER BY 1;
```

Expected output: one successful run per interval in the range. Rerun any failed dates individually rather than redoing the entire backfill.

## Variant phrasings

### airflow backfill command example
Step 2 is the canonical form. Dag id, start date, end date, that is the whole command, everything else is optional tuning.

### airflow catchup vs backfill
Catchup fills gaps automatically when a DAG with a past start_date gets unpaused. Backfill is the manual command for when you want explicit control over the range.

### airflow rerun failed tasks in a date range
Add --rerun-failed-tasks to the backfill command, or clear the failed tasks for those dates first and then backfill normally.

## Why it happens
Airflow only creates runs for intervals the scheduler processes while the DAG is active. History from before the DAG existed, or intervals lost to downtime, simply have no runs. The backfill command synthesizes those runs in date order, which is exactly what the scheduler would have done had it been running the whole time.

## Edge cases
- depends_on_past=True makes backfills strictly sequential, which is slow but correct. Do not disable it just to go faster.
- Backfills ignore paused DAGs entirely. Unpause first or nothing happens.
- A backfill over hundreds of intervals can take days. Chunk the range and monitor between chunks.
- Rerunning a backfill over already-successful dates is pure waste. Scope the range to what is actually missing.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_qX4ob7y_6fntw8m4RnE5JA
