## TL;DR
catchup=False means a new DAG starts from now and never runs missed historical intervals: the safe default for most pipelines. A backfill is a deliberate manual run over a date range, created when you actually need history. The classic outage is deploying with catchup=True on a daily DAG with a start_date a year ago, which queues 365 runs at once.

## The query
```text
airflow backfill catchup false vs manual runs
```

## Use this when
- A new DAG just queued hundreds of unexpected historical runs
- Planning how to load history for a new pipeline
- Deciding the catchup setting for a DAG

## Not for
- Tuning sensors or retry delays
- Testing a single run manually (use the UI trigger)
- Pausing a DAG temporarily

## Steps

1. Set catchup=False on nearly every DAG unless you have a specific reason to process history automatically. Pair it with a start_date close to the deployment date.

Expected output: deploying the DAG creates exactly one run for the current interval, not hundreds.

2. If you are already in a run flood, pause the DAG first, then clear or mark the unwanted runs. Do not let the scheduler chew through a year of backfill you never wanted.

Expected output: the run queue drains and only intended runs remain.

3. When history is genuinely needed, do a deliberate backfill: use the CLI backfill command or the UI with an explicit date range, a limited number of runs, and monitoring. Backfills compete with live runs for pool slots, so run them with a capped parallelism.

Expected output: historical intervals fill in a controlled batch without starving the live schedule.

4. Make backfills idempotent. Any DAG that might ever be backfilled must handle re-running an interval safely, usually by writing to date-partitioned targets with overwrite semantics.

Expected output: re-running any interval produces the same data, so backfills are safe to repeat.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_lq6Kb3W4nU6gC-EaU-fujw
