Airflow scheduler "heartbeat timeout" what it means
Explains what an Airflow scheduler heartbeat timeout means and how to fix it. Use when the UI warns about scheduler heartbeats, when DAGs stop scheduling and the scheduler looks wedged, or when sizing scheduler resources. Not for task failures, dead databases, or worker issues.
TL;DR
The scheduler is supposed to write a heartbeat timestamp to the metadata database every few seconds, and something stopped it from doing that. A heartbeat timeout means the scheduler is overloaded, stuck, or dead, not that your DAGs are broken. Check whether the scheduler process is alive and responsive first, then look at whatever is starving it.
Airflow scheduler heartbeat timeout: the scheduler has not sent a heartbeat in over 30 secondsUse this when
- the UI warns about scheduler heartbeat timeouts
- DAGs stop scheduling and the scheduler looks wedged
- you are sizing scheduler resources for growth
Not for this skill when
- tasks fail but scheduling continues, that is a task problem
- the metadata database is down entirely, fix the DB first
- you see worker heartbeat issues, workers are a different component
Steps
- Check whether the scheduler process is alive right now:
ps aux | grep -i "airflow scheduler" | grep -v grepExpected output: a running process with recent CPU time. No process means it died outright, restart it and then investigate why it died.
- Read the tail of the scheduler log for what the loop was doing when it stalled:
ls ~/airflow/logs/scheduler/ | tail -3Expected output: recent scheduler log files. Open the newest and look for long garbage collection pauses, database errors, or parsing storms right before the silence.
- Check the heartbeat configuration to see whether the timeout is even realistic:
airflow config get-value scheduler scheduler_heartbeat_secExpected output: the heartbeat interval, 5 seconds by default. If someone set aggressive timeouts, normal load under a busy scheduler can trip them spuriously.
- Look for DAG parsing overload, which is the most common cause of a starved scheduler:
airflow dags list-import-errors | head -20Expected output: ideally an empty list. Broken or enormous DAG files make the parser spin, and a spinning parser starves the heartbeat loop of CPU time.
- Check metadata DB responsiveness, since a slow database delays every heartbeat write:
SELECT COUNT(*) FROM dag_run WHERE state = 'running';Expected output: a quick count returned promptly. If even this simple query hangs, the database is the bottleneck, not the scheduler process.
- Restart the scheduler cleanly after addressing the cause, not before:
pkill -f "airflow scheduler"; sleep 5; airflow scheduler -DExpected output: a fresh scheduler process that starts heartbeating normally. Watch the UI warning clear within a minute of the restart.
Variant phrasings
airflow scheduler seems unhealthy heartbeat
Same signal, different wording. The scheduler is alive enough to be watched but not responsive enough to heartbeat, which points at a stalled loop.
airflow scheduler timeout under load
More DAGs means more parsing and slower heartbeats. Reduce parsing load or add scheduler resources, do not just raise the timeout.
scheduler heartbeat vs liveness probe
The heartbeat is the scheduler reporting itself healthy to the database. A Kubernetes liveness probe watches the process, the heartbeat watches the loop, and they catch different failures.
Why it happens
The scheduler runs a tight loop: parse DAGs, create runs, queue tasks, write heartbeat. Anything that stalls the loop, a wedged parser, a locked database, CPU exhaustion, delays the heartbeat write. The timeout is just the monitoring noticing the silence, which is why the fix is always about whatever stalled the loop and never about the heartbeat setting itself.
Edge cases
- Raising the timeout hides the symptom without fixing the stall. Tune it only after the underlying cause is addressed.
- Multiple schedulers in HA mode each heartbeat separately. One timing out while the others are fine is a single-host problem.
- Clock skew between hosts can fake a timeout. Keep clocks synced or the timestamps lie.
- Very large DAG files with thousands of generated tasks parse slowly by design. Generate fewer tasks or split the DAG.
Provenance
Resolved from the public thread: https://vectle.com/posts/pstpxBYJQqkxs_xZ-5odkjbg
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.