## TL;DR
The heartbeat is how Airflow proves each component is alive; when it times out, the scheduler, a worker, or the metadata database is the culprit, in that order of likelihood. Check the scheduler logs for heartbeat errors first, then database connection latency. Raising the timeout without finding the cause just delays the same stall.

```text
airflow heartbeat timeout
```

## Use this when
- The Airflow UI shows the scheduler as unhealthy or tasks stop being scheduled.
- Logs mention heartbeat timeouts on the scheduler or workers.
- Tasks sit in queued or scheduled state far longer than usual.

## Not for this skill when
- A single task times out inside its own code. That is a task timeout, not a heartbeat.
- The webserver is slow. Heartbeats are scheduler and executor health, not the UI.

## Steps
1. Check the scheduler logs for heartbeat entries and note the last successful beat versus now. Confirm whether the scheduler process is actually running. Verify: you know if the scheduler died or just went quiet
2. Measure metadata database latency with a simple query from the scheduler host. Note whether queries that should take milliseconds take seconds. Verify: database slowness is confirmed or ruled out
3. Look at scheduler CPU and memory during the gap. Note whether the box was saturated. Verify: resource starvation is confirmed or ruled out
4. If the database is slow, check for long-running queries or lock contention on the Airflow tables. Note the blocking query. Verify: you found the actual bottleneck
5. Only after the cause is fixed, tune scheduler_heartbeat_sec and related timeouts to fit your database latency. Note heartbeats resume steadily. Verify: the timeout matches reality instead of masking it

## Variant phrasings
### airflow scheduler heartbeat timeout
The scheduler-specific phrasing.
### airflow worker heartbeat timed out
The executor-side version of the same failure.

Compatibility: Airflow 2.x, all executors. Heartbeat settings live in airflow.cfg under [scheduler]; exact keys shifted slightly between 2.0 and 2.8.

## Why it happens
Airflow coordinates through the metadata database, and the heartbeat is the liveness check that keeps components from stepping on each other. When the database slows down or the scheduler host runs out of CPU, heartbeats arrive late and Airflow marks components unhealthy. Everything downstream of that, stalled scheduling, zombie tasks, looks like an Airflow bug but is usually a resource problem one layer down.

## Edge cases / pitfalls
- A scheduler that was killed and restarted quickly can show phantom heartbeat timeouts while the old row in the db is still marked alive. Check process age.
- Over-aggressive heartbeat intervals on a slow database make the problem worse: every missed beat triggers recovery work that loads the db more.
- In HA scheduler setups, a partitioned network makes two schedulers both think the other is dead. Fix the network, not the timeout.
- Celery workers that OOM get heartbeat timeouts as a side effect. The OOM is the real error.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_myc7ZZaL-00nSJuslE0NEA
