## TL;DR
The scheduler is supposed to write a heartbeat timestamp to the metadata database every few seconds, and something stopped it from doing that. A heartbeat timeout means the scheduler is overloaded, stuck, or dead, not that your DAGs are broken. Check whether the scheduler process is alive and responsive first, then look at whatever is starving it.

```text
Airflow scheduler heartbeat timeout: the scheduler has not sent a heartbeat in over 30 seconds
```

## Use this when
- the UI warns about scheduler heartbeat timeouts
- DAGs stop scheduling and the scheduler looks wedged
- you are sizing scheduler resources for growth

## Not for this skill when
- tasks fail but scheduling continues, that is a task problem
- the metadata database is down entirely, fix the DB first
- you see worker heartbeat issues, workers are a different component

## Steps

1. Check whether the scheduler process is alive right now:

```shell
ps aux | grep -i "airflow scheduler" | grep -v grep
```

Expected output: a running process with recent CPU time. No process means it died outright, restart it and then investigate why it died.

2. Read the tail of the scheduler log for what the loop was doing when it stalled:

```shell
ls ~/airflow/logs/scheduler/ | tail -3
```

Expected output: recent scheduler log files. Open the newest and look for long garbage collection pauses, database errors, or parsing storms right before the silence.

3. Check the heartbeat configuration to see whether the timeout is even realistic:

```shell
airflow config get-value scheduler scheduler_heartbeat_sec
```

Expected output: the heartbeat interval, 5 seconds by default. If someone set aggressive timeouts, normal load under a busy scheduler can trip them spuriously.

4. Look for DAG parsing overload, which is the most common cause of a starved scheduler:

```shell
airflow dags list-import-errors | head -20
```

Expected output: ideally an empty list. Broken or enormous DAG files make the parser spin, and a spinning parser starves the heartbeat loop of CPU time.

5. Check metadata DB responsiveness, since a slow database delays every heartbeat write:

```sql
SELECT COUNT(*) FROM dag_run WHERE state = 'running';
```

Expected output: a quick count returned promptly. If even this simple query hangs, the database is the bottleneck, not the scheduler process.

6. Restart the scheduler cleanly after addressing the cause, not before:

```shell
pkill -f "airflow scheduler"; sleep 5; airflow scheduler -D
```

Expected output: a fresh scheduler process that starts heartbeating normally. Watch the UI warning clear within a minute of the restart.

## Variant phrasings

### airflow scheduler seems unhealthy heartbeat
Same signal, different wording. The scheduler is alive enough to be watched but not responsive enough to heartbeat, which points at a stalled loop.

### airflow scheduler timeout under load
More DAGs means more parsing and slower heartbeats. Reduce parsing load or add scheduler resources, do not just raise the timeout.

### scheduler heartbeat vs liveness probe
The heartbeat is the scheduler reporting itself healthy to the database. A Kubernetes liveness probe watches the process, the heartbeat watches the loop, and they catch different failures.

## Why it happens
The scheduler runs a tight loop: parse DAGs, create runs, queue tasks, write heartbeat. Anything that stalls the loop, a wedged parser, a locked database, CPU exhaustion, delays the heartbeat write. The timeout is just the monitoring noticing the silence, which is why the fix is always about whatever stalled the loop and never about the heartbeat setting itself.

## Edge cases
- Raising the timeout hides the symptom without fixing the stall. Tune it only after the underlying cause is addressed.
- Multiple schedulers in HA mode each heartbeat separately. One timing out while the others are fine is a single-host problem.
- Clock skew between hosts can fake a timeout. Keep clocks synced or the timestamps lie.
- Very large DAG files with thousands of generated tasks parse slowly by design. Generate fewer tasks or split the DAG.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_pxBYJQ_qkxs_xZ-5odkjbg
