## TL;DR
A dead-man switch inverts normal alerting: instead of alerting when something bad happens, a job sends a regular heartbeat and you alert when the heartbeat stops. This catches the failures normal alerts miss: the cron job that never ran, the pipeline that died silently, the monitoring agent that stopped reporting. Every critical scheduled thing needs one.

## The query
```text
how to set up dead-man-switch alerts for pipeline health
```

## Use this when
- A scheduled job or pipeline could fail silently
- You need to know that monitoring itself is working
- Cron jobs have no built-in failure alerting
- Auditing that critical periodic tasks actually ran

## Not for when
- Threshold alerts on metrics (normal alerting covers those)
- Alerting on application errors
- One-off job monitoring

## Steps

### Step 1: List what must heartbeat
Inventory the scheduled things whose silence is dangerous: backup jobs, data pipeline runs, cert renewals, report generation, the monitoring agents themselves. If nobody would notice it not running, it needs a dead-man switch.
Expected output: a list of heartbeats, each with its expected interval.

### Step 2: Emit a heartbeat on success
Have each job ping a heartbeat endpoint (or write a timestamp metric) when it completes successfully. Heartbeat on success, not on start: a job that starts and hangs should trigger the alert.
Expected output: a fresh heartbeat timestamp after every successful run.

### Step 3: Alert on missing heartbeats
Configure the alert to fire when no heartbeat arrives within interval plus grace period. The grace period absorbs normal jitter; too tight and you get noise, too loose and detection lags. Start with 2x the expected interval.
Expected output: an alert that fires only when the job genuinely missed its run.

### Step 4: Make the heartbeat itself reliable
The heartbeat mechanism must not share fate with the thing it monitors: if the job and the heartbeat sender die together from the same cause, use an external heartbeat service. At minimum, heartbeat delivery failures should be visible.
Expected output: heartbeats that survive the failures they are meant to detect.

### Step 5: Test by breaking it
Deliberately stop one job and confirm the alert fires within the expected time. A dead-man switch you have never seen fire is a hope, not a monitor.
Expected output: a confirmed end-to-end test: silence produces a page within the designed window.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_xgPp3WpqFofMnm7D_xIRxw
