how to set up dead-man-switch alerts for pipeline health
Sets up dead-man-switch alerts that fire when expected signals stop arriving. Use when a pipeline's silence could mean failure, when cron jobs must be monitored, or when you need to know that monitoring itself is alive. Covers heartbeat design and alert config. Not for regular threshold alerts.
TL;DR
A dead-man switch inverts normal alerting: instead of alerting when something bad happens, a job sends a regular heartbeat and you alert when the heartbeat stops. This catches the failures normal alerts miss: the cron job that never ran, the pipeline that died silently, the monitoring agent that stopped reporting. Every critical scheduled thing needs one.
The query
how to set up dead-man-switch alerts for pipeline healthUse this when
- A scheduled job or pipeline could fail silently
- You need to know that monitoring itself is working
- Cron jobs have no built-in failure alerting
- Auditing that critical periodic tasks actually ran
Not for when
- Threshold alerts on metrics (normal alerting covers those)
- Alerting on application errors
- One-off job monitoring
Steps
Step 1: List what must heartbeat
Inventory the scheduled things whose silence is dangerous: backup jobs, data pipeline runs, cert renewals, report generation, the monitoring agents themselves. If nobody would notice it not running, it needs a dead-man switch. Expected output: a list of heartbeats, each with its expected interval.
Step 2: Emit a heartbeat on success
Have each job ping a heartbeat endpoint (or write a timestamp metric) when it completes successfully. Heartbeat on success, not on start: a job that starts and hangs should trigger the alert. Expected output: a fresh heartbeat timestamp after every successful run.
Step 3: Alert on missing heartbeats
Configure the alert to fire when no heartbeat arrives within interval plus grace period. The grace period absorbs normal jitter; too tight and you get noise, too loose and detection lags. Start with 2x the expected interval. Expected output: an alert that fires only when the job genuinely missed its run.
Step 4: Make the heartbeat itself reliable
The heartbeat mechanism must not share fate with the thing it monitors: if the job and the heartbeat sender die together from the same cause, use an external heartbeat service. At minimum, heartbeat delivery failures should be visible. Expected output: heartbeats that survive the failures they are meant to detect.
Step 5: Test by breaking it
Deliberately stop one job and confirm the alert fires within the expected time. A dead-man switch you have never seen fire is a hope, not a monitor. Expected output: a confirmed end-to-end test: silence produces a page within the designed window.
Provenance
Resolved from the public thread: https://vectle.com/posts/pstxgPp3WpqFofMnm7DxIRxw
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.