## TL;DR
Rotate weekly among at least 4 people, define exactly what pages (urgent access, outages, security) vs what waits until morning, and compensate the time. Publish the schedule where everyone can see it. On-call fails on ambiguity about what deserves a page.

## The error
```text
(Process task; no error.)
```

## Steps
1. Define pageable events narrowly: production outages, security incidents, urgent executive issues. Expected: list. Everything else waits until morning; write that down.
2. Set the rotation: weekly shifts, minimum 4 people, no back-to-back weeks. Expected: schedule published. Two-person rotations burn out in months.
3. Build the escalation: primary has 15 minutes, then secondary, then the manager. Expected: documented. Pages that nobody answers are worse than no on-call.
4. Compensate: pay, time off, or both, per local law and policy. Expected: defined. Uncompensated on-call is a retention problem.
5. Review quarterly: page volume, false pages, and participant feedback. Expected: reviewed. Tune the pageable list; most rotations page too much at first.

## When to use
- New on-call setup
- Burnout or complaints about the current rotation

## When not to use
- Follow-the-sun coverage (different model)
- Incident response procedures (separate runbooks)

## Compatibility
- PagerDuty, Opsgenie, or any on-call tool; the principles are universal

## Variants
### Small team
Share with an adjacent team or use a managed service for nights.
### Low urgency environment
A next-business-day model may beat on-call entirely; do not build on-call you do not need.

## Why it happens
On-call exists for the rare truly urgent event. Vague paging criteria turn it into 24/7 availability, which destroys the team it was meant to protect.

## Edge cases
- Holidays need explicit coverage planned in advance, not assumed.
- New hires shadow for a cycle before taking pages solo.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_kjiN3mWziy2TfjWpQ5ooSA
