on-call handoff template for SRE teams
Standardizes on-call rotation handoffs for SRE teams. Use when context is lost between rotations, the outgoing on-call keeps getting pinged, or handoffs need a shared template. Covers active incidents, fragile systems, in-flight changes, and contact maps. Not for escalation policies or permanent knowledge transfers.
TL;DR
A good on-call handoff tells the incoming engineer what is broken, what is fragile, and what is already in flight, in writing, before the rotation changes. Cover: active incidents, known flaky systems, in-progress changes, and who to call for what. The test is simple: the new on-call should be able to handle the first page without pinging the old one.
Error / query
on-call handoff template for SRE teamsUse this skill when
- Rotations change weekly and context gets lost between on-calls
- The outgoing on-call keeps getting pinged "just in case" after handoff
- You are standardizing handoffs across teams
- An agent needs to summarize or consume handoff notes
Not for this skill when
- You need an escalation policy (who gets paged when; different document)
- This is a permanent team knowledge transfer (deeper than a rotation handoff)
- You are writing incident status updates (those go to stakeholders, not the next on-call)
- The rotation is follow-the-sun with heavy overlap (lighter handoff suffices)
Steps
Step 1: List active and recently resolved incidents
echo "## Active incidents"
echo "- [INC-123]: [one-line status], next step [X], owner [name]"
echo "## Resolved this week (watch for recurrence)"
echo "- [INC-118]: [what broke], [what fixed it], [what to watch]"Expected: the incoming on-call knows what is still burning and what might reignite. Anything not written down here will be rediscovered at page time.
Step 2: Note fragile systems and in-flight changes
echo "## Fragile right now"
echo "- [service]: [why, e.g. 'deploy paused mid-rollout, do not deploy']"
echo "## In-flight changes"
echo "- [change]: [who is driving it], [expected completion], [rollback plan link]"Expected: the new on-call does not step on a half-finished migration or deploy into a known-bad state. This section prevents the most common handoff failures.
Step 3: Record the contact map
echo "## Who to call"
echo "- [service/team]: [person], [how to reach], [when, e.g. 'only if DB is involved']"
echo "- Escalation: [manager], [exec on-call]"Expected: no guessing about who owns what at 3am. Include the "when" for each contact; a name without a trigger condition just creates hesitation.
Step 4: Do the handoff live, then write it down
echo "15-min sync: outgoing walks through the notes, incoming asks questions."
echo "Then the notes go in [shared location] and the rotation flips."Expected: questions get asked while the outgoing on-call is still available. The written notes are the artifact; the call is what makes them complete.
Step 5: Set the expectation for post-handoff pings
echo "After [time], the old on-call is unreachable except for SEV1. Everything else goes to the new on-call."Expected: a clean break. Without an explicit cutoff, the outgoing on-call stays shadow on-call forever and the handoff never really happens.
Variant phrasings
"sre shift handover checklist"
Same template. Steps 1-3 are the checklist; step 4 is the ritual that makes it stick.
"on-call handoff doc example"
Use the section headers from steps 1-3 verbatim; fill each with one-liners, not paragraphs.
"how to stop getting paged after rotation ends"
Step 5: an explicit cutoff, plus making sure steps 1-3 were actually complete so the new on-call never needs you.
Why it happens
Context lives in the outgoing on-call's head: which alerts are noisy, which deploy is half done, which service is one bad deploy from paging. Without a written handoff, the new on-call learns all of it from the first page, which is the most expensive possible way to transfer knowledge. A template makes the transfer cheap and repeatable.
Edge cases and pitfalls
- Handoffs written as paragraphs do not get read; keep every item to one line.
- If there is "nothing to hand off," write that down explicitly; silence is ambiguous, "all clear" is information.
- Time zones: state the exact handoff time with zone; "EOD" means different things to different people.
- During incident-heavy weeks, do a mid-week mini-handoff, not just the scheduled one.
- Archive old handoffs; they become a useful history of what was fragile when.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_bCLOeZW1EIM81vII9cSaVA
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.