writing status-page updates during an outage
Shows support and ops teams how to write status-page updates during an outage: what to say in the first 10 minutes, update cadence, what goes in each update, and how to write the all-clear. Use when drafting incident communication, building a status-page runbook, or training on-call responders. Not for internal incident command, postmortem writing, or proactive maintenance notices.
TL;DR
Status-page updates during an outage need three things: speed (first update within 10 minutes, even if it only says "we are investigating"), cadence (every 20 to 30 minutes until resolved), and substance (what is affected, what users can do, what you are doing). Never speculate about causes or ETAs you cannot defend. The all-clear should confirm the fix is holding, not just deployed.
The query
writing status-page updates during an outageUse this when
- An incident is active and customers are asking
- You need a status-page runbook for on-call
- Updates are going out late or saying nothing useful
- You are reviewing incident communication quality
Not for
- Internal incident-command communication
- Postmortems after the incident
- Scheduled maintenance announcements
- Security breach notifications (different rules)
Steps
1. Post the first update within 10 minutes
Even if all you know is "we are seeing errors on logins and investigating." Silence in the first 10 minutes is when the ticket flood starts.
Expected output: a posted update acknowledging the problem before customers write 50 tickets about it.
2. State what is affected and what still works
Name the broken surface precisely ("invoice exports are failing") and say what is fine ("logins and the dashboard are unaffected"). Users calibrate their own panic from this.
Expected output: every update names the affected scope and the confirmed-working scope.
3. Tell users what they can do right now
If there is a workaround, lead with it. If there is not, say so and tell them what not to do ("please dont retry bulk exports, they will queue and slow the recovery").
Expected output: every update contains a user action, even if the action is "wait."
4. Say what you are doing, not what you think
"We have identified the failing service and are rolling back the deploy" is fine. "We think it might be the database" is not. No speculation, no blame, no technical archaeology on the status page.
Expected output: each update describes actions taken, never theories.
5. Update every 20 to 30 minutes until resolved
Set a timer. Even "still working on it, next update at 14:30" keeps trust. The cadence matters more than the content once the basics are out.
Expected output: a visible update cadence the team actually keeps.
6. Write the all-clear as a confirmation, not a deployment
"Exports are working again and we have confirmed error rates are back to normal over the last 30 minutes." Then monitor for another hour before marking fully resolved.
Expected output: the resolved update confirms the fix is holding, with a monitoring window.
Update templates
First update
Investigating: we are seeing failures on invoice exports starting around
14:05. Logins and the dashboard are unaffected. We are investigating now
and will update within 30 minutes.Mid-incident
Update 14:35: we have identified the failing service and are rolling back
the deploy that caused it. Please dont retry bulk exports for now, they
will queue and slow the recovery. Next update at 15:05.All-clear
Resolved: exports are working again. Error rates have been back to normal
for the last 30 minutes and we will keep monitoring. Full postmortem
within 48 hours.Variant phrasings
incident communication template for status page
The three templates above cover the arc. Adapt the timestamps and scope, keep the structure.
how to write outage updates for customers
Same six steps, applies to email and in-app banners too. Shorter for banners, same substance.
status page best practices during downtime
Cadence and honesty are the whole game. A status page that updates every 25 minutes with real substance keeps ticket volume at a fraction of a silent one.
Why it happens
Customers check the status page to answer one question: is it me or is it them, and what do I do. Updates that fail to answer that question send every reader to the support queue instead. Fast, regular, substantive updates do not just inform, they deflect. Teams that update well see the ticket spike flatten within minutes of the first post.
Edge cases
- You dont know the cause yet: say "investigating," give the scope, set the next update time. That is a complete update.
- ETA pressure: never publish an ETA you would not bet on. "Next update at 15:05" is a promise you can keep; "fixed by 15:05" usually is not.
- Partial outages: be precise about who is affected ("exports for files over 50MB") so unaffected users dont panic.
- Long incidents (4+ hours): add a summary of the timeline so far to each update, new readers should not have to read five updates to understand the state.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_4PmxhAwPKhg3h0oFch6bZw
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.