status page templates for degraded performance
Copy-paste status page templates for degraded performance: investigating, identified, monitoring, and resolved, each in a three-part shape with timestamps and a 30-minute update clock. Use when service is slow but up, when a feature is partially broken, or when building incident presets. Not for full outages, scheduled maintenance, or internal war-room comms.
TL;DR
A good status page update names what is degraded, who it affects, and what you are doing, in that order. Keep each update under 80 words, timestamp everything, and post on a schedule even when nothing changed. Customers check the status page to decide whether to file a ticket, so every unclear update costs you ten tickets.
The query
status page templates for degraded performanceUse this when
- Response times or error rates are elevated but the service is up
- A feature is partially broken for some customers
- You need pre-written templates before the next incident
- Updates are going out inconsistent or too slow
Not for
- Full outages (different severity, different wording)
- Scheduled maintenance announcements
- Internal incident war-room comms
- Postmortem writeups
Steps
1. Post the first update within 15 minutes of detection
Speed beats completeness. "We are investigating degraded performance in [area]" with a timestamp is enough. The investigation details come later.
Expected output: first post live within 15 minutes, every time.
2. Follow the same three-part shape every update
What is degraded, who is affected, what is happening now. Same order, every time. Regulars learn to scan your updates in seconds.
Expected output: no update goes out missing any of the three parts.
3. Timestamp and version your updates
Every post carries a time and a short status word: Investigating, Identified, Monitoring, Resolved. Customers scrolling a long incident need to find the latest state instantly.
Expected output: status word plus timestamp at the top of every update.
4. Update on a clock, not on news
Every 30 minutes during degraded performance, with or without progress. "Still investigating, next update at 3:30pm" is a complete update.
Expected output: no gap longer than 30 minutes without a post.
5. Close the loop with a resolved note
When metrics are back to normal, say so and note how long it lasted. "Resolved at 4:12pm after 2 hours of degraded response times." Then stop posting.
Expected output: a final resolved update on every incident, no dangling "monitoring" forever.
Ready-to-use template
INVESTIGATING - [time, home_tz]
We are seeing slower than normal [response times / error rates] in [area].
Affected: [who, e.g. customers in the EU region / API users].
Our team is investigating. Next update at [time].
IDENTIFIED - [time, home_tz]
We found the cause: [one-line cause, plain words]. [Area] is still slower
than normal for [who]. We are [fix in progress / rolling back / scaling up].
Next update at [time].
MONITORING - [time, home_tz]
[Area] is back to normal speeds as of [time]. We are watching the numbers
closely for the next hour. Next update at [time], or sooner if anything changes.
RESOLVED - [time, home_tz]
Resolved. [Area] ran slower than normal from [start] to [end] ([duration]).
We will publish a summary of what happened within [timeframe].Variant phrasings
degraded performance status page wording
The four templates cover the full lifecycle. Copy them into your status tool as presets.
how to write a status update for slow service
Steps 2 and 3. The three-part shape plus the status word is the whole craft.
status page copy for partial outage
Same templates, but say "partial" plainly: "Some customers in [region] are affected." Dont let people guess whether it is them.
Why it happens
Status pages go wrong when engineers write them like changelogs, full of internal detail and missing the customer's question: "is it me, and should I wait?" The three-part shape answers exactly that. And the 30-minute clock exists because silence during an incident reads as "they dont know," even when the team is working hard.
Edge cases
- Degraded but you dont know the scope yet: say so. "We are still determining who is affected" is better than guessing wrong and correcting later.
- It degrades, recovers, degrades again: post each change as its own update. A flapping incident with honest posts beats one "monitoring" post that is quietly wrong.
- Customers file tickets anyway: that is fine. Link the ticket back to the status page and close the loop there.
- The cause is embarrassing (a bad deploy, a config typo): say it plainly in the summary. Customers forgive mistakes; they dont forgive cover stories.
- Degraded performance that lasts days: switch from 30-minute updates to twice daily, and say you switched. "We are moving to twice-daily updates on this" manages the expectation.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_D0TwWaYh0CxG6XAiA8fHeA
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.