incident severity auto-triage: rules that work
Designs auto-triage rules that assign incident severity correctly. Use when severity is inconsistent, when every incident becomes a SEV1, or when building auto-triage into paging. Covers rule design and the human override. Not for severity definitions themselves.
TL;DR
Auto-triage works when the rules are simple, based on measurable signals (customer impact, error rate, affected region), and wrong in a safe direction (over-triage beats under-triage for the first pass). Start with three rules covering your most common incidents, let humans correct the severity, and feed the corrections back into the rules. Complexity is the enemy; simple rules people trust beat clever ones nobody understands.
The query
incident severity auto-triage: rules that workUse this when
- Severity assignment is inconsistent between responders
- Everything gets called SEV1 (severity inflation)
- Building severity into paging automation
- Reviewing whether auto-triage is worth it
Not for when
- Defining what SEV1 vs SEV2 means (do that first)
- Manual severity assessment training
- Postmortem severity debates
Steps
Step 1: Write severity definitions in measurable terms
Before automating, define each severity with numbers: SEV1 means customer-facing impact or data loss risk; SEV2 means degraded but working; SEV3 means internal only. If humans cannot apply the definitions consistently, automation cannot either. Expected output: severity definitions a new hire could apply without asking.
Step 2: Start with three rules for the common cases
Pick the three most frequent incident types and write one rule each: e.g. checkout error rate above X becomes SEV1; single-region degradation becomes SEV2; internal tool outage becomes SEV3. Ship only these. Expected output: three rules covering the majority of incidents, each simple enough to explain in one sentence.
Step 3: Default to over-triage, correct down fast
When signals are ambiguous, assign the higher severity. It is cheaper to downgrade a SEV1 in ten minutes than to discover a SEV3 was customer-facing for an hour. Make downgrading easy and blameless. Expected output: no incident discovered to be worse than its initial severity hours later.
Step 4: Let humans override and log the correction
Every auto-assigned severity must be one click to change, and every change gets logged with the reason. The log is the training data for better rules. Expected output: a correction log showing where the rules are wrong.
Step 5: Review and refine monthly
Monthly, review the corrections: which rules misfire, which incidents had no rule. Adjust thresholds, add rules for new patterns, retire rules for dead services. Auto-triage is a living system. Expected output: triage accuracy improving month over month; rules matching the current architecture.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_NTjTzAnymsV1WwbPo8H5Sg
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.