## TL;DR
Auto-triage works when the rules are simple, based on measurable signals (customer impact, error rate, affected region), and wrong in a safe direction (over-triage beats under-triage for the first pass). Start with three rules covering your most common incidents, let humans correct the severity, and feed the corrections back into the rules. Complexity is the enemy; simple rules people trust beat clever ones nobody understands.

## The query
```text
incident severity auto-triage: rules that work
```

## Use this when
- Severity assignment is inconsistent between responders
- Everything gets called SEV1 (severity inflation)
- Building severity into paging automation
- Reviewing whether auto-triage is worth it

## Not for when
- Defining what SEV1 vs SEV2 means (do that first)
- Manual severity assessment training
- Postmortem severity debates

## Steps

### Step 1: Write severity definitions in measurable terms
Before automating, define each severity with numbers: SEV1 means customer-facing impact or data loss risk; SEV2 means degraded but working; SEV3 means internal only. If humans cannot apply the definitions consistently, automation cannot either.
Expected output: severity definitions a new hire could apply without asking.

### Step 2: Start with three rules for the common cases
Pick the three most frequent incident types and write one rule each: e.g. checkout error rate above X becomes SEV1; single-region degradation becomes SEV2; internal tool outage becomes SEV3. Ship only these.
Expected output: three rules covering the majority of incidents, each simple enough to explain in one sentence.

### Step 3: Default to over-triage, correct down fast
When signals are ambiguous, assign the higher severity. It is cheaper to downgrade a SEV1 in ten minutes than to discover a SEV3 was customer-facing for an hour. Make downgrading easy and blameless.
Expected output: no incident discovered to be worse than its initial severity hours later.

### Step 4: Let humans override and log the correction
Every auto-assigned severity must be one click to change, and every change gets logged with the reason. The log is the training data for better rules.
Expected output: a correction log showing where the rules are wrong.

### Step 5: Review and refine monthly
Monthly, review the corrections: which rules misfire, which incidents had no rule. Adjust thresholds, add rules for new patterns, retire rules for dead services. Auto-triage is a living system.
Expected output: triage accuracy improving month over month; rules matching the current architecture.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_NTjTzAnymsV1WwbPo8H5Sg
