## TL;DR
Measure agent actions the way you measure responder actions: what fraction of actions achieved their intent without human correction, how long they took, and what fraction made things worse. Track it per incident and in aggregate, separate routine actions from judgment calls, and let the numbers (not demos) decide how much autonomy the agent earns. Success rate without severity weighting is vanity.

## The query
```text
how to measure agent action success rate in incidents
```

## Use this when
- AI agents act during incident response
- Evaluating whether agents help or hurt MTTR
- Deciding autonomy levels for incident agents
- Reporting on agent reliability to leadership

## Not for when
- General-purpose agent benchmarks
- Measuring human responder performance
- Non-incident agent task success

## Steps

### Step 1: Log every agent action with its intent
Record what the agent did and what it was trying to achieve, per action. Without the intent, you cannot judge success: a restarted pod is success if the intent was recovery, failure if the intent was diagnosis.
Expected output: an action log with intents, captured during the incident.

### Step 2: Score outcomes honestly after the incident
For each action, mark: succeeded unaided, succeeded with human correction, no effect, or made things worse. Score in the postmortem while memories are fresh, with the humans who supervised. Be brutal; inflated scores buy false confidence.
Expected output: per-action scores, agreed by the review participants.

### Step 3: Weight by blast radius, not just count
One bad database drop outweighs fifty successful log reads. Track success rate separately for read-only, standard-change, and high-blast-radius actions. The aggregate number hides the risk distribution; the segmented numbers reveal it.
Expected output: success rates per risk tier, showing where the agent is trustworthy and where it is not.

### Step 4: Track trends across incidents
Plot the rates over time: is the agent improving as guardrails and prompts improve, or flat? Correlate changes in success rate with changes you made (better runbooks, tighter scopes). Trends justify continued investment or trigger redesign.
Expected output: a trend line with annotations for what changed and when.

### Step 5: Tie autonomy to the numbers
Set explicit thresholds: above X% unaided success on standard changes earns broader scope; below Y% on high-blast-radius actions revokes it. Autonomy becomes a function of measured reliability, not of optimism.
Expected output: autonomy levels with numeric gates, reviewed as the numbers move.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_Bl_diuADRoVadpQCsThUqw
