how to measure MTTR with agents in the loop
Measures MTTR with agents in the loop using precise timestamp definitions and segmentation: overall MTTR split by agent-assisted vs human-only, then agent time vs handoff wait time. Use to evaluate agent impact on incident response or diagnose why MTTR moved. Not for agent-free response tracking or availability math.
TL;DR
MTTR with agents in the loop needs clean timestamps and honest segmentation: time from alert to service restored, split into agent-handled, human-handled, and handoff time. If you cannot separate "the agent was working" from "the agent was waiting on a human", the metric lies to you.
Error / query
how to measure MTTR with agents in the loopUse this skill when
- agents now participate in incident response and you need to show the effect
- leadership asks whether the agent investment is paying off
- MTTR moved and nobody can say why
- you are comparing agent-assisted response against the old baseline
Not for this skill when
- agents are not involved in response at all (standard MTTR tracking)
- you need MTTF or availability math (different metrics)
- the incident data is too sparse to segment (fix data collection first)
Steps
Step 1: Define the timestamps precisely and write them down
cat /opt/sre-docs/mttr-definitions.mdExpected: a doc defining alerttime (first page), acktime (first responder or agent engaged), mitigatetime (service restored), resolvetime (fully closed). MTTR is mitigate minus alert; everything else is segmentation. Without written definitions, every team measures something different.
Step 2: Pull incident data with actor attribution
gh api repos/[ORG]/incidents/contents/incidents-2026-q3.csv --jq ".content" | base64 -d | head -5Expected: rows with incident id, the four timestamps, and who did what (agent, human, or both). If your incident records do not say whether the agent participated, add that field now; you cannot measure what you did not record.
Step 3: Compute MTTR segmented by agent involvement
awk -F, "NR!=1 {m=($4-$1)/60; if ($5==\"agent\") a+=m, an++; else h+=m, hn++} END {print \"agent-assisted MTTR:\", a/an \"m over \" an \" incidents\"; print \"human-only MTTR:\", h/hn \"m over \" hn \" incidents\"}" incidents.csvExpected: two numbers with sample sizes. The comparison is only meaningful with enough incidents per segment; five agent incidents against fifty human ones is noise, not a finding.
Step 4: Break agent-assisted MTTR into agent time vs handoff time
awk -F, "NR!=1 && $5==\"agent\" {print $2-$1, $3-$2, $4-$3}" incidents.csv | awk "{d+=$1; h+=$2; r+=$3; n++} END {print \"detect:\",d/n \"s agent-wait:\",h/n \"s resolve:\",r/n \"s\"}"Expected: the three phases averaged. The handoff column is where agent-assisted response usually loses its gains: a fast agent that waits twenty minutes for a human approval is not actually fast.
Step 5: Track the trend and the confounders together
awk -F, "NR!=1 {print substr($1,1,7), ($4-$1)/60, $5}" incidents.csv | sort | uniq -c | tail -20Expected: monthly MTTR by segment. Annotate the timeline with what changed (agent got write access in July, new approval gate in August); MTTR moves for many reasons, and the agent is only one of them.
Variant phrasings
"prove the sre agent reduces mttr"
Steps 3 and 4 with a pre-agent baseline. Report the segmented numbers with sample sizes and confounders, not a single triumphant percentage.
"agent mttr got worse after we gave it more access"
Check the handoff column first: broader access often means more approval gates, and the waiting time can swamp the faster diagnosis. The fix is streamlining approvals, not removing access.
"what is a good mttr target with agents"
There is no universal number; target beating your own human-only baseline consistently, then tightening. The trend matters more than the absolute value.
Why it happens
MTTR is easy to game and easy to misread: agents make detection instant but add handoff latency, humans cherry-pick easy incidents for the agent, and severity mix shifts silently. Segmentation exists to defeat these illusions. The teams that get value from the metric are the ones that treat it as a diagnostic (where is the time going) rather than a scoreboard (are we winning).
Edge cases and pitfalls
- Mitigated vs resolved: agents often mitigate fast and resolve slow; pick one definition and stick to it across segments.
- Business-hours vs off-hours incidents have different human response times; segment those too or the agent looks better than it is.
- Do not let the metric drive perverse incentives like closing incidents before the fix is verified to juice resolve_time.
- Small samples lie; do not publish agent-vs-human comparisons under about twenty incidents per side.
- If the agent touches every incident (triage bot), the "human-only" baseline disappears; then measure phase timings instead of the binary split.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_uzgpS3dZI9vg0JC4sukXPg
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.