building a service health dashboard executives trust
Shows how to build a service health dashboard executives trust: one status color, three headline numbers, 30-day trends, all from the same SLO queries engineering uses. Use when leadership doubts the status page or asks for health in Slack. Not for engineering debug views or per-customer impact.
TL;DR
Executives trust a dashboard with four things: one big status (green/yellow/red), the three numbers they already ask about (uptime, error rate, latency), a 30-day trend so today has context, and zero metrics they have to interpret. Build it from the same SLO queries engineering uses, so the dashboard can never disagree with the on-call. If it needs a legend, it is too complicated.
Error / query
building a service health dashboard executives trustUse this skill when
- Leadership does not believe the status page because it disagreed with reality once
- You get asked "is everything OK" in Slack and answer from a different dashboard each time
- The current exec dashboard has 40 panels and nobody looks at it
- You need one shared view for incidents, reviews, and board updates
Not for this skill when
- You need an engineering debug dashboard; this is the opposite, it hides detail on purpose
- There are no SLOs yet; the dashboard needs agreed numbers to display
- The audience is support or customer success; they need per-customer impact, not aggregate health
Steps
Step 1: Confirm the SLO queries the dashboard will reuse
curl -s -G 'https://example.com/api/v1/query' --data-urlencode 'query=sum(rate(http_requests_total{status!~"5.."}[30d])) / sum(rate(http_requests_total[30d]))' | head -c 200Expected: the 30-day success ratio; this exact query powers the dashboard's headline number, so it can never disagree with engineering.
Step 2: Check the latency number executives will see
curl -s -G 'https://example.com/api/v1/query' --data-urlencode 'query=histogram_quantile(0.95, sum(rate(http_request_duration_seconds_bucket[30d])) by (le))' | head -c 200Expected: the 30-day p95 latency; one number, no percentiles soup.
Step 3: Verify error budget remaining for the status color
curl -s -G 'https://example.com/api/v1/query' --data-urlencode 'query=1 - ((1 - sum(rate(http_requests_total{status!~"5.."}[30d])) / sum(rate(http_requests_total[30d]))) / (1 - 0.995))' | head -c 200Expected: budget fraction remaining; green above 0.5, yellow 0.2-0.5, red below 0.2, same thresholds as the on-call policy.
Step 4: Check the dashboard renders from the same data source
curl -s https://example.com/api/dashboards/health-exec | grep -c 'example.com/api/v1/query'Expected: a nonzero count; every panel queries the same metrics backend the alerts use, no separate pipeline to drift.
Step 5: Confirm the dashboard loads without login friction for execs
curl -s -o /dev/null -w '%{http_code}' https://example.com/d/health-execExpected: 200 (or your SSO redirect working); a dashboard that needs VPN plus three clicks gets checked never.
Variant phrasings
"What metrics do executives actually want"
Uptime percentage, error rate, latency p95, and incident count, each with a 30-day trend; everything else is engineering detail.
"Status page vs internal exec dashboard"
The status page is for customers and must be conservative; the exec dashboard can show yellow earlier, but the two must never contradict on red.
"How to stop execs from asking for more panels"
Agree the four numbers up front and put everything else one click away; panel creep is how trusted dashboards die.
Why it happens
Exec dashboards lose trust the first time they disagree with reality or with the on-call's numbers. Reusing the exact SLO queries guarantees consistency, and ruthless simplicity guarantees it gets looked at.
Edge cases and pitfalls
- A dashboard that is green during a real incident destroys trust permanently; wire the status color to the same alerts that page.
- Do not average across services with wildly different SLOs; show the worst offender, not the mean.
- Timezone confusion on "today" panels causes false alarms; label the timezone on the dashboard.
- Refresh the trend data on the same cadence as SLO windows; a stale 30-day number is worse than none.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_-DCvw0MV94kcwYzRt8l1sQ