how to set SLOs that matter: the basics
Shows an SRE or on-call engineer how to set SLOs that matter: pick user-facing SLIs, set targets from measured history, and alert on burn rate. Use when a service needs its first SLO or an existing one is ignored. Not for external SLA contracts or internal batch jobs with no user impact.
TL;DR
Pick one or two SLIs that measure what users actually feel (availability first, latency second), set the target from your last 30 days of measured data instead of guessing 99.9%, and alert on burn rate so you get paged hours before the budget runs out. An SLO the on-call cannot explain in one sentence is decoration.
Error / query
how to set SLOs that matter: the basicsUse this skill when
- A service has no SLO and you need to define one from scratch
- The current SLO is either never breached (too loose to matter) or permanently red (nobody believes it)
- You want to move from "is the service down" paging to burn-rate alerts
- Leadership asks "how reliable is this" and nobody has a number
Not for this skill when
- You are drafting an external SLA contract with penalties; that is legal work, not engineering
- The workload is an internal batch job with no user-facing latency or availability
- You need per-tenant SLOs; start with the aggregate service SLO first
Steps
Step 1: Find your top user-facing endpoints by traffic
curl -s -G 'https://example.com/api/v1/query' --data-urlencode 'query=topk(10, sum(rate(http_requests_total[24h])) by (route))' | head -c 1500Expected: JSON listing your 10 busiest routes; those are your SLI candidates.
Step 2: Measure the real success ratio over the last 30 days
curl -s -G 'https://example.com/api/v1/query' --data-urlencode 'query=sum(rate(http_requests_total{route="/api/checkout",status!~"5.."}[30d])) / sum(rate(http_requests_total{route="/api/checkout"}[30d]))'Expected: a ratio like 0.9987; set the target just under reality, not at a wish number.
Step 3: Write the SLO to a file your alerting can read
printf 'slo:\n name: checkout-availability\n sli: good_requests / total_requests\n target: 0.995\n window: 30d\n' > /tmp/slo.yaml && cat /tmp/slo.yamlExpected: the file prints back with target 0.995, the measured value rounded down slightly.
Step 4: Check the current burn rate against that target
curl -s -G 'https://example.com/api/v1/query' --data-urlencode 'query=(1 - sum(rate(http_requests_total{status!~"5.."}[1h])) / sum(rate(http_requests_total[1h]))) / (1 - 0.995)'Expected: a burn-rate number; anything above 1 means you are spending budget faster than the 30-day pace allows.
Step 5: Confirm the alert rule is loaded
curl -s https://example.com/api/v1/rules | grep -c checkout-availabilityExpected: a count of 1 or more; the SLO and its alert are live.
Variant phrasings
"What should my SLO be for a brand new service"
Start at 99.5% over 30 days and tighten after a month of real data; new services almost never deserve 99.99%.
"SLI vs SLO vs SLA, what is the difference"
SLI is the measurement (success ratio), SLO is the target (99.5%), SLA is the contract (what you owe customers when you miss). This skill covers the first two.
"How many SLOs should one service have"
One per user-facing concern, availability first, latency second. Past three, nobody remembers them.
Why it happens
Teams either skip SLOs or copy 99.99% from a blog post, so alerts fire on noise or never at all. Grounding the target in measured history and alerting on burn rate turns the SLO from a vanity metric into an on-call tool.
Edge cases and pitfalls
- Do not pick 99.99% because it sounds good; one bad deploy eats the whole budget.
- Latency SLOs need percentiles (p99), never averages.
- Do not set tight SLOs on dependencies you do not control; split third-party APIs into their own looser SLO.
- Revisit targets quarterly; traffic mix drifts and last year's 99.9% can quietly become this year's 99.5%.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst9QGPzmggl2poScrZr_glg
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.