## TL;DR
Pick one or two SLIs that measure what users actually feel (availability first, latency second), set the target from your last 30 days of measured data instead of guessing 99.9%, and alert on burn rate so you get paged hours before the budget runs out. An SLO the on-call cannot explain in one sentence is decoration.

## Error / query
```text
how to set SLOs that matter: the basics
```

## Use this skill when
- A service has no SLO and you need to define one from scratch
- The current SLO is either never breached (too loose to matter) or permanently red (nobody believes it)
- You want to move from "is the service down" paging to burn-rate alerts
- Leadership asks "how reliable is this" and nobody has a number

## Not for this skill when
- You are drafting an external SLA contract with penalties; that is legal work, not engineering
- The workload is an internal batch job with no user-facing latency or availability
- You need per-tenant SLOs; start with the aggregate service SLO first

## Steps

### Step 1: Find your top user-facing endpoints by traffic
```bash
curl -s -G 'https://example.com/api/v1/query' --data-urlencode 'query=topk(10, sum(rate(http_requests_total[24h])) by (route))' | head -c 1500
```
Expected: JSON listing your 10 busiest routes; those are your SLI candidates.

### Step 2: Measure the real success ratio over the last 30 days
```bash
curl -s -G 'https://example.com/api/v1/query' --data-urlencode 'query=sum(rate(http_requests_total{route="/api/checkout",status!~"5.."}[30d])) / sum(rate(http_requests_total{route="/api/checkout"}[30d]))'
```
Expected: a ratio like 0.9987; set the target just under reality, not at a wish number.

### Step 3: Write the SLO to a file your alerting can read
```bash
printf 'slo:\n  name: checkout-availability\n  sli: good_requests / total_requests\n  target: 0.995\n  window: 30d\n' > /tmp/slo.yaml && cat /tmp/slo.yaml
```
Expected: the file prints back with target 0.995, the measured value rounded down slightly.

### Step 4: Check the current burn rate against that target
```bash
curl -s -G 'https://example.com/api/v1/query' --data-urlencode 'query=(1 - sum(rate(http_requests_total{status!~"5.."}[1h])) / sum(rate(http_requests_total[1h]))) / (1 - 0.995)'
```
Expected: a burn-rate number; anything above 1 means you are spending budget faster than the 30-day pace allows.

### Step 5: Confirm the alert rule is loaded
```bash
curl -s https://example.com/api/v1/rules | grep -c checkout-availability
```
Expected: a count of 1 or more; the SLO and its alert are live.

## Variant phrasings

### "What should my SLO be for a brand new service"
Start at 99.5% over 30 days and tighten after a month of real data; new services almost never deserve 99.99%.

### "SLI vs SLO vs SLA, what is the difference"
SLI is the measurement (success ratio), SLO is the target (99.5%), SLA is the contract (what you owe customers when you miss). This skill covers the first two.

### "How many SLOs should one service have"
One per user-facing concern, availability first, latency second. Past three, nobody remembers them.

## Why it happens
Teams either skip SLOs or copy 99.99% from a blog post, so alerts fire on noise or never at all. Grounding the target in measured history and alerting on burn rate turns the SLO from a vanity metric into an on-call tool.

## Edge cases and pitfalls
- Do not pick 99.99% because it sounds good; one bad deploy eats the whole budget.
- Latency SLOs need percentiles (p99), never averages.
- Do not set tight SLOs on dependencies you do not control; split third-party APIs into their own looser SLO.
- Revisit targets quarterly; traffic mix drifts and last year's 99.9% can quietly become this year's 99.5%.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_9QGPzmgg_l2poScrZr_glg
