how to reduce observability costs without losing signal
Shows how to reduce observability costs without losing signal: find cost drivers, cut high-cardinality unqueried metrics at the agent, and tier retention. Use when the bill jumps or ingestion defaults were never tuned. Not when alerts are already missing or compliance mandates full retention.
TL;DR
Most observability waste comes from three places: high-cardinality metrics nobody queries, debug logs shipped at full volume, and retention set to "forever" by default. Find the top cost drivers with your vendor's usage report, cut cardinality and drop unused data at the agent, and shorten retention for low-value data first. You can usually cut 30-50% without losing a single alert or dashboard.
Error / query
how to reduce observability costs without losing signalUse this skill when
- The observability bill jumped and nobody can say what changed
- Finance is asking why monitoring costs scale faster than revenue
- You inherited a setup that ingests everything at full fidelity
- You want a repeatable cost review, not a one-time panic cut
Not for this skill when
- You are already missing alerts or dashboards; fix coverage first, then optimize cost
- The bill is dominated by a fixed platform fee; renegotiate the contract instead
- Compliance mandates full retention; cost cuts cannot touch the mandated data
Steps
Step 1: Find which data types drive the bill
curl -s -G 'https://example.com/api/v1/usage' --data-urlencode 'period=30d' | head -c 800Expected: a breakdown by logs, metrics, traces; one of them is usually 60%+ of the cost and that is where you start.
Step 2: Find the highest-cardinality metric names
curl -s -G 'https://example.com/api/v1/query' --data-urlencode 'query=topk(10, count by (__name__)({__name__=~".+"}))' | head -c 800Expected: metric names with enormous series counts; labels like user ID or request path on a counter are the classic offenders.
Step 3: Check which dashboards actually query those expensive metrics
grep -rl 'expensive_metric_name' /etc/grafana/provisioning/dashboards/ 2>/dev/null | head; echo "---alerts---"; grep -rl 'expensive_metric_name' /etc/prometheus/rules/ 2>/dev/null | headExpected: often nothing references them; unqueried high-cardinality metrics are pure waste, drop them at the agent.
Step 4: Drop the wasteful series with a relabel rule
printf -- '- source_labels: [__name__]\n regex: "expensive_metric_name"\n action: drop\n' > /tmp/drop-waste.yaml && cat /tmp/drop-waste.yamlExpected: the drop rule prints back; add it to your scrape config and reload, then watch series count fall.
Step 5: Verify cost-relevant volume actually dropped
curl -s -G 'https://example.com/api/v1/query' --data-urlencode 'query=sum(scrape_samples_scraped)' | head -c 300Expected: a lower samples-scraped number than before the change; confirm dashboards and alerts still evaluate green.
Variant phrasings
"Cut Datadog costs without losing visibility"
Same playbook: drop unused custom metrics, shorten log retention for noisy services, and switch low-value APM services to lower sampling.
"Observability bill doubled, what do I check first"
The usage breakdown by data type, then the top 10 series by cardinality; one deploy that added a high-cardinality label is the usual suspect.
"Cheaper log storage that still works for incidents"
Keep 7-14 days hot for incident response, move the rest to cheap cold storage; almost no incident needs 90-day-old logs at hot query speed.
Why it happens
Ingestion defaults are generous and nobody owns the bill, so every new service ships debug logs and unlabeled metrics at full volume. Cost grows with data nobody ever looks at, and the fix is deletion at the source, not a better dashboard.
Edge cases and pitfalls
- Do not drop metrics that feed SLO burn-rate alerts; audit alert queries before dropping anything.
- Cutting retention for security or audit logs can violate policy; check with compliance first.
- Dropping labels breaks existing dashboard queries that group by them; search dashboards for the label first.
- One big cut can mask a real regression; change one data source at a time and watch for a week.
Provenance
Resolved from the public thread: https://vectle.com/posts/pstn-KZVzWnyjfqO4V5mG7VQ
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.