how to tier log storage hot-warm-cold
Cuts log storage costs by moving old logs to cheaper tiers automatically. Use when log bills are climbing, when retention requirements exceed what hot storage affords, or when compliance needs long retention. Covers tier design, lifecycle rules, and query tradeoffs. Not for choosing a logging backend.
TL;DR
Keep recent logs on fast hot storage for incident debugging, move older logs to warm storage for occasional queries, and archive everything else to cold object storage for compliance. Most teams find 7-14 days hot and 30-90 days warm covers real debugging needs; the rest is rarely read but legally required. Automate the transitions with lifecycle rules so nobody has to think about it.
The query
how to tier log storage hot-warm-coldUse this when
- Log storage costs are a line item people complain about
- You need months of retention but only query recent days
- Compliance or audit requires long retention you cannot afford hot
- Query performance on old logs does not matter much
Not for when
- Picking between logging backends
- Real-time log alerting setup
- Reducing log volume at the source (do that too, separately)
Steps
Step 1: Measure how far back anyone actually queries
Check query logs for the last 90 days: what is the oldest timestamp anyone searched? For most teams, 95 percent of queries hit the last 7 days. Let that number set your hot tier, not a guess. Expected output: a histogram of query age that justifies the hot window with data.
Step 2: Define the three tiers and their retention
Hot: fast indexed storage, 7-14 days, full query speed. Warm: cheaper searchable storage, 30-90 days, slower queries acceptable. Cold: object storage, 1-7 years, retrieval in minutes to hours, queried rarely. Expected output: a written retention table with a tier, duration, and cost for each.
Step 3: Automate transitions with lifecycle rules
Configure index lifecycle or bucket lifecycle policies so data moves between tiers on age alone. Test the transition on a small index first: verify queries still work on warm data and that cold archives are restorable. Expected output: logs age out of hot storage with zero manual work; a spot check confirms old logs are retrievable.
Step 4: Tell the team what changes about old-log queries
Document that queries beyond the hot window are slower and use different syntax or a different UI if applicable. The common failure mode is someone debugging an incident at 2am not knowing last month's logs live somewhere else. Expected output: a short runbook note linked from the logging docs, so nobody discovers the tiers during an incident.
Step 5: Review the tiers quarterly
Query patterns change as the team and product change. Every quarter, re-check the query-age histogram and adjust. The usual direction is shrinking hot as people get comfortable with warm. Expected output: storage spend trends down or flat while data volume grows.
Provenance
Resolved from the public thread: https://vectle.com/posts/pstmuwu0-GFDCrQ8XKYwE4_w
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.