## TL;DR
Keep recent logs on fast hot storage for incident debugging, move older logs to warm storage for occasional queries, and archive everything else to cold object storage for compliance. Most teams find 7-14 days hot and 30-90 days warm covers real debugging needs; the rest is rarely read but legally required. Automate the transitions with lifecycle rules so nobody has to think about it.

## The query
```text
how to tier log storage hot-warm-cold
```

## Use this when
- Log storage costs are a line item people complain about
- You need months of retention but only query recent days
- Compliance or audit requires long retention you cannot afford hot
- Query performance on old logs does not matter much

## Not for when
- Picking between logging backends
- Real-time log alerting setup
- Reducing log volume at the source (do that too, separately)

## Steps

### Step 1: Measure how far back anyone actually queries
Check query logs for the last 90 days: what is the oldest timestamp anyone searched? For most teams, 95 percent of queries hit the last 7 days. Let that number set your hot tier, not a guess.
Expected output: a histogram of query age that justifies the hot window with data.

### Step 2: Define the three tiers and their retention
Hot: fast indexed storage, 7-14 days, full query speed. Warm: cheaper searchable storage, 30-90 days, slower queries acceptable. Cold: object storage, 1-7 years, retrieval in minutes to hours, queried rarely.
Expected output: a written retention table with a tier, duration, and cost for each.

### Step 3: Automate transitions with lifecycle rules
Configure index lifecycle or bucket lifecycle policies so data moves between tiers on age alone. Test the transition on a small index first: verify queries still work on warm data and that cold archives are restorable.
Expected output: logs age out of hot storage with zero manual work; a spot check confirms old logs are retrievable.

### Step 4: Tell the team what changes about old-log queries
Document that queries beyond the hot window are slower and use different syntax or a different UI if applicable. The common failure mode is someone debugging an incident at 2am not knowing last month's logs live somewhere else.
Expected output: a short runbook note linked from the logging docs, so nobody discovers the tiers during an incident.

### Step 5: Review the tiers quarterly
Query patterns change as the team and product change. Every quarter, re-check the query-age histogram and adjust. The usual direction is shrinking hot as people get comfortable with warm.
Expected output: storage spend trends down or flat while data volume grows.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_muwu_0-GFDCrQ8XKYwE4_w
