## TL;DR
Rebuilding a timeline after the fact is a join on timestamps: pull logs from each involved service for the incident window, normalize to one timezone, and anchor on known markers (the alert time, the deploy time, the first error). Sort everything into one chronological list and the causal chain usually reads itself. Work backward from impact to trigger, not forward from the start.

## Error / query
```text
how to reconstruct an incident timeline from logs after the fact
```

## Use this skill when
- The incident is over and the sequence is murky
- Building the timeline section of a postmortem
- Correlating events across multiple services
- Nobody wrote a live timeline during the incident

## Not for this skill when
- The incident is still active (timeline live, not after)
- You need metrics rather than log events
- Logs were never collected (pipeline problem)

## Steps

### Step 1: Fix the incident window and timezone
```bash
# anchor markers: alert fired at [T0], impact started [T-?], resolved [T1]
# pull logs from T0-30m to T1+15m, all in UTC
```
Expected: a bounded window in a single timezone. Mixing timezones is the classic timeline corruption; convert everything to UTC before merging.

### Step 2: Pull the key log streams for the window
```bash
kubectl logs --since=[duration] -n [namespace] deploy/[service] > service-a.log
# repeat for each involved service, plus ingress/gateway logs
```
Expected: one file per source covering the window. Include the load balancer or ingress logs; they show user-facing impact timing better than app logs.

### Step 3: Extract timestamped event lines and merge
```bash
grep -hE "ERROR|FATAL|exception|timeout" service-*.log | sort > merged-events.log
wc -l merged-events.log
```
Expected: a single chronologically sorted list of significant events. Start with errors and exceptions; you can widen to warnings if the chain has gaps. The merge is where the story appears.

### Step 4: Walk backward from impact to trigger
```text
Find the first user-facing error in the merged log.
Walk backward: what changed just before it? A deploy, a config push,
a dependency error, a traffic spike? That is your trigger candidate.
```
Expected: a trigger event preceding the impact with a plausible causal link. Validate the candidate against deploy records and config change logs; correlation in the timeline is a hypothesis until confirmed.

## Variant phrasings

### "incident timeline from logs"
Steps 1-4. Window, pull, merge, walk backward.

### "postmortem timeline reconstruction"
Same process. The merged log (step 3) becomes the postmortem's timeline section.

## Why it happens
During incidents nobody has time to write the timeline, and afterward memory is unreliable about ordering: people remember what they did, not when. Logs are the only objective record, and merging them restores the sequence that memory loses. Working backward from impact avoids the trap of starting at the beginning and drowning in noise.

## Edge cases and pitfalls
- Clock skew between hosts can misorder events by seconds; for sub-second causality, use request IDs to chain events instead of timestamps.
- Log sampling drops events under load; the absence of a log line is not proof the event did not happen.
- Correlate with deploy and config-change records, not just app logs; the trigger is often outside the application's own logs.
- Very verbose services drown the merge; filter to error-level first, then widen only around the interesting minutes.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_3Qeg2XXi-9DlMIAKl6tCAA
