# which log fields actually matter for security investigations

## TL;DR

Every log line that matters for security answers five questions: when, who, from where, what did they do, and did it work. That means a normalized UTC timestamp, a user or service identity, a source IP, an action name, the target resource, and a success or failure result. Add a request or correlation ID and you can trace a whole attack across services. Everything else is nice to have.

```text
which log fields actually matter for security investigations
```

## Use this when

- You are deciding what to log in a new service
- Your logs cannot answer "who did what from where" and you need to fix that
- You are writing a logging standard for your team
- An investigation just failed for lack of fields and you want it to never happen again

## Not for this skill when

- You need full log schema design for compliance (different, stricter skill)
- You want application performance telemetry (different fields, different skill)
- You are choosing a log storage backend (storage is separate from fields)
- You need help parsing logs you already have (parsing is downstream of this)

## Steps

### 1. Require the six core fields on every security-relevant event

Timestamp (UTC, ISO format), identity (user, service account, or API key id), source IP, action or event name, target resource, and result (success, failure, denied). If any one is missing, the event is half a clue.

```bash
jq -r 'select(.event=="login") | [.ts, .user, .ip, .action, .resource, .result] | @tsv' app.log | head -5
```

Expected: neat columns with all six fields populated on every line. Blank columns are your backlog: each blank is a question you cannot answer in an investigation. (jq 1.6+.)

### 2. Normalize timestamps to UTC at write time

Mixed timezones and local times make correlation a guessing game. Write ISO 8601 UTC everywhere; convert for display only.

```bash
jq -c 'select(.ts | test("Z$|[+-][0-9]{2}:[0-9]{2}$") | not) | .ts' app.log | head -5
```

Expected: empty output means every timestamp carries timezone info. Any lines printed are malformed timestamps to fix at the source. Do not fix this in the SIEM; fix it where the log is written.

### 3. Add a correlation or request ID that crosses service boundaries

One ID, generated at the edge, passed through every service call. This is what turns five separate log lines in five services into one story.

```bash
jq -r 'select(.request_id=="[incident-request-id]") | [.ts, .service, .action, .result] | @tsv' app.log | sort
```

Expected: a single chronological story of one request across services. If your services each mint their own IDs, agree on one header name and propagate it; this is a one-line change per service with outsized payoff.

### 4. Log the security extras on auth and access events

For logins: auth method, MFA result, user agent, and failure reason (bad password vs unknown user vs locked). For access decisions: the policy or role evaluated and the allow or deny outcome. These are the fields that separate "someone logged in" from "someone is brute-forcing".

```bash
jq -c 'select(.event=="login" and .result=="failure") | {user, ip, method, reason}' app.log | head -5
```

Expected: each failed-login line names the account, the IP, the method, and why it failed. If your auth logs just say "login failed", extend them before the next incident.

### 5. Test with the five questions

Pick a real recent event and try to answer: when exactly, who, from where, what did they do, did it work. Time yourself; if it takes more than a few minutes, the fields are not doing their job.

```bash
printf 'when: \nwho: \nfrom where: \nwhat: \ndid it work: \n' | tee investigation-checklist.txt
```

Expected: five answers in under five minutes from logs alone, no SSH-ing into boxes. Whatever question you could not answer becomes the next logging ticket.

### Variant: you inherited logs you cannot change

Build the mapping in your SIEM: parse and rename fields on ingest, derive missing ones where possible (geo from IP, service from port), and document what cannot be answered. Then file the upstream fixes with the owning teams; ingest-time duct tape is not a strategy.

### Variant: high-volume services where every byte costs

Log the six core fields always; sample the verbose context. Keep full payloads for auth failures, privilege changes, and admin actions; those are rare and precious. Sampled-out context on a routine read is fine; sampled-out context on a privilege escalation is a hole.

### Variant: client-side and frontend events

Same six fields, adapted: timestamp, a stable user or session id, IP as seen by your backend, action, resource, result. Never trust client-reported identity alone; join it to the backend session.

## Why this happens

Developers log for debugging ("why did this crash"), which produces stack traces and variable dumps but no identity, no source, no result. Security needs a different cut: the who-from-where-what-worked record of every sensitive action. Nobody goes back to add these fields after the service ships, so investigations drown in logs that cannot answer the basic questions. Deciding the fields up front is cheap; reconstructing them after an incident is impossible.

## Edge cases and pitfalls

- PII in logs: IPs and user agents are useful and also personal data; know your retention and access rules before you centralize.
- Do not log secrets: passwords, tokens, and full auth headers must be redacted at write time, not filtered later.
- Clock skew between services makes "when" unreliable; NTP everywhere is a prerequisite, not a nice-to-have.
- Case and format drift ("user" vs "userId" vs "username") breaks queries; one canonical name per field, enforced by a shared schema or lint.
- Result values need a closed vocabulary: success, failure, denied. Free-text results ("kinda worked", "timeout-ish") cannot be counted or alerted on.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_FAc7zQs0BdmnrFRYh7gp_g
