how to read Kubernetes events for root cause
Reads Kubernetes events to find root causes of pod and cluster problems. Use when you need the timeline of what happened, events are the fastest signal, or describe output is too narrow. Triggers: kubectl get events, sorting and filtering events, event spam, finding root cause from warnings. Not for: application log analysis (use container logs), long-term history (events expire), metrics-based investigation.
TL;DR
Kubernetes events are the cluster's own timeline: kubectl get events --sort-by shows what happened and when, newest last. Filter by namespace and warning type to cut the noise, and correlate the timestamps with your incident. Most root causes show up as a Warning event minutes before everything else broke.
The query
how to read Kubernetes events for root causeEvents expire after an hour by default, so read them during the incident, not after.
Use this when
- something broke and you need the timeline
kubectl describeoutput is too narrow- you want warnings across a whole namespace at once
- correlating pod, node, and scheduler activity
Not for
- application logs (thats
kubectl logs) - history older than the event TTL (ship events to a collector for that)
- performance metrics (use monitoring, not events)
Steps
- Get the timeline, newest last:
kubectl get events -n [namespace] --sort-by=.metadata.creationTimestampExpected: chronological list. Read bottom-up during an incident: the last Warning before things broke is usually the cause, everything after is fallout.
- Cut the noise to warnings only:
kubectl get events -n [namespace] --sort-by=.metadata.creationTimestamp | grep -v NormalExpected: only Warning events remain. Normal events (Scheduled, Pulled, Created) are the happy path; warnings are where causes live.
- Scope to the object you care about:
kubectl get events -n [namespace] --field-selector involvedObject.name=[pod-name]Expected: only events for that pod. Also useful: involvedObject.kind=Node for node-level issues, or no selector for the full namespace picture.
- Watch events live while you reproduce:
kubectl get events -n [namespace] --sort-by=.metadata.creationTimestamp -wExpected: new events stream in as they happen. Trigger the failing action (reapply, restart) and watch what the cluster says in real time.
- Read the reason strings; theyre a controlled vocabulary:
- FailedScheduling: scheduler couldnt place the pod (read the detail)
- FailedMount / FailedAttachVolume: storage problems
- Unhealthy: probe failures (liveness/readiness)
- BackOff: container crashing repeatedly
- Evicted: kubelet reclaimed the pod (resource pressure)
- OOMKilling: memory limit hit
Expected: the reason names the subsystem; the message names the specifics. Start from the reason, not the message.
- Note the count and age columns:
kubectl get events -n [namespace] --sort-by=.metadata.creationTimestamp | tail -20Expected: high COUNT with recent age means its happening right now and repeatedly. A single old event is history, not the cause.
Variant: events are empty but things are clearly broken
Events expired (default TTL 1 hour) or the namespace is wrong. Check -A for all namespaces, and check whether something is deleting events.
Variant: one object spams thousands of events
A crashlooping pod or flapping probe floods the stream. Filter it out with grep -v on the name, or fix the underlying loop.
Variant: need events from before the TTL window
Too late for kubectl. Long term, ship events with an event exporter to your logging stack; thats the durable fix for next time.
Why it happens
Controllers and kubelets emit events for everything they do: scheduling, pulling, probing, killing, evicting. Theyre the highest-signal, lowest-effort diagnostic in the cluster because theyre already correlated by object and timestamp. Most incidents have their cause sitting in the event stream; people just dont look there first.
Edge cases
- Event TTL is 1 hour by default; increase
--event-ttlon the API server if you need longer windows. - Duplicate events get aggregated with COUNT, so a count of 500 means one problem happening 500 times, not 500 problems.
- Some managed clusters restrict event access; check RBAC if
get eventsreturns Forbidden. - Events dont include application logs; a pod can log errors for hours with zero Warning events if the process stays alive.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_okbpUtnokY1-s2CIGtG8gg
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.