out of memory" killer: dmesg reading guide
Reads Linux dmesg output to diagnose OOM-killer events: find the kill entries, read the full report block, check the victim's oom_score, correlate with container limits, and track pressure over time. Use when a process dies with exit 137 or a container restarts OOMKilled. Not for app-level heap errors, full machine hangs, or preemptive heap tuning.
TL;DR
When the kernel kills a process for memory, it logs exactly why in dmesg: which process, how much it used, and the oom_score that picked it. Grep for the oom-killer lines, read the victim's memory numbers, then decide whether the app leaked, the limit was too tight, or the box is just undersized.
Error / query
"out of memory" killer: dmesg reading guideUse this skill when
- a process dies and the logs say "Killed" or exit code 137
- a container restarts with OOMKilled status
- the box gets slow then a service disappears with no crash log
- you need to prove to a vendor that the kernel killed their app, not the other way around
Not for this skill when
- the app logs its own out-of-memory error (check heap settings, not the kernel)
- the whole machine hangs (that is thrashing or a different panic)
- you are tuning JVM or Go heap sizes preemptively
Steps
Step 1: Find the oom-killer events in the kernel log
dmesg | grep -i -E "oom-killer|killed process" | tail -10Expected: lines like "Out of memory: Killed process [PID] ([NAME])". No matches means the kernel did not kill anything recently and your 137 came from somewhere else (a container limit or a manual kill).
Step 2: Read the full oom-killer report block
dmesg | grep -B5 -A40 "oom-killer" | tail -60Expected: a block starting with "oom-killer:" showing total memory, free memory, and a per-process table with rss values. The "Killed process" line names the victim; the table shows what everything else was using at that moment.
Step 3: Check the victim's oom_score to see why IT was picked
cat /proc/[PID]/oom_score || echo "process gone, check the dmesg table"
dmesg | grep -A30 "oom-killer" | grep -E "rss|oom_score_adj" | tail -15Expected: the score and adjustment. The killer picks the highest score, which is mostly about rss size plus oomscoreadj. A victim with a modest rss but a high adj was deliberately made killable; a victim with huge rss was just the biggest hog.
Step 4: Correlate with container or cgroup limits
kubectl get pod [POD] -n [NAMESPACE] -o jsonpath="{.status.containerStatuses[0].lastState.terminated.reason}"
cat /sys/fs/cgroup/memory.max || cat /sys/fs/cgroup/memory/memory.limit_in_bytesExpected: reason "OOMKilled" if a container limit fired, and the cgroup limit in bytes. Container OOMs kill without a host dmesg entry sometimes, so the pod status is the source of truth there.
Step 5: Check whether memory pressure was building over time
journalctl -k --since "24 hours ago" | grep -i -c -E "oom|memory.*pressure|killed process"
free -hExpected: a count of memory events over the day and current free memory. One isolated kill after months of quiet is a spike or leak onset; repeated kills daily is chronic undersizing or a steady leak.
Variant phrasings
"container keeps getting OOMKilled but the app looks fine"
The container limit is below the app's real working set. Compare the limit against the app's peak rss from metrics, then raise the limit or lower the heap; do not just keep restarting it.
"dmesg shows oom-killer but i dont know which process caused it"
The killer kills the biggest scorer, not necessarily the root cause. Read the per-process table in step 2; ten medium processes adding up to exhaustion is a capacity problem, not one bad app.
"exit code 137 with no dmesg entry"
Check the container runtime and the orchestrator: Kubernetes OOMKilled, Docker memory limits, and systemd MemoryMax all produce 137s with their own logs instead of host dmesg.
Why it happens
The kernel overcommits memory and when it truly runs out, the oom-killer sacrifices one process to save the rest. It scores candidates by memory use adjusted by oomscoreadj, so the victim is the "cheapest" big process to kill, which is often an innocent app that happened to be large while the real culprit is a slow leak elsewhere.
Edge cases and pitfalls
- dmesg is a ring buffer; on busy boxes the oom block scrolls away fast, so ship kernel logs to your log aggregator with persistent journald.
- oomscoreadj of -1000 makes a process unkillable, which just moves the kill to the next victim; use it sparingly and deliberately.
- Memory cgroup v1 vs v2 paths differ; the step 4 command tries both, but know which one your distro uses.
- Swap being enabled changes oom behavior; a box thrashing on swap looks slow long before the killer fires.
- A kill during a deploy may be the new version's higher footprint, not a leak; compare rss before and after the release.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_DvG7umAtcFJA1OoZXsLhmw
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.