how to find which process is eating CPU on Linux
Shows how to find which process is eating CPU on Linux: top for the overview, ps sorted by CPU, pidstat for the user vs system split. Use when load is high and the culprit is unknown. Not when CPU is idle (check disk or memory) or for profiling a known process.
TL;DR
Start with top to see the hog, then use ps sorted by CPU to name the exact process, and pidstat to see whether it is user CPU (your code) or system CPU (kernel work). Once you know which process and which kind of CPU, you know whether to look at the app, the disk, or the network. Do not kill the top process reflexively; identify first, act second.
Error / query
how to find which process is eating CPU on LinuxUse this skill when
- Load average is high and you do not know why
- An alert says CPU is pegged on a host
- A deploy made a box slow and you need the culprit process
- You need to tell user-space burn from kernel/system CPU burn
Not for this skill when
- The box is slow but CPU is idle; check disk IO or memory instead
- You already know the process; use a profiler on it directly
- It is a container host; check per-container stats first, the host view misleads
Steps
Step 1: See the overall picture and the top CPU consumer
top -b -n 1 | head -15Expected: the %CPU column shows the hog at the top; note the PID and whether the load is in us (user) or sy (system).
Step 2: List processes sorted by CPU with full command lines
ps -eo pid,ppid,%cpu,%mem,comm,args --sort=-%cpu | head -8Expected: the exact process and its arguments; the args column often reveals which worker or script it is.
Step 3: Check whether it is user or system CPU over time
pidstat -p [pid-from-step-2] 1 3Expected: %usr vs %system columns; high %usr means the app is burning CPU, high %system means syscalls (IO, locks, network).
Step 4: See what the process is doing right now
cat /proc/[pid-from-step-2]/status | grep -E 'State|Threads'; ls /proc/[pid-from-step-2]/task | wc -lExpected: state R (running) plus the thread count; hundreds of threads all in R points at a thread-pool or fork bomb.
Step 5: Check which container it belongs to, if any
cat /proc/[pid-from-step-2]/cgroup | head -3Expected: a cgroup path naming the container or slice; on shared hosts the hog is often a neighbor container, not your service.
Variant phrasings
"High load average but CPU looks idle"
Uninterruptible sleep (disk IO wait); use iostat, not CPU tools, the processes are stuck on disk.
"Which thread inside the process is hot"
pidstat -t -p [pid] breaks CPU down per thread; then map the hot thread to a stack with the app's profiler.
"CPU spike every hour on the dot"
A cron job; check the spike time against crontabs and systemd timers before blaming the app.
Why it happens
Something is either computing hard (user CPU: hot loop, expensive query, crypto), syscalling hard (system CPU: lock contention, excessive IO), or there is simply more work than cores. The us/sy split in step 1 tells you which world you are in within seconds.
Edge cases and pitfalls
- Short-lived processes vanish before you catch them; use pidstat in continuous mode or an always-on profiler.
- Steal time on VMs looks like your process burning CPU; check %st in top, if it is high the hypervisor is the problem.
- A process at 100% on a 32-core box is one thread; check per-core view (press 1 in top) before declaring an emergency.
- Do not nice or kill production processes to "test"; identify, then fix the cause (bad deploy, runaway job, noisy neighbor).
Provenance
Resolved from the public thread: https://vectle.com/posts/pst__zkBIkTRe7KP-7732YerZw