## TL;DR
The OOM killer picks victims by score: every process gets one based on memory usage, and oom_score_adj lets you nudge it. Protect critical processes (SSH, monitoring agents) with negative adjustments, and make expendable workers (batch jobs, caches) more killable with positive ones. This does not prevent OOMs; it chooses who dies when they happen.

## The query
```text
Linux OOM killer score: how to protect critical processes
```

## Use this when
- The OOM killer killed the wrong process
- SSH or monitoring died during memory pressure
- Tuning which processes survive OOM events
- Designing services for memory-constrained hosts

## Not for when
- Fixing the memory leak or overcommit causing OOMs
- Container OOMKills (cgroup limits, related but separate)
- Swap tuning

## Steps

### Step 1: Read the current scores
Check each process's oom_score and oom_score_adj. The score combines memory usage with the adjustment; the highest score dies first. Know the baseline before changing anything.
Expected output: current scores for the important processes, showing who the killer would pick today.

### Step 2: Protect the must-survive processes
Set negative oom_score_adj on processes that must survive memory pressure: sshd, the monitoring agent, the orchestrator agent. These are the processes you need to diagnose and recover; losing them turns an OOM into a blind outage.
Expected output: critical infrastructure processes deprioritized as victims.

### Step 3: Mark the expendable processes
Set positive oom_score_adj on processes designed to die and restart: workers, caches, batch jobs. They should be the first victims, and they should come back cleanly via their supervisor.
Expected output: expendable processes listed as preferred victims, with restart behavior verified.

### Step 4: Never set -1000 casually
An adjustment of -1000 makes a process unkillable by the OOM killer, which can wedge the entire system (the kernel cannot free memory and everything stalls). Reserve it for nothing, or nearly nothing. Prefer modest negative values.
Expected output: no unkillable processes except by deliberate, documented exception.

### Step 5: Persist the settings
OOM adjustments set at runtime vanish on restart. Put them in the service's systemd unit, the container spec, or the startup script so they survive reboots and redeploys. An OOM policy that resets on reboot is not a policy.
Expected output: scores verified correct after a restart, not just after manual tuning.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_DxqhivsJ2bOLUhL2Kdysfg
