## TL;DR

An OOMKilled pod means a container exceeded its memory limit (or the node ran dry). Check `kubectl describe pod` for the last state, find the culprit with memory usage stats, then fix by raising limits or fixing the leak. The kill almost always comes from a limit you set, not the node itself.

## The error

```text
Last State: Terminated
Reason: OOMKilled
Exit Code: 137
```

Also shows up as: `kubectl get pods` shows `OOMKilled` under STATUS, or events say `Out of memory: Kill process` from the kernel.

## Use this when

- pod status is OOMKilled or exit code 137
- a pod restarts repeatedly under load
- you suspect a memory leak in a container
- requests/limits were set long ago and never revisited

## Not for

- pods killed by liveness probes (different reason string)
- node-level memory pressure where the kubelet evicts pods (check node events instead)
- CPU problems; memory and CPU limits behave very differently

## Steps

1. Confirm the OOMKilled and which container died:

```bash
kubectl describe pod [pod-name] -n [namespace]
```

Expected: under Containers, `Last State: Terminated, Reason: OOMKilled, Exit Code: 137`. Note the container name and the memory request/limit in the same block.

2. Check how the pod is behaving over time:

```bash
kubectl get pod [pod-name] -n [namespace] -w
kubectl top pod [pod-name] -n [namespace]
```

Expected: `kubectl top` shows current memory vs the limit. If usage climbs steadily toward the limit before each restart, its a leak or growth, not a one-off spike.

3. See if the node itself is under pressure (rules out node-level kills):

```bash
kubectl describe node [node-name] | grep -A 5 Conditions
kubectl get events -n [namespace] --sort-by=.metadata.creationTimestamp | grep -i memory
```

Expected: node Conditions show MemoryPressure=False if the kill came from the container limit. Events will say which limit was exceeded.

4. Check for a leak pattern inside the app:

```bash
kubectl logs [pod-name] -n [namespace] --previous | tail -100
```

Expected: look for heap dumps, GC errors, or the app logging its own memory stats right before dying. Repeat OOM kills right after startup point at a startup burst; kills after hours point at a slow leak.

5. Fix it. Pick one:

```yaml
resources:
  requests:
    memory: "512Mi"
  limits:
    memory: "1Gi"
```

Raise the limit to something sane (a headroom of ~20 percent over observed peak), set a request so the scheduler places the pod on a node with room, and deploy. If the app leaks, fix the leak in code instead of endlessly raising the limit.

6. Apply and verify the pod stays up:

```bash
kubectl apply -f deployment.yaml
kubectl rollout status deployment/[name] -n [namespace]
```

Expected: rollout completes, new pods show STATUS Running, no OOMKilled in describe output after a full load cycle.

### Variant: OOMKilled but the container limit looks fine

The node ran out of memory and the kernel picked your pod. Check node MemoryPressure, lower total commitments on the node, or add requests so the scheduler stops overpacking.

### Variant: exit code 137 without OOMKilled in the reason

Same fix path. 137 = SIGKILL, which in Kubernetes is almost always the OOM killer. Confirm via `kubectl describe pod`.

### Variant: Java app OOMKilled with plenty of headroom in the limit

JVM heap defaults plus off-heap can exceed the container limit. Set `-Xmx` to about 75 percent of the container limit so heap plus overhead fits.

## Why it happens

Kubernetes enforces memory limits with cgroups. When a container touches more memory than its limit, the kernel OOM killer terminates it instantly, no graceful shutdown. Exit code 137 is the signature. Most of the time the limit was set too low for real workload, or the app has a slow leak nobody noticed because restarts masked it.

## Edge cases

- No limit set at all means the container can eat the node and the kernel picks victims by oom_score; always set limits.
- Bursty workloads (batch jobs, image processing) need limits far above the request; use a VerticalPodAutoscaler in recommendation mode to find real numbers.
- OOMKilled during init or startup often means the request is fine but the limit is below the startup peak; watch `kubectl top` during a fresh deploy.
- If many pods die at once, suspect node pressure or a bad deploy, not individual leaks.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_tbJdavdXVNdMEZhjCgDLMA
