pod memory requests vs limits: a sizing recipe
Sizes Kubernetes pod memory requests and limits with a repeatable recipe. Use when setting resources on new workloads, tuning after OOMKills, or standardizing requests/limits across a fleet. Covers measuring real usage, request/limit ratios, and QoS implications. Not for CPU sizing, node capacity planning, or fixing an active OOMKilled loop.
TL;DR
Set requests from measured steady-state usage plus headroom, and set limits from measured peak plus a smaller buffer: requests around p50 usage times 1.2, limits around p99 usage times 1.3. Start new workloads with a generous limit and tighten after a week of data. Requests that are too low cause evictions; limits that are too low cause OOMKills; get both from data, not guesses.
Error / query
pod memory requests vs limits: a sizing recipeUse this skill when
- You are writing resource blocks for a new deployment
- Pods get OOMKilled or evicted and the numbers look arbitrary
- You want a standard requests/limits policy across services
- VPA or Goldilocks-style recommendations need a sanity check
Not for this skill when
- A pod is actively crash-looping on OOMKilled (fix the immediate limit first)
- You are sizing CPU (different dynamics: throttling, not killing)
- The node itself is undersized (capacity planning, not pod sizing)
Steps
Step 1: Measure actual memory usage over a representative window
kubectl top pod -n [namespace] -l app=[app] --containersExpected: current per-container usage. For the recipe you need history, so pull the same metric from Prometheus (container_memory_working_set_bytes) over 7+ days covering peak traffic, deploys, and batch jobs.
Step 2: Set the request from steady state plus headroom
resources:
requests:
memory: "[p50-usage x 1.2]"Expected: requests land near typical usage with 20 percent headroom. The scheduler uses requests for placement, so honest requests keep nodes from being overcommitted into eviction territory.
Step 3: Set the limit from peak plus a smaller buffer
resources:
limits:
memory: "[p99-usage x 1.3]"Expected: limits sit above real peaks with 30 percent buffer. The limit is a kill switch, not a target: it should almost never be hit, but must absorb genuine spikes (deploys, traffic bursts, GC pauses).
Step 4: Deploy, watch for a week, then tighten
kubectl get events -n [namespace] --field-selector reason=OOMKilling --sort-by=.lastTimestampExpected: zero OOMKilling events and no evictions. If clean, you can tighten the limit toward p99 x 1.2. If you see kills, the peak measurement missed something; widen the measurement window before raising blindly.
Variant phrasings
"kubernetes memory request limit best practice"
Requests from typical usage, limits from peak usage, both from measured data. Steps 2 and 3 are the recipe.
"how to set pod memory limits"
Measure peak first (step 1), then limit at peak x 1.3 (step 3). Never copy limits from another service.
Why it happens
Requests and limits do different jobs: requests drive scheduling and eviction priority, limits drive the OOM killer. Teams that set both to the same guess either waste capacity (limits too high to matter) or get killed on normal spikes (limits too tight). Measuring separates the two jobs and gives each number a reason.
Edge cases and pitfalls
- JVM and other runtimes need heap flags set below the limit, or the runtime will happily allocate past it and get OOMKilled.
- Init containers with big memory needs can block scheduling; their requests count during scheduling even though they are short-lived.
- Setting request equal to limit gives Guaranteed QoS (evicted last); leaving request unset gives BestEffort (evicted first). Choose deliberately.
- Memory usage that grows unboundedly is a leak, not a sizing problem; no limit recipe fixes a leak, it just delays the kill.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst2HStBYVYJFU95Lx0ew6pw
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.