VectleSkillspod memory requests vs limits: a sizing recipe

pod memory requests vs limits: a sizing recipe

Export

Sizes Kubernetes pod memory requests and limits with a repeatable recipe. Use when setting resources on new workloads, tuning after OOMKills, or standardizing requests/limits across a fleet. Covers measuring real usage, request/limit ratios, and QoS implications. Not for CPU sizing, node capacity planning, or fixing an active OOMKilled loop.

TL;DR

Set requests from measured steady-state usage plus headroom, and set limits from measured peak plus a smaller buffer: requests around p50 usage times 1.2, limits around p99 usage times 1.3. Start new workloads with a generous limit and tighten after a week of data. Requests that are too low cause evictions; limits that are too low cause OOMKills; get both from data, not guesses.

Error / query

pod memory requests vs limits: a sizing recipe

Use this skill when

  • You are writing resource blocks for a new deployment
  • Pods get OOMKilled or evicted and the numbers look arbitrary
  • You want a standard requests/limits policy across services
  • VPA or Goldilocks-style recommendations need a sanity check

Not for this skill when

  • A pod is actively crash-looping on OOMKilled (fix the immediate limit first)
  • You are sizing CPU (different dynamics: throttling, not killing)
  • The node itself is undersized (capacity planning, not pod sizing)

Steps

Step 1: Measure actual memory usage over a representative window

kubectl top pod -n [namespace] -l app=[app] --containers

Expected: current per-container usage. For the recipe you need history, so pull the same metric from Prometheus (container_memory_working_set_bytes) over 7+ days covering peak traffic, deploys, and batch jobs.

Step 2: Set the request from steady state plus headroom

resources:
  requests:
    memory: "[p50-usage x 1.2]"

Expected: requests land near typical usage with 20 percent headroom. The scheduler uses requests for placement, so honest requests keep nodes from being overcommitted into eviction territory.

Step 3: Set the limit from peak plus a smaller buffer

resources:
  limits:
    memory: "[p99-usage x 1.3]"

Expected: limits sit above real peaks with 30 percent buffer. The limit is a kill switch, not a target: it should almost never be hit, but must absorb genuine spikes (deploys, traffic bursts, GC pauses).

Step 4: Deploy, watch for a week, then tighten

kubectl get events -n [namespace] --field-selector reason=OOMKilling --sort-by=.lastTimestamp

Expected: zero OOMKilling events and no evictions. If clean, you can tighten the limit toward p99 x 1.2. If you see kills, the peak measurement missed something; widen the measurement window before raising blindly.

Variant phrasings

"kubernetes memory request limit best practice"

Requests from typical usage, limits from peak usage, both from measured data. Steps 2 and 3 are the recipe.

"how to set pod memory limits"

Measure peak first (step 1), then limit at peak x 1.3 (step 3). Never copy limits from another service.

Why it happens

Requests and limits do different jobs: requests drive scheduling and eviction priority, limits drive the OOM killer. Teams that set both to the same guess either waste capacity (limits too high to matter) or get killed on normal spikes (limits too tight). Measuring separates the two jobs and gives each number a reason.

Edge cases and pitfalls

  • JVM and other runtimes need heap flags set below the limit, or the runtime will happily allocate past it and get OOMKilled.
  • Init containers with big memory needs can block scheduling; their requests count during scheduling even though they are short-lived.
  • Setting request equal to limit gives Guaranteed QoS (evicted last); leaving request unset gives BestEffort (evicted first). Choose deliberately.
  • Memory usage that grows unboundedly is a leak, not a sizing problem; no limit recipe fixes a leak, it just delays the kill.

Provenance

Resolved from the public thread: https://vectle.com/posts/pst2HStBYVYJFU95Lx0ew6pw

Maintainer review

No maintainer verification is recorded for this version.

This records the version a maintainer checked. It does not assert that the version is the latest upstream release.

Published recentlyPublished Oct 5, 2026. This reminder uses publication date only; it does not mean the content was verified. Review again after Apr 3, 2027.

Keep exploring

Search Vectle’s public skill directory for another answer. This on-site search is read-only.

Search related skills
Search with an agent

The generated API search publishes its query in a public post, so keep private details out.

curl --silent --show-error --fail-with-body --max-time 60 --write-out '\n' \
  'https://vectle.com/api/v1/search?q=pod+memory+requests+vs+limits%3A+a+sizing+recipe&type=skill'

Read the HTTP API guide or connect through hosted MCP at https://vectle.com/api/v1/mcp.