## TL;DR
Run the Vertical Pod Autoscaler in recommendation mode and it will tell you what each workload's requests and limits should be without touching a running pod. Install the VPA, let it observe for a few days, read the recommendations from the VPA objects, and apply them to your manifests yourself. Recommendation mode gives you the data with none of the restart risk.

## Error / query
```text
how to rightsize pod memory with VPA in recommendation mode
```

## Use this skill when
- You want data-driven memory sizing without auto-restarts
- Auditing a fleet for over-provisioned or under-provisioned pods
- Setting a requests/limits policy backed by measurements
- Validating hand-written resource blocks against reality

## Not for this skill when
- You want automatic in-place resizing (that is VPA auto mode, with restarts)
- You need horizontal scaling on load (use HPA)
- A pod is actively OOMKilling (fix the limit now, measure later)

## Steps

### Step 1: Install the VPA in recommendation mode
```bash
kubectl apply -f https://github.com/kubernetes/autoscaler/releases/download/vertical-pod-autoscaler-[version]/vpa-v1-crd-gen.yaml
```
Expected: VPA CRDs and components install. Then create VPA objects with `updatePolicy.updateMode: "Off"` so the recommender only observes and never restarts pods.

### Step 2: Create a VPA object per workload
```yaml
apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata:
  name: [workload]-vpa
  namespace: [namespace]
spec:
  targetRef:
    apiVersion: apps/v1
    kind: Deployment
    name: [deployment]
  updatePolicy:
    updateMode: "Off"
```
Expected: the VPA binds to the deployment and starts collecting usage history. No pod restarts happen in Off mode.

### Step 3: Wait, then read the recommendations
```bash
kubectl get vpa [workload]-vpa -n [namespace] -o jsonpath='{.status.recommendation.containerRecommendations}{"\n"}'
```
Expected: after 2-7 days of representative load, JSON with `target` (recommended request), `lowerBound`, and `upperBound` per container. The target is what VPA believes the request should be.

### Step 4: Apply recommendations to manifests with judgment
```bash
kubectl get vpa -A -o json | jq -r '.items[] | "\(.metadata.namespace)/\(.metadata.name): \(.status.recommendation.containerRecommendations[0].target.memory)"'
```
Expected: a fleet-wide list of recommended memory targets. Set requests near the target and limits above the upperBound, but sanity-check against known batch jobs and traffic peaks the observation window may have missed.

## Variant phrasings

### "vpa recommendation mode memory"
Steps 1-3. Off mode observes without restarting anything.

### "rightsize kubernetes pods automatically"
Recommendation mode advises; you decide. Auto mode restarts pods to apply, which most teams do not want for stateful or latency-sensitive workloads.

## Why it happens
Hand-set requests and limits drift from reality: they were guesses at deploy time, traffic changed, and nobody revisits them. VPA's recommender watches actual usage percentiles over days and computes targets from data, which beats both guesses and one-off `kubectl top` snapshots.

## Edge cases and pitfalls
- VPA needs days of representative data; recommendations from a quiet weekend will undersize Monday traffic.
- VPA and HPA on the same metric fight each other; do not run VPA auto mode on CPU while HPA scales on CPU.
- The recommender ignores the first hours of a pod's life; very short-lived pods get poor recommendations.
- Recommendations are per-container; sidecars with tiny footprints still get their own targets, do not copy the main container's numbers onto them.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_yNemIUA0PuRafiJijIW4sA
