## TL;DR
rollout restart redeploys the current revision (new pods, same config): use it to bounce pods, pick up ConfigMap or secret changes, or clear wedged state. rollout undo returns to the previous revision: use it when the current revision itself is bad. Restart never fixes a bad image; undo never picks up a config change. Picking the wrong one wastes the incident.

## The query
```text
kubectl rollout restart vs rollout undo: when to use each
```

## Use this when
- Recovering from a bad deployment
- Pods need bouncing without config changes
- Rolling back a bad image or bad code
- Teaching deployment recovery patterns

## Not for when
- First-time deployment failures (nothing to roll back to)
- StatefulSet or DaemonSet recovery (similar but check specifics)
- Database migrations (need their own rollback story)

## Steps

### Step 1: Diagnose: is the revision bad or just stale
Check whether the running revision is broken (bad image, bad code) or merely stale (needs a config refresh, pods wedged). Bad revision means undo; stale-but-good revision means restart.
Expected output: the situation classified correctly before acting.

### Step 2: Use rollout restart for refresh scenarios
When the deployment spec is fine but pods need recreating (config change via env vars, cleared caches, recovered nodes), rollout restart creates new pods from the current spec with zero config change.
Expected output: fresh pods on the same revision, problem cleared.

### Step 3: Use rollout undo for bad revisions
When the current revision is broken, rollout undo moves back to the previous ReplicaSet. Check the rollout history first to confirm what the previous revision was; undo is only as good as your history.
Expected output: the previous known-good revision serving traffic.

### Step 4: Verify the recovery, not just the command
After either command, watch the rollout status and verify the application actually recovered (health checks, error rates). A completed rollout that serves errors is not a recovery.
Expected output: application healthy on the intended revision.

### Step 5: Record what you did for the postmortem
Note which command you ran and why. Incident reviews need to know whether the fix was a restart (transient issue) or an undo (bad revision); they imply very different follow-ups.
Expected output: the recovery action documented with its rationale.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_NIOslQPK8Aibtb4sqzQD6g
