## TL;DR

Rollback when you can reverse the deploy in minutes and the bad version is clearly identified. Roll forward when the fix is a one-line change you can ship faster than a revert, or when the deploy included migrations that cant be cleanly undone. The default is rollback, because reverting restores a state you already tested in production. Decide with a 5-minute timer, not a debate.

## Error / query

```text
rollback vs rollforward: decision framework
```

A deploy caused the incident and the team is split: half want to revert, half want to hotfix forward. You need a decision in minutes.

## Use this skill when

- a deploy caused the incident and the team is arguing about revert vs fix
- you need decision criteria written down before the next bad deploy
- a rollback is being considered but the deploy included data migrations
- you want a default policy the on-call can apply without a meeting

## Not for this skill when

- the incident was not caused by a deploy (this framework doesnt apply)
- you need data recovery or backup restore procedures
- the question is feature flag strategy (related, but a separate decision)

## Steps

### 1. Confirm the deploy is actually the cause (5-minute check)

Find the most recent deploy and line it up with the incident start. If they dont line up, stop: this is not a deploy problem.

```bash
git log --oneline -10
kubectl rollout history deployment/[app-name] 2>/dev/null | tail -5
```

Expected: you can point at a specific deploy whose timestamp matches the incident start. If nothing lines up, work the incident normally instead.

### 2. Score rollback vs rollforward in under 2 minutes

Write the decision down with a deadline. The checklist forces the tradeoffs into the open instead of a circular argument.

```bash
cat > /tmp/rollback-decision.txt <<'EOF'
ROLLBACK if: bad version clearly identified, revert takes under 10 min, no irreversible migrations in the deploy
ROLLFORWARD if: fix is a small known change, revert would break data written since deploy, hotfix ships faster than revert
DECIDE BY: [timestamp, 5 minutes from now]
EOF
cat /tmp/rollback-decision.txt
```

Expected: a written decision with a deadline. No open-ended debate; the on-call picks one and executes.

### 3. Execute and verify on the dashboard

Run the rollback (or push the hotfix), then watch the error rate, not the deploy logs. The incident is over when the metrics say so.

```bash
kubectl rollout undo deployment/[app-name]
kubectl rollout status deployment/[app-name] --timeout=120s
```

Expected: rollout status reports success, and the error rate drops on the dashboard within minutes. If it doesnt, the deploy wasnt the whole story.

## Variant phrasings

### when to roll back a deployment

Same fix: roll back when the bad version is clear and the revert is fast. Thats the default branch of the decision in step 2.

### is it better to roll back or fix forward

Same fix: it depends on reversibility and speed, which is exactly what the step 2 checklist scores. There is no universal answer, only the faster safe path.

### how to undo a bad release

Same fix: confirm the deploy caused it (step 1), pick rollback or rollforward (step 2), execute and verify (step 3).

## Why it happens

Teams freeze because both options feel risky and nobody wants to own the call. Defaulting to rollback works because reverting restores a known-good state that already ran in production. Rolling forward is a bet that your untested fix is correct under pressure, which is a worse bet more often than people admit.

## Edge cases and pitfalls

- Database migrations shipped in the deploy: rollback may need a forward migration, not a revert. Plan for this before you need it.
- Canary deploy gone bad: just shift traffic back to the stable version; a full rollback is overkill.
- No rollback path ever tested: rehearse reverts in staging first. An untested rollback is its own incident.
- Hotfix rolled forward but made things worse: timebox the attempt, then roll back. Two failed forwards in a row means stop forwarding.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_0moXUaLSxQPvftGRRgUNgQ
