rollback vs rollforward: decision framework
Offers a decision framework for rollback vs rollforward after a bad deploy: rollback when the bad version is clear and reversible, roll forward when the fix is smaller or data cannot be un-written. Use when a deploy caused the incident and the team is debating. Does not cover non-deploy incidents.
TL;DR
Rollback when you can reverse the deploy in minutes and the bad version is clearly identified. Roll forward when the fix is a one-line change you can ship faster than a revert, or when the deploy included migrations that cant be cleanly undone. The default is rollback, because reverting restores a state you already tested in production. Decide with a 5-minute timer, not a debate.
Error / query
rollback vs rollforward: decision frameworkA deploy caused the incident and the team is split: half want to revert, half want to hotfix forward. You need a decision in minutes.
Use this skill when
- a deploy caused the incident and the team is arguing about revert vs fix
- you need decision criteria written down before the next bad deploy
- a rollback is being considered but the deploy included data migrations
- you want a default policy the on-call can apply without a meeting
Not for this skill when
- the incident was not caused by a deploy (this framework doesnt apply)
- you need data recovery or backup restore procedures
- the question is feature flag strategy (related, but a separate decision)
Steps
1. Confirm the deploy is actually the cause (5-minute check)
Find the most recent deploy and line it up with the incident start. If they dont line up, stop: this is not a deploy problem.
git log --oneline -10
kubectl rollout history deployment/[app-name] 2>/dev/null | tail -5Expected: you can point at a specific deploy whose timestamp matches the incident start. If nothing lines up, work the incident normally instead.
2. Score rollback vs rollforward in under 2 minutes
Write the decision down with a deadline. The checklist forces the tradeoffs into the open instead of a circular argument.
cat > /tmp/rollback-decision.txt <<'EOF'
ROLLBACK if: bad version clearly identified, revert takes under 10 min, no irreversible migrations in the deploy
ROLLFORWARD if: fix is a small known change, revert would break data written since deploy, hotfix ships faster than revert
DECIDE BY: [timestamp, 5 minutes from now]
EOF
cat /tmp/rollback-decision.txtExpected: a written decision with a deadline. No open-ended debate; the on-call picks one and executes.
3. Execute and verify on the dashboard
Run the rollback (or push the hotfix), then watch the error rate, not the deploy logs. The incident is over when the metrics say so.
kubectl rollout undo deployment/[app-name]
kubectl rollout status deployment/[app-name] --timeout=120sExpected: rollout status reports success, and the error rate drops on the dashboard within minutes. If it doesnt, the deploy wasnt the whole story.
Variant phrasings
when to roll back a deployment
Same fix: roll back when the bad version is clear and the revert is fast. Thats the default branch of the decision in step 2.
is it better to roll back or fix forward
Same fix: it depends on reversibility and speed, which is exactly what the step 2 checklist scores. There is no universal answer, only the faster safe path.
how to undo a bad release
Same fix: confirm the deploy caused it (step 1), pick rollback or rollforward (step 2), execute and verify (step 3).
Why it happens
Teams freeze because both options feel risky and nobody wants to own the call. Defaulting to rollback works because reverting restores a known-good state that already ran in production. Rolling forward is a bet that your untested fix is correct under pressure, which is a worse bet more often than people admit.
Edge cases and pitfalls
- Database migrations shipped in the deploy: rollback may need a forward migration, not a revert. Plan for this before you need it.
- Canary deploy gone bad: just shift traffic back to the stable version; a full rollback is overkill.
- No rollback path ever tested: rehearse reverts in staging first. An untested rollback is its own incident.
- Hotfix rolled forward but made things worse: timebox the attempt, then roll back. Two failed forwards in a row means stop forwarding.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_0moXUaLSxQPvftGRRgUNgQ
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.