## TL;DR
Dont try to make staging behave like prod by guessing. Write down the exact production context for one failing case, list every way prod differs from staging, and eliminate the differences one at a time until the bug shows up. Production-only bugs are always a difference you havent found yet, and the repro runbook below is how you find it.

## The query

```text
how to reproduce a bug that only happens in production
```

## Use this when

- A ticket comes back "cannot reproduce" but customers keep hitting it
- Staging is green and production is red for the same code
- An agent needs a repro before engineering will take the escalation
- The bug is intermittent and you suspect data, timing, or config

## Not for

- Bugs you can already reproduce locally
- Performance or load testing
- Deciding whether a bug is worth fixing
- Writing the fix itself

## Steps

### 1. Freeze one failing case in writing

Pick a single incident, not the general complaint. Record the account, plan tier, request path, timestamp, region, feature flag states, and the exact error text. One case gives you something to chase; the general complaint gives you nothing.

Expected output: a one-paragraph incident snapshot at the top of the ticket.

### 2. List every way prod differs from staging

Data volume and shape, real user input vs fixtures, flag values, third-party responses, cron timing, traffic mix, cache warmth, region. Most production-only bugs live in this list.

Expected output: a written difference list with at least five entries.

### 3. Replay with production-like input, not your fixtures

Rebuild the failing case using an anonymized production record or request payload for that exact account. Fixtures are built from assumptions; production input carries the thing your assumptions missed.

Expected output: a replay attempt using real-shaped input, logged with its result.

### 4. Instrument the path instead of guessing

Add temporary logging or tracing around the failing path in a staging or canary environment: what value entered, what branch ran, what the third party returned. Remove it after. One instrumented run beats ten theories.

Expected output: log lines that show the exact divergence point.

### 5. Mirror prod config in your repro environment

Set the same flags, the same limits, the same integrations to the same sandbox endpoints. A bug that needs flag X on and data shape Y will never appear in an environment that has neither.

Expected output: a checklist showing which prod settings are now mirrored.

### 6. Lock the repro into a regression test

Once it reproduces, write the automated test before the fix. The test proves the repro and protects it forever. A repro that only exists in a chat thread will regress within a year.

Expected output: a failing test in the suite that captures this exact case.

## Repro runbook template

```text
Bug: [one-line summary]
First seen: [date, ticket links]
Failing case: [account or identifier, plan, region, timestamp]
Error text: [paste the exact error]
Prod context: [flag states, data shape, third party involved]
Differences from staging:
1. [e.g. real user payloads, not fixtures]
2. [e.g. flag checkout_v2 is ON in prod]
3. [e.g. cron runs every 5 min in prod, hourly in staging]
Hypotheses (testable, one at a time):
1. [hypothesis], test it by: [how you will test it]
2. [hypothesis], test it by: [how you will test it]
Repro attempts:
- [date] [what you tried] then [result]
Repro confirmed: [yes/no, date, test location]
```

## Variant phrasings

### cant reproduce a bug that only happens in prod

Steps 1 through 3 are the whole answer. Support agents usually skip the failing-case snapshot and jump to theorizing; the snapshot is what makes the rest work.

### how to debug production-only errors as a support agent

Your job is steps 1 and 2: freeze the case and list the differences. That alone turns an escalation engineering ignores into one they can act on.

### staging works but production fails

That sentence is the difference list in disguise. Write the list down instead of saying it.

## Why it happens

Production differs from every other environment in dozens of small ways at once: real input shapes, real volumes, real timing, real third-party behavior, flags that were flipped for one cohort. Locally reproducible bugs are usually logic errors; production-only bugs are usually context errors, the code is fine in a context that never exists in prod. Reproducing is just hunting the context difference.

## Edge cases

- Bugs that need real user input: anonymize carefully. Keep the shape of the input (length, encoding, nesting) and strip the content. Shape is usually what matters.
- Time-based bugs: check cron schedules, timezone handling, and daylight-saving transitions. These reproduce only when the clock cooperates.
- Third-party flakiness: record the exact third-party response in the failing case. If you cant capture it, mock the recorded failure response in staging.
- Multi-tenant bugs: the failing case may need the account's full state (plan history, entitlements), not just the failing request.
- One-off prod incidents: sometimes you truly cant reproduce. Say so explicitly, ship the monitoring from step 4 permanently, and close the loop on the next occurrence.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_L18f3KfPuKR0xWyaHYM_Kw
