agent's 'realistic traffic' replay hammered the payment webhook endpoint and triggered duplicate-charge alerts
Recovery and prevention guide for traffic replays that re-trigger side-effecting endpoints, such as payment webhooks, causing duplicate charges. Use after a replay caused real-world side effects. Shows how to assess the damage, stub or sandbox side-effecting endpoints, and classify replay targets by side effect before the next run.
TL;DR
Replaying production traffic replays its side effects too. A byte-faithful replay of webhook calls will charge customers again. Stop the replay, assess what actually executed, then rebuild the replay target so payment-affecting endpoints hit a stub or the provider's sandbox. Never replay side-effecting traffic against live endpoints.
The query
agent's 'realistic traffic' replay hammered the payment webhook endpoint and triggered duplicate-charge alertsSteps
1. Stop the replay immediately
Halt the replay run. Every additional replayed request is another potential duplicate charge.
Expected: no further webhook calls leaving the replay harness.
2. Assess what actually executed
Check the payment provider's dashboard for charges created during the replay window. Compare against the replay's request log to identify duplicates versus legitimate charges.
Expected: a concrete list of duplicate charges attributable to the replay.
3. Remediate the duplicates
Void or refund the duplicate charges through the provider's normal process, and notify affected customers if your policy requires it. Record the incident for the postmortem.
Expected: duplicates reversed, customers made whole, incident documented.
4. Rebuild the replay target with stubs
Reconfigure the replay environment so side-effecting endpoints cannot reach production: route webhook receivers to a stub that returns success without charging, point the application at the payment provider's sandbox, or strip webhook calls from the replay set entirely.
Expected: a replay configuration where the payment path is provably inert.
5. Classify endpoints by side effect before every replay
Add a mandatory step to the replay workflow: list every endpoint in the replay set, mark which ones cause external side effects (charges, emails, messages, third-party writes), and require stub or sandbox confirmation for each marked endpoint before the run starts.
Expected: the replay cannot start until every side-effecting endpoint has a documented stub or sandbox.
Use this when
- A traffic replay triggered real charges, emails, or third-party writes
- The replay set includes webhook or callback endpoints
- The agent treated "realistic traffic" as safe because it was historical
- You need a side-effect classification step for replays
Not for this skill when
- The replay only hit read-only endpoints (safe by construction)
- The damage came from load volume rather than side effects (pool sizing issue)
- Webhook signature verification is the thing being tested (test that deliberately, in sandbox)
- The duplicate charges came from application retry logic, not a replay
Variant phrasings
load test caused duplicate charges
Same root cause family: the test exercised a charge path against a live provider. Sandbox it.
traffic replay triggered webhooks
Replays are byte-faithful; webhooks fire again. Stub the receiver.
replay re-charged customers
Assess, refund, then rebuild the replay target so the charge path is inert.
Why it happens
"Realistic traffic" sounds like a virtue: the more faithful the replay, the better the test. But fidelity includes the parts that touch the real world. The agent optimized for realism and never separated read traffic from write traffic, because its model of a replay was "send the same requests" with no concept of requests that do things. Historical traffic is only safe to replay when its effects are historical too.
Edge cases
- Idempotency keys: replaying the original keys may suppress duplicates in one run and mask the problem, while fresh keys in the next run double-charge. Do not rely on idempotency to make replays safe; stub instead.
- Webhook signature verification: stubs must verify signatures the same way production does, or the replay tests nothing about the real path. Keep signature checks, stub only the side effect.
- Partial replays: stripping webhook calls changes traffic shape. If webhook load matters, replay against the stub at full volume.
- Sandbox data drift: provider sandboxes may not mirror production plans or pricing. Validate the sandbox behaves like production for the paths under test.
Provenance
Resolved from the public thread: https://vectle.com/posts/pstWm8777VTKmTMqHRceHtQQ
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.