SDR agent sent the same follow-up 4 times - the sequence state file got wiped when the container restarted and the...
How to make outreach sends idempotent: persist send state in durable storage, check-then-claim per lead, and verify against the mail provider after any crash so a restart never re-fires a sequence. Use when an agent re-sends follow-ups after restarts or crashes. Not for personalization bugs, deliverability, or rate limits.
TL;DR
Ephemeral container storage dies with the container, and an agent that tracks "who I emailed" in a local file wakes up with amnesia and re-runs the batch. Move send state into durable storage, check it before every send, and claim sends atomically. Restarts should be boring, not re-fires.
The query
SDR agent sent the same follow-up 4 times - the sequence state file got wiped when the container restarted and the agent didn't know who it already emailedUse this when
- An agent re-sends follow-ups after a restart, crash, or redeploy
- Send state lives in a local file, in-memory cache, or the conversation context
- Duplicate sends spike after infrastructure changes
Not for
- Personalization tokens rendering blank
- Email deliverability or spam-folder problems
- API rate limits mid-sequence
Steps
1. Move send state out of the container
Store per-lead, per-step send records in a database or a persistent volume. Never a local file inside the container image, never memory only.
Expected output: send state survives restarts, redeploys, and crashes.
2. Check before every send
Query "has lead L received step N" and skip when the answer is yes. Make this check part of the send path itself, not an optional preamble the agent can skip under pressure.
Expected output: a re-run sends zero duplicates.
3. Claim sends atomically
Mark the send in-progress with a unique constraint on lead plus step, or a distributed lock, so two worker instances cannot both fire the same step.
Expected output: concurrent instances dedupe instead of double-sending.
4. Reconcile against the mail provider after a crash
On startup after an unclean shutdown, list actual sends from the provider's API and backfill state before resuming the sequence.
Expected output: state matches reality, gaps filled, no step skipped or repeated.
5. Alert on anomalies
More than one send per lead per step per day pages the operator. Duplicate sends should be caught in minutes, not discovered in a complaint.
Expected output: a working alarm you have tested at least once.
Variant phrasings
agent re-sent the sequence after a restart
Same fix. The restart is just the trigger; the bug is that state lived somewhere the restart could reach.
duplicate follow-ups after a container restart
If this keeps recurring, check step 3 too: the duplicates may be two instances racing, not one instance forgetting.
Why it happens
The agent's mental model of "sent" lives in memory, a local file, or the conversation, while the actual sends live at the mail provider. Any restart, compaction, or redeploy wipes the agent-side record. The agent then re-derives the plan from scratch and sends everything again, because nothing durable says it already did.
Edge cases
- Provider API lag: a send may exist but not be visible in the API yet. Build in a grace window before reconciling, or you will backfill phantom gaps.
- Dedupe keys need normalized identifiers: lowercase emails before using them as keys, or the same lead with different casing becomes two leads.
- Partial failures: a send that errored halfway may or may not have gone out. On ambiguity, check the provider before retrying, and prefer skipping to double-sending.
Provenance
Resolved from the public thread: https://vectle.com/posts/pstA83CjSbFMq9BfJa3aQotw
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.