two agent instances processed the same lead CSV at once and both emailed the full list - no distributed lock on the...
Adds a distributed lock to an SDR send job so two agent instances can never process the same lead CSV at once. Use when duplicate outreach happened because parallel workers shared a send queue with no coordination. Not for single-instance duplicate sends, sequence-state loss after restarts, or CRM write-back failures.
TL;DR
Two workers, one CSV, no lock: both read "unsent" and both sent. The fix is a distributed lock around the claim step. Each worker atomically claims a batch of leads (status moves from queued to in-progress in a single transaction) before sending anything. Claims expire so a crashed worker does not block the queue forever.
The query
two agent instances processed the same lead CSV at once and both emailed the full list - no distributed lock on the send jobUse this when
- Duplicate outreach happened because parallel workers shared a send queue.
- Two agent instances (or a retry plus the original) processed the same lead list.
- The send job has no coordination between workers.
Not for
- Single-instance duplicate sends (that is a state-loss or idempotency problem).
- Sequence state wiped on container restart (that is a persistence problem).
- Contacted state recorded locally but never written to the CRM (that is a write-back problem).
Steps
Step 1: Confirm the double-send from mail logs
grep "[lead email]" mail-send.log | grep -o "msg-id=[^ ]*" | sort | uniq -cExpected output: two message IDs for the same lead minutes apart, from different worker IDs. That confirms parallel workers, not a retry. A retry would show one worker sending twice.
Step 2: Add an atomic claim step before any send
UPDATE leads
SET status = 'in-progress', claimed_by = '[worker id]', claimed_at = NOW()
WHERE lead_id IN (
SELECT lead_id FROM leads
WHERE status = 'queued'
LIMIT [batch size]
FOR UPDATE SKIP LOCKED
)
RETURNING lead_id, email;Expected output: each worker's claim returns a disjoint set of leads. FOR UPDATE SKIP LOCKED makes the claim atomic: two workers running this at the same instant cannot claim the same row. The check and the act are now one step.
Step 3: Send only claimed leads, then mark them done
echo "worker sends to its claimed batch, then: UPDATE leads SET status='sent' WHERE claimed_by='[worker id]'"Expected output: the worker's send loop iterates only over rows it claimed. On completion the rows move to sent. No worker ever touches a row it did not claim.
Step 4: Expire stale claims so crashes do not block the queue
UPDATE leads
SET status = 'queued', claimed_by = NULL
WHERE status = 'in-progress'
-- plus an age condition: only claims older than 30 minutesExpected output: leads claimed by a crashed worker return to the queue after 30 minutes. Without expiry, a dead worker's batch sits in-progress forever and those leads never get emailed.
Step 5: Add a job-level lock as a second layer
redis-cli SET send-job-lock [worker id] NX EX 3600Expected output: only one worker holds the job lock at a time. If the lock is already held, the second worker exits cleanly instead of processing the CSV. The row-level claim in step 2 is the real protection. The job lock is defense in depth.
Step 6: Dedupe at send time as the last line of defense
grep -q "[lead email]" sent-log.txt || send_email "[lead email]"Expected output: even if both locking layers failed, the sent-log check stops the second send. Three layers (job lock, atomic claim, send-time dedupe) make a repeat of this incident practically impossible.
Variant phrasings
outreach webhook fired twice and both runs sent the full sequence
Same no-dedupe shape at the trigger level. An idempotency check on the webhook payload precedes the claim step.
two outreach agents emailed the same candidate list
Same missing claim-lock. The claim-expired variant: the lock existed but expired mid-run, so both workers picked up the shard.
scheduled cron restarted the agent mid-run and it re-enrolled 800 leads
Same state-sharing failure. The checkpoint row must commit before the sends it covers, or the restart replays them.
Why it happens
"Check then act" across two processes is a race. Worker A reads the CSV, sees unsent leads. Worker B reads the CSV a millisecond later, sees the same unsent leads. Both act. Each worker's view was correct at read time and wrong by send time, because nothing made the read and the send atomic. CSV files have no transactions, which is why the queue must live in a database (or a real lock service) where atomic claims exist.
Edge cases
SKIP LOCKEDis PostgreSQL syntax. On MySQL useFOR UPDATEwith a short lock timeout. On datastores without row locking, use the Redis lock from step 5 as the primary mechanism.- The claim-expiry window must exceed the longest legitimate batch send time, or slow workers lose their claims mid-send and the batch gets double-processed by the reaper.
- A worker that crashes between sending and marking sent leaves the lead in-progress until expiry, then it gets re-emailed. The send-time dedupe in step 6 covers exactly this window.
- Do not "fix" this by running a single worker. That trades a correctness bug for a throughput ceiling and a single point of failure. Fix the coordination instead.
- Log every claim, send, and expiry with the worker ID. When duplicates happen again (new code path, new tool), the logs show which layer failed.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_U6xBCC1v0oWR8jGqjT2qsg
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.