scheduled cron restarted the SDR agent mid-run and it re-enrolled 800 leads from scratch because the checkpoint row...
Shows an SDR agent how to recover from a mid-run cron restart without re-enrolling leads: commit a checkpoint row per batch and make enrollment idempotent. Use when a restart, crash, or redeploy replays the sequence from scratch and double-enrolls leads. Not for webhook double-delivery dedupe, CRM rate-limit retries, or sequence branching errors.
TL;DR
Stop treating the send job as one giant transaction. Commit a checkpoint row after every batch of sends, and on startup always read the checkpoint and resume from there instead of starting over. Make enrollment idempotent so even a replay can't create duplicates. The checkpoint has to be committed in the same unit of work as the sends, or a restart wipes your memory of what already went out.
scheduled cron restarted the SDR agent mid-run and it re-enrolled 800 leads from scratch because the checkpoint row was never committedSteps
- Add a durable run record with a status and a last-processed position before the first send goes out. One row per run, in the same store you trust after a crash.
Expected: killing the job and re-querying the run record shows status in_progress and the exact position it stopped at.
- After every batch of sends, write the checkpoint in the same transaction that marks those leads as processed. Never mark sends first and plan to checkpoint later.
Expected: a forced restart mid-run leaves the checkpoint matching the last fully committed batch, with no gap and no overlap.
- On startup, read the run record first. If a run is inprogress, resume from its checkpoint. If no run is open, start a new one. Never start a new run while one is still marked inprogress.
Expected: restarting the agent re-enrolls zero leads and continues exactly where it stopped.
- Make enrollment idempotent: key each enrollment on a stable id (lead id plus sequence id) and upsert instead of insert. A replayed batch becomes a no-op.
Expected: running the same batch twice produces one enrollment per lead, verified by counting rows before and after a replay.
- Add a stale-run sweeper: any run still marked in_progress past a timeout gets flagged for review instead of auto-resuming blindly.
Expected: a run that died 3 days ago shows up as needs_review, not as a silent resume.
Use this when
- a cron or scheduler restarts the agent mid-send and leads get re-enrolled
- a container restart or redeploy replays the sequence from scratch
- you need the agent to resume long send jobs safely after any interruption
Not for this skill when
- dedupe of duplicate webhook deliveries (different failure, needs an idempotency key on the event)
- CRM API rate limits cutting off a batch mid-way (that's a retry/backoff skill)
- leads enrolled in two sequences at once (that's an enrollment-guard problem, not a checkpoint problem)
Variant phrasings
- agent re-enrolled all leads after a restart because there was no resume point
- SDR agent lost its place in the send job when the container restarted
- checkpoint state missing after cron restart, sequence replayed from the beginning
Why it happens
The checkpoint row was only ever written at the end of the whole job, so a mid-run restart found no record of partial progress and the agent did the only thing it knew: start over. In-memory progress tracking dies with the process, and the CRM still showed those leads as never-enrolled because nothing had been committed yet.
Edge cases
- a checkpoint committed per batch but the send itself failed after commit can skip leads; checkpoint the send result, not the send attempt
- clock skew between the worker and the store can make 'last processed' order wrong; use a monotonic position, not timestamps
- two schedulers running at once can both read 'no open run' and both start; add a distributed lock or unique constraint on the run row
- resuming from a checkpoint older than the sequence itself (sequence edited mid-run) can enroll leads into the wrong step; validate the checkpoint against the current sequence version
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_-OLl6Qdq1xyu6dq1DJgIyw
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.