bulk-import agent created 300 duplicate candidates - the retry loop on a timed-out POST never checked that the first...
Fixes bulk imports that create duplicate ATS candidates when a timed-out create call is retried without checking whether the first write succeeded. Use when an import or sync retries POST requests. Key trigger: duplicate candidate records appearing right after a slow or timed-out import run.
TL;DR
Treat every timeout as unknown and verify before retrying: after a timed-out create, look the candidate up by email before re-posting. A timeout means the response was lost, not that the write failed, so blind retries create duplicates. Merge the existing dupes on email, then make the import idempotent with check-before-create plus a client-generated idempotency key.
The error
Bulk import finished with 300 extra candidates - the retry loop re-POSTed creates that timed out, but the first writes had already succeededSteps
- Pause the import loop so no more duplicates are created.
Expected: the candidate count stops growing.
- Find the duplicates: query the ATS for candidates sharing the same email address.
Expected: groups of 2+ records with identical emails, created minutes apart.
- Merge each group into one record, keeping the richest profile and all attached notes and activities.
Expected: one candidate per email, with no notes or history lost.
- Change the retry logic: on timeout, look the candidate up by email first and only create when the lookup finds nothing.
Expected: re-running a timed-out batch creates zero new records.
- Attach a client-generated idempotency key (a stable value per candidate, e.g. derived from the source row) to every create call so a repeated request is deduplicated server-side.
Expected: the ATS returns the existing record instead of creating a second one on a duplicate key.
- Set sane timeouts and log every retry decision with its verification result.
Expected: future timeouts are audited as verified-absent-and-reposted or verified-present-and-skipped.
Use this when
- A bulk import or sync creates candidates via POST with retry logic
- Duplicates appear after slow, flaky, or timed-out runs
- An agent syncs candidates from a spreadsheet, CSV, or external system
Not for this skill when
- Duplicates come from two source systems with genuinely different emails (that is a matching problem, not a retry problem)
- The source data itself has dupes (dedupe the source before importing)
- Records were intentionally created twice, e.g. separate applications per job (check your dedup policy first)
Variant phrasings
- retry loop created duplicate candidates after POST timeout
- bulk import double-created candidates on a slow API
- timeout but the write succeeded, duplicate ATS records
- import agent re-submitted candidates that were already created
Why it happens
An HTTP timeout only means the client never saw the response; the server may have completed the write. The retry loop assumed timeout equals failure and re-posted unconditionally, so every timed-out create ran twice. Without a check-before-create or an idempotency key, nothing stops the second write.
Edge cases
- Server 500s that partially wrote: same rule, verify before retrying on any ambiguous response.
- Email missing or malformed in the source row: fall back to an external source ID for the lookup key.
- Concurrent imports: two workers can both pass the check, so the idempotency key is the real guard.
- Merge order matters: keep the record with the most complete history as the survivor.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_KXr4mdqKfT6Bdkfpgg6DhQ
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.