## TL;DR
Treat every timeout as unknown and verify before retrying: after a timed-out create, look the candidate up by email before re-posting. A timeout means the response was lost, not that the write failed, so blind retries create duplicates. Merge the existing dupes on email, then make the import idempotent with check-before-create plus a client-generated idempotency key.

## The error
```text
Bulk import finished with 300 extra candidates - the retry loop re-POSTed creates that timed out, but the first writes had already succeeded
```

## Steps
1. Pause the import loop so no more duplicates are created.
   Expected: the candidate count stops growing.
2. Find the duplicates: query the ATS for candidates sharing the same email address.
   Expected: groups of 2+ records with identical emails, created minutes apart.
3. Merge each group into one record, keeping the richest profile and all attached notes and activities.
   Expected: one candidate per email, with no notes or history lost.
4. Change the retry logic: on timeout, look the candidate up by email first and only create when the lookup finds nothing.
   Expected: re-running a timed-out batch creates zero new records.
5. Attach a client-generated idempotency key (a stable value per candidate, e.g. derived from the source row) to every create call so a repeated request is deduplicated server-side.
   Expected: the ATS returns the existing record instead of creating a second one on a duplicate key.
6. Set sane timeouts and log every retry decision with its verification result.
   Expected: future timeouts are audited as verified-absent-and-reposted or verified-present-and-skipped.

## Use this when
- A bulk import or sync creates candidates via POST with retry logic
- Duplicates appear after slow, flaky, or timed-out runs
- An agent syncs candidates from a spreadsheet, CSV, or external system

## Not for this skill when
- Duplicates come from two source systems with genuinely different emails (that is a matching problem, not a retry problem)
- The source data itself has dupes (dedupe the source before importing)
- Records were intentionally created twice, e.g. separate applications per job (check your dedup policy first)

## Variant phrasings
- retry loop created duplicate candidates after POST timeout
- bulk import double-created candidates on a slow API
- timeout but the write succeeded, duplicate ATS records
- import agent re-submitted candidates that were already created

## Why it happens
An HTTP timeout only means the client never saw the response; the server may have completed the write. The retry loop assumed timeout equals failure and re-posted unconditionally, so every timed-out create ran twice. Without a check-before-create or an idempotency key, nothing stops the second write.

## Edge cases
- Server 500s that partially wrote: same rule, verify before retrying on any ambiguous response.
- Email missing or malformed in the source row: fall back to an external source ID for the lookup key.
- Concurrent imports: two workers can both pass the check, so the idempotency key is the real guard.
- Merge order matters: keep the record with the most complete history as the survivor.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_KXr4mdqKfT6Bdkfpgg6DhQ
