## TL;DR
Search endpoints rank and reindex while you crawl, so the same record appears on multiple pages. Dedupe everything by record ID on the way in, and for full syncs prefer the provider's stable list endpoint (or a snapshot/export) over search. Search is for finding things, not for enumerating them.

## The query

```text
agent paginated a search endpoint  -  the search index updated mid-crawl and the generated sync duplicated hundreds of records
```

## Steps

### 1. Confirm the duplicates come from index movement

Compare the duplicated IDs: if the same record appears on two different pages of one crawl, and the total collected exceeds the index size, the index shifted under the crawl. Rule out retry-caused duplicates by checking request logs first.

Expected: duplicate IDs proven to come from overlapping pages, not from retried requests.

### 2. Dedupe by record ID immediately

Add a seen-ID set to the crawl: skip any record whose ID was already collected. This makes the sync correct (no duplicate processing downstream) even though the crawl itself is still unstable.

Expected: zero duplicate records in the collected set, regardless of index movement.

### 3. Switch full syncs to the stable list endpoint

Check whether the provider offers a list or export endpoint for the same records - list endpoints with cursors are built for enumeration; search endpoints are built for relevance. Move bulk syncs to the stable endpoint and reserve search for actual searches.

Expected: the sync runs against an endpoint whose pages do not shift mid-crawl.

### 4. If only search exists, shrink the crawl window

When search is the only enumeration path, page with the narrowest time or ID window the endpoint supports, and sort by an immutable field. Smaller windows shift less. Re-crawl and dedupe to converge.

Expected: duplicate rate drops; dedupe from step 2 absorbs the remainder.

### 5. Check for a snapshot or point-in-time option

Some search providers offer point-in-time or snapshot reads that freeze the index for the crawl. If available, open the snapshot first and paginate within it.

Expected: a frozen index view the crawl pages against, eliminating movement entirely.

## Use this when

- A crawl of a search endpoint returns the same records on multiple pages
- Collected counts exceed the index size with no errors
- The generated sync paginates a search endpoint for bulk enumeration
- Duplicates appear in bursts matching index update cycles

## Not for this skill when

- Duplicates come from retried requests (idempotency problem)
- The endpoint is a stable list endpoint (cursor or keyset problem instead)
- Records genuinely changed between crawls (that is new data, not duplication)
- The search query itself is wrong (relevance problem)

## Variant phrasings

### sync duplicated hundreds of records overnight

Overnight syncs span index update cycles, which is when movement is worst. Steps 2 and 3 together fix it.

### search pagination returned overlapping pages

Overlapping pages are the mechanism; shifting index is the cause. Dedupe (step 2) treats the symptom, the stable endpoint (step 3) removes the cause.

## Why it happens

The agent needed "all records" and the search endpoint was the one it knew, so it paginated it like a list endpoint. Search indexes are relevance-ranked and continuously updated: a reindex between page one and page two reshuffles which records land on which page, and records near the boundary appear twice. List endpoints exist precisely because search cannot promise stable enumeration, but nothing in the generated code knew the difference.

## Edge cases

- Provider's search is the only read path: steps 2 and 4 are the whole fix. Document that full syncs are approximate and converge via dedupe.
- Records deleted mid-crawl: the crawl may also MISS records, not just duplicate them. Dedupe does not fix misses; only the stable endpoint or snapshot does.
- Sort by relevance: the least stable ordering. If you must crawl search, sort by an immutable field like creation time, never by relevance.
- Very large result sets: deep pages on shifting indexes degrade fastest. Prefer time-windowed crawls (step 4) over deep offset paging.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_Pw0FesU0YW0Yt2N7QHx7Pg
