## TL;DR

A batch job that cannot resume is a batch job that will eventually lose data. Find where the upsert stopped by comparing input rows to HubSpot records, resume from the last committed offset after the rate window clears, and change the job so it checkpoints progress after every chunk.

```
agent's batch upsert to HubSpot got cut off at the rate limit boundary and it never resumed - half the list has no sequence assignment
```

## Steps

1. Identify the cutoff: take the original input list and query HubSpot for the records the job was supposed to create or update. Sort both by the same key (email or record id).
   Expected: a clean split point where HubSpot records stop and unprocessed rows begin.

2. Verify the failure mode: check the job log around the cutoff time for 429s. If the job logged success and moved on, note that as the silent-drop bug to fix in step 4.
   Expected: confirmed the stop was throttling, not bad data.

3. Wait out the rate window (respect Retry-After), then re-run the upsert for ONLY the unprocessed rows, starting from the cutoff. Do not re-send the whole list.
   Expected: remaining rows process without duplicates.

4. Make the job resumable: after each chunk (e.g. every 50 records), write the last completed offset to a checkpoint store. On startup, the job reads the checkpoint and skips everything before it.
   Expected: a future cutoff resumes from the chunk boundary, not from zero.

5. Reconcile at the end: count input rows vs HubSpot records vs checkpoint offset. All three must agree before the job reports success.
   Expected: "half the list missing" becomes impossible to miss.

## Use this when

- A HubSpot batch upsert processed only part of the input
- Sequence assignment counts do not match the upload list
- A batch job hit a rate limit and reported completion anyway

## Not for this skill when

- Single-record calls are 429ing one at a time (per-request retry handling instead)
- The batch failed on validation errors (bad data needs fixing, not resuming)
- The job is a Salesforce Bulk API job (different job/offset model)
- Nothing was processed at all (the job never started; check scheduling, not resumption)

## Variant phrasings

- "HubSpot batch upsert stopped halfway at the rate limit, never resumed"
- "partial batch upsert to HubSpot, missing sequence assignments"
- "agent batch job cut off by HubSpot throttling with no checkpoint"

## Why it happens

The job was written as a single pass over the input: iterate, upsert, done. Rate limiting turned a transient pause into a permanent stop because there was no concept of "not yet done" - the loop ended when the API said slow down, and the code reported the partial run as complete. Checkpoints turn a batch into a restartable unit of work.

## Edge cases

- Rows near the boundary may have partially succeeded (record created but sequence step not assigned). Reconcile at the record level, not just the count level.
- If the input file changed between runs, offsets shift. Checkpoint on a stable key (email) rather than a row number.
- HubSpot's batch endpoint has its own per-batch size limits. Smaller chunks checkpoint more often but consume more rate budget; 50 to 100 per chunk is a sane tradeoff.
- A second run of the same job while the first is paused can double-process. Take a lock on the job id before resuming.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_PamKt4q8TmlDZ-QyGba4xw
