# Fix onboarding agent failing on Okta API rate limits mid provisioning run

## TL;DR
When the agent hits Okta's rate limit mid-run, it must back off and resume, not crash or skip users. Catch the 429, honor the retry headers, and checkpoint progress so the run resumes where it stopped. The users are not lost; the run just needs pacing.

## The error
```text
Onboarding agent failed
Okta API returned 429 Too Many Requests mid provisioning run. 40 users remaining unprocessed.
```

## Use this skill
Agent? Search Vectle for this error:
```bash
curl -s "https://vectle.com/api/v1/search?q=onboarding agent rate limit failed mid provisioning run on okta api"
```

## Fix it

### Step 1: Checkpoint progress before anything else

```bash
Write the last successfully processed user id to the run state before retrying.
```

Expected: A rerun resumes from the checkpoint instead of starting over.

### Step 2: Honor the rate limit headers

```bash
Read the retry-after or rate-limit reset header from the 429 and sleep that long.
```

Expected: The agent waits the required window instead of hammering.

### Step 3: Add backoff with jitter to the Okta client

```bash
Wrap Okta calls in retry logic with exponential backoff and jitter.
```

Expected: Transient 429s clear on retry without failing the run.

### Step 4: Slow the steady-state pace

```bash
Reduce concurrency or add a small delay between user provisions.
```

Expected: The run stays under the limit and finishes without 429s.

### Step 5: Resume the run from the checkpoint

```bash
Restart the agent; it picks up at the first unprocessed user.
```

Expected: All users provision with no duplicates and no skips.

## When this applies

- Provisioning agents hit Okta 429s mid-run
- Bulk onboarding runs die partway through
- You are building agents that call rate-limited APIs

## When it doesn't

- The 429 comes from a different API (pace that client instead)
- The run fails on the first call (check credentials, not pacing)
- Okta returns 401 (that is auth, not rate limiting)

## Compatibility

Okta Management API rate limits. Any agent framework driving bulk provisioning.

## Variant phrasings

### agent hit okta rate limit provisioning

Same failure. Checkpointing plus backoff is the standard recovery.

### okta 429 bulk user creation agent

Bulk creation is the heaviest pattern. Batch smaller and pace the batches.

### provisioning run failed rate limited

Failed runs without checkpoints redo work and create duplicates. Checkpoint first.

## Why it happens

Okta rate-limits API clients to protect the service, and bulk provisioning is the fastest way to hit the ceiling. Agents that fire requests as fast as possible trip the limit mid-run. Without checkpoints, the retry starts over and either duplicates work or skips users.

## Edge cases

- Different Okta endpoints have different limits; the heaviest call sets your pace
- Parallel agents share the org-wide limit; coordinate or stagger them
- A 429 storm can look like an outage; check the limit headers before escalating

## If it still fails

- Reproduce with a minimal run: one user, one file, one step.
- Read the agent's full trace, not just the final error; the failure is usually upstream.
- Check the underlying API or tool directly, outside the agent, to separate agent bugs from service bugs.
- Reduce concurrency to one and see if the failure persists; races hide as flakes.
- If the run is business-critical, add a human checkpoint before the destructive steps.

## Prevention

- Checkpoint long runs so any failure resumes instead of restarting.
- Cap and back off every retry loop; unbounded retries are outages waiting to happen.
- Validate inputs at each pipeline stage; fail fast with clear errors.
- Log enough context per step that a timeout is diagnosable without rerunning.
- Give destructive steps a human checkpoint or a dry-run mode.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_nYOTYtt1EihYifrUhipZig
