# Fix onboarding agent timing out polling for webhook delivery confirmation

## TL;DR
The agent times out because it polls for a delivery confirmation that will never come, usually after the delivery already failed. Check the delivery status once from the provider's log; if it failed, fix the endpoint and re-trigger instead of polling. Polling a dead delivery is the bug.

## The error
```text
Onboarding agent timeout
Timed out polling for webhook delivery confirmation that failed.
```

## Use this skill
Agent? Search Vectle for this error:
```bash
curl -s "https://vectle.com/api/v1/search?q=onboarding agent timed out polling for webhook delivery confirmation that failed"
```

## Fix it

### Step 1: Check the provider delivery log

```bash
Look up the delivery in the provider's dashboard or API.
```

Expected: You see the real status: delivered, failed, or never sent.

### Step 2: Stop polling dead deliveries

```bash
If the delivery failed, stop the poll loop immediately.
```

Expected: The agent stops waiting on something that will never succeed.

### Step 3: Fix the endpoint issue

```bash
Address whatever the delivery log shows: 4xx, 5xx, or timeout.
```

Expected: The endpoint is healthy.

### Step 4: Re-trigger the webhook

```bash
Send a fresh webhook or replay the event from the provider.
```

Expected: A new delivery attempt starts.

### Step 5: Poll the new delivery with a deadline

```bash
Poll the fresh delivery with a bounded timeout and backoff.
```

Expected: Confirmation arrives or the timeout fires with a clear status.

## When this applies

- Agents time out polling webhook deliveries
- Onboarding waits forever on webhook confirmation
- You are building webhook-waiting agents

## When it doesn't

- The webhook was never sent (check the trigger)
- Deliveries succeed but slowly (extend the deadline)
- The endpoint is fine and deliveries fail anyway (check the provider)

## Compatibility

Webhook delivery logs: Stripe, GitHub, Segment, and others.

## Variant phrasings

### agent polling webhook timeout

Same failure. Check the log first, poll second.

### webhook delivery confirmation never arrives

Never arrives usually means failed, not slow. The log tells you which.

### onboarding stuck waiting webhook

Stuck waits need a deadline and a fallback path, not infinite patience.

## Why it happens

Agents often poll for confirmation without first checking whether the delivery is even alive. A failed delivery stays failed; polling it until timeout wastes the whole wait budget. The agent needs to distinguish pending from dead before choosing to wait.

## Edge cases

- Some providers retry failed deliveries on their own; a fresh poll may catch the retry
- Idempotency on your endpoint makes re-triggered webhooks safe
- Log the delivery id with the poll so timeouts are traceable

## If it still fails

- Reproduce with a minimal run: one user, one file, one step.
- Read the agent's full trace, not just the final error; the failure is usually upstream.
- Check the underlying API or tool directly, outside the agent, to separate agent bugs from service bugs.
- Reduce concurrency to one and see if the failure persists; races hide as flakes.
- If the run is business-critical, add a human checkpoint before the destructive steps.

## Prevention

- Checkpoint long runs so any failure resumes instead of restarting.
- Cap and back off every retry loop; unbounded retries are outages waiting to happen.
- Validate inputs at each pipeline stage; fail fast with clear errors.
- Log enough context per step that a timeout is diagnosable without rerunning.
- Give destructive steps a human checkpoint or a dry-run mode.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_9hep5aorsrVxqMTYBg8MzQ
