# Fix PDF parser failing during vendor onboarding while the agent times out on the SOC2 report

## TL;DR
The parser fails and the agent times out because SOC2 PDFs are long, scanned, or image-heavy. Split the work: extract text with a proper PDF tool, fall back to OCR for scanned pages, and give the agent a longer deadline for document steps. Parsing and waiting both need to match the document reality.

## The error
```text
Agent failed
PDF parser failed during vendor onboarding; agent timed out on SOC2 report.
```

## Use this skill
Agent? Search Vectle for this error:
```bash
curl -s "https://vectle.com/api/v1/search?q=pdf parser failed during vendor onboarding while agent timed out on soc2 report"
```

## Fix it

### Step 1: Check the PDF type

```bash
Determine whether the report is text-based or scanned images.
```

Expected: You know whether text extraction or OCR is needed.

### Step 2: Extract with the right tool

```bash
Use text extraction for text PDFs and OCR for scanned pages.
```

Expected: You get usable text from the report.

### Step 3: Process in sections

```bash
Split the document into sections and process each within its own budget.
```

Expected: No single step times out on document size.

### Step 4: Extend the document deadline

```bash
Give document steps a longer timeout than API steps.
```

Expected: The agent finishes without timing out.

### Step 5: Verify the extracted data

```bash
Check the key fields the onboarding needs against the PDF.
```

Expected: The vendor record is complete and accurate.

## When this applies

- Agents fail parsing vendor SOC2 PDFs
- Document steps time out in onboarding
- You are building document-processing onboarding

## When it doesn't

- The PDF is corrupt (get a fresh copy)
- The parser works but extraction is wrong (tune the extraction)
- The document is not a PDF (use the right parser)

## Compatibility

PDF text extraction and OCR tools. Vendor onboarding flows.

## Variant phrasings

### pdf parser failed soc2 report agent

Same failure. Right tool plus sectioned processing fixes it.

### agent timeout large pdf onboarding

Large documents need sectioned work and longer deadlines.

### scanned pdf parsing failed vendor onboarding

Scanned pages need OCR. Text extraction alone returns nothing.

## Why it happens

SOC2 reports are long and often scanned, which breaks naive text extraction and blows through agent step timeouts sized for API calls. The parser and the timeout were both built for a different kind of document.

## Edge cases

- Password-protected PDFs from vendors need the password before any parsing works
- OCR quality varies; verify critical fields against the source visually for key vendors
- Store the extracted text so re-parsing is not needed on retry

## If it still fails

- Reproduce with a minimal run: one user, one file, one step.
- Read the agent's full trace, not just the final error; the failure is usually upstream.
- Check the underlying API or tool directly, outside the agent, to separate agent bugs from service bugs.
- Reduce concurrency to one and see if the failure persists; races hide as flakes.
- If the run is business-critical, add a human checkpoint before the destructive steps.

## Prevention

- Checkpoint long runs so any failure resumes instead of restarting.
- Cap and back off every retry loop; unbounded retries are outages waiting to happen.
- Validate inputs at each pipeline stage; fail fast with clear errors.
- Log enough context per step that a timeout is diagnosable without rerunning.
- Give destructive steps a human checkpoint or a dry-run mode.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_eODCKDRAFpZjCMQ-MHhxmg
