# Fix onboarding agent hitting context overflow on 200-row CSV import

## TL;DR
The agent overflows because it loads the whole CSV into context at once. Process the file in small batches, streaming rows and summarizing progress between batches. The import is a streaming job, not a reading task.

## The error
```text
Onboarding agent failed
Context overflow on 200-row CSV import. Run aborted partway.
```

## Use this skill
Agent? Search Vectle for this error:
```bash
curl -s "https://vectle.com/api/v1/search?q=onboarding agent hit context overflow on 200-row csv import failed"
```

## Fix it

### Step 1: Switch to streaming reads

```bash
Read the CSV row by row or in small chunks instead of loading it whole.
```

Expected: Memory and context stay flat regardless of file size.

### Step 2: Batch the work

```bash
Process rows in batches with a checkpoint after each batch.
```

Expected: A failure loses at most one batch of work.

### Step 3: Summarize between batches

```bash
After each batch, compress progress to a short summary and drop the raw rows.
```

Expected: Context stays bounded across the whole file.

### Step 4: Validate rows before processing

```bash
Check required fields per row and quarantine bad rows instead of failing the run.
```

Expected: One bad row no longer kills the import.

### Step 5: Verify the full import

```bash
Compare imported user count against the CSV row count.
```

Expected: Counts match and bad rows are listed for review.

## When this applies

- Agents overflow context on CSV imports
- Large file imports die partway through
- You are building file-processing agents

## When it doesn't

- The CSV itself is malformed (fix the file)
- The import API rejects rows (check the API errors)
- Small files fail too (different problem)

## Compatibility

CSV imports generally. Any agent framework with bounded context.

## Variant phrasings

### agent context overflow csv import

Same failure. Streaming plus batching is the fix.

### csv import too large agent

Too large for context is normal. Stream it; do not shrink the business requirement.

### agent failed large file processing

Large files need streaming, checkpointing, and summarization together.

## Why it happens

Language-model agents have finite context, and a 200-row CSV with wide columns easily exceeds it when loaded whole. The agent then aborts mid-import with partial work done. Treating the file as a stream instead of a document keeps context bounded.

## Edge cases

- Wide rows overflow faster than many narrow rows; batch by token estimate, not row count
- Quarantined bad rows need a human-readable report, not a silent skip
- Resume from the last checkpoint, never from the start, on retry

## If it still fails

- Reproduce with a minimal run: one user, one file, one step.
- Read the agent's full trace, not just the final error; the failure is usually upstream.
- Check the underlying API or tool directly, outside the agent, to separate agent bugs from service bugs.
- Reduce concurrency to one and see if the failure persists; races hide as flakes.
- If the run is business-critical, add a human checkpoint before the destructive steps.

## Prevention

- Checkpoint long runs so any failure resumes instead of restarting.
- Cap and back off every retry loop; unbounded retries are outages waiting to happen.
- Validate inputs at each pipeline stage; fail fast with clear errors.
- Log enough context per step that a timeout is diagnosable without rerunning.
- Give destructive steps a human checkpoint or a dry-run mode.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_hfVXuMLF_NvLbRZZtloGCQ
