# Fix the extraction agent loop re-extracting the same catalog after failure

## TL;DR

The agent retries the whole extraction instead of resuming, so make it checkpoint per file and skip completed work. A failure near the end currently throws away all progress and starts over, which looks like a loop. Write a progress file as each unit completes, and on retry only process what is missing.

## The error

```text
Extraction agent log shows pass 1, pass 2, pass 3 over the same 800 files
each pass fails at 95 percent and restarts from zero
```

## Fix it

### Step 1: Confirm the restart pattern in the logs

```bash
grep -c "starting extraction pass" logs/agent.log
```

Expected: The count is high, proving full restarts.

### Step 2: Add per-file checkpointing

```bash
node -e "console.log('pattern: after each file, append its path to .extract-progress; on start, skip paths already listed')"
```

Expected: Progress survives a crash.

### Step 3: Cap retries with backoff

```bash
node -e "console.log('max 3 attempts per file, then mark failed and continue -- never restart the whole run')"
```

Expected: One bad file cannot loop the run.

### Step 4: Re-run and watch it resume

```bash
node scripts/extract-strings.js | tail -3
```

Expected: The run completes, resuming past the old failure point.

## When to use this

- An extraction agent loops re-processing the same files
- A late failure restarts the whole run

## When NOT to use this

- The agent loops on one file forever, that is a parser hang on that file
- The loop is in translation retries, different fix

## Tool and version compatibility

- Any agent-based extraction with retry logic
- Progress files or a small state store

## Variant phrasings

### loop detected by a watchdog that kills it

The watchdog is right. Fix the resume logic, do not just raise the watchdog limit.

### progress file itself gets corrupted

Write it atomically: write temp then rename. Partial writes are what corrupt it.

## Why it happens

Retry logic written as try-the-whole-thing-again treats a 95-percent-complete run like a zero-percent one. Without checkpoints, every retry repeats the completed work, hits the same late failure, and retries again, which is the loop.

## Edge cases

- Checkpoint per file, not per batch, batches still repeat work
- Clean the progress file on a config change, stale skips hide new strings
- Log which files were skipped on resume so the run is auditable

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_V0J-wBIuu6qijM2l8HL1bw
