TL;DR
The agent reported extraction "complete" while the last 30 pages of schedules never entered context, because the completion flag was set by the loop finishing, not by coverage. Gate the "complete" flag on an independent page count: pages processed must equal pages in the document, verified separately from the processing loop. Anything less is "partial", reported as partial.

The misleading report:
```text
extraction status: complete (120 of 120 pages claimed, 90 actually processed)
```

## Steps
1. Confirm the gap. Get the document's true page count independently and compare it to the number of pages your extraction actually processed.
   Command: read the page count from the PDF metadata and compare it against your extraction log's processed-page count.
   Expected: a mismatch, like 120 pages in the file versus 90 processed, proving the trailing schedules were never read.
2. Add an independent page-count check at intake: record the document's page count from the file itself before extraction starts, and store it where the processing loop cannot overwrite it.
   Expected: you have a ground-truth page count that no downstream bug can fudge.
3. Gate the completion flag on coverage: set status to "complete" only when processed pages equal the intake page count, otherwise set "partial" with the missing ranges listed.
   Expected: the same run that used to report "complete" now reports "partial, pages 91-120 unprocessed".
4. Find why the trailing pages were skipped, usually a loop bound, a pagination bug, or an early exit on a blank page, and fix the root cause so the common case is genuinely complete.
   Expected: re-running extraction processes all 120 pages and the complete flag is earned, not assumed.

## Use this when
- Extraction reports complete but trailing pages, schedules, or exhibits are missing from the output
- Your completion flag is set by loop termination rather than coverage verification
- Long documents lose their endings silently with no error
- You cannot independently verify what fraction of a document was processed

## Not for this skill when
- The pages are missing from the source file, verify the input document's page count first
- The trailing pages were processed but produced no extractable content, like blank pages, which is fine if logged
- The incompleteness is already reported as partial, your problem is the root cause, not the reporting
- The extraction is streaming and "complete" refers to the stream ending, define coverage for streams separately

## Variant phrasings
- extraction said complete but the last 30 pages were never processed
- how to verify full document coverage in extraction pipelines
- agent reported success while schedules never entered context
- false complete status in contract extraction
- pages-in versus pages-processed mismatch

## Why it happens
Most pipelines set their done flag when the processing loop exits normally, which conflates "the loop finished" with "the document was fully processed". A loop that stops early, because of a bad bound, an exception swallowed somewhere, or a pagination quirk, still exits, and the flag still flips to complete. The trailing schedules are the usual victim because the bug only shows up near the end of long documents, exactly where nobody spot-checks.

## Edge cases
- Blank or image-only trailing pages may process to zero content, count them as processed but content-empty rather than missing.
- Encrypted or permission-restricted page ranges can block processing mid-document, report those ranges explicitly instead of silently skipping.
- Multi-file portfolios need per-file coverage, a portfolio-level complete flag hides per-file gaps.
- If the PDF's own page count metadata is wrong, which happens with some generators, count pages by rendering and use the higher of the two counts.
- Downstream consumers may cache the old false-complete outputs, re-run extraction for any document whose coverage was never verified.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_y9p-ClyrRaA9k9SO5-vtAA
