## TL;DR

A golden dataset is the durable asset: versioned invoices with ground-truth labels, covering your vendor mix, ugly scans, and edge cases. Build it from production samples with human labeling, version it like code, and run it on every change. Grow it with every production error: each miss becomes a new test case.

## Steps

1. Sample production invoices across vendors, layouts, and quality.
   Expected: A representative corpus.
2. Label fields, matches, and expected routings with review.
   Expected: Ground truth.
3. Version the dataset.
   Expected: Reproducible tests.
4. Add every production miss as a new case.
   Expected: The set learns from failures.
5. Run on every change; gate on scores.
   Expected: Living regression protection.

## When to use

- Ongoing QA programs
- Team-owned quality
- Multi-engine comparisons

## When not to use

- One-off vendor evals
- Production monitoring
- Labeling methodology research

## Compatibility

Framework-agnostic.

## Variant phrasings

### golden dataset invoices

### regression test AP agent

### labeled invoice corpus

## Root cause

Without a stable labeled set, every change is a gamble and every vendor comparison is apples-to-oranges. The dataset turns quality into an asset.

## Edge cases

- PII in the dataset needs protection like production
- Label drift: re-verify labels when policies change
- Size for statistical power, not just coverage

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_YoKzVZ9GDCoyeBbM4pNSPQ
