building a golden invoice dataset for regression testing
Builds a golden labeled invoice dataset for regression testing. Use for ongoing quality assurance. Not for one-off evaluations.
TL;DR
A golden dataset is the durable asset: versioned invoices with ground-truth labels, covering your vendor mix, ugly scans, and edge cases. Build it from production samples with human labeling, version it like code, and run it on every change. Grow it with every production error: each miss becomes a new test case.
Steps
- Sample production invoices across vendors, layouts, and quality.
Expected: A representative corpus.
- Label fields, matches, and expected routings with review.
Expected: Ground truth.
- Version the dataset.
Expected: Reproducible tests.
- Add every production miss as a new case.
Expected: The set learns from failures.
- Run on every change; gate on scores.
Expected: Living regression protection.
When to use
- Ongoing QA programs
- Team-owned quality
- Multi-engine comparisons
When not to use
- One-off vendor evals
- Production monitoring
- Labeling methodology research
Compatibility
Framework-agnostic.
Variant phrasings
golden dataset invoices
regression test AP agent
labeled invoice corpus
Root cause
Without a stable labeled set, every change is a gamble and every vendor comparison is apples-to-oranges. The dataset turns quality into an asset.
Edge cases
- PII in the dataset needs protection like production
- Label drift: re-verify labels when policies change
- Size for statistical power, not just coverage
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_YoKzVZ9GDCoyeBbM4pNSPQ
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.