## TL;DR

Test the agent like software: a labeled set of invoices with expected extractions, matches, and routings, run on every change. Score field-level accuracy, match correctness, and routing decisions separately; gate deployment on the scores. Keep a holdout set the team never trains on for honest measurement.

## Steps

1. Build a labeled set covering your vendor and layout variety.
   Expected: A representative test set.
2. Define scoring: field accuracy, match correctness, routing.
   Expected: Separate metrics per capability.
3. Run the suite on every prompt, model, or code change.
   Expected: Regression protection.
4. Gate deployment on score thresholds.
   Expected: Quality bars.
5. Keep a holdout set for honest evaluation.
   Expected: No overfitting illusions.

## When to use

- Pre-deployment testing
- Model or prompt changes
- Vendor mix changes

## When not to use

- Production monitoring (different system)
- Load testing
- Security testing

## Compatibility

Framework-agnostic.

## Variant phrasings

### test invoice extraction agent

### labeled invoice dataset

### AP agent regression testing

## Root cause

Agents change behavior with prompts and models in ways unit tests miss. Labeled end-to-end tests are the only reliable gate.

## Edge cases

- Labeling is expensive; sample strategically across vendors and layouts
- Ugly scans must be in the set, not just clean PDFs
- Score routing separately from extraction; both can fail independently

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_D1_PEUfYvTF7cchsmdZUzw
