## TL;DR

Confidence scores are the cheapest triage signal: high-confidence straight-through, medium gets targeted rechecks, low goes to humans with the weak fields highlighted. Calibrate the bands on labeled data, track the review queue's hit rate, and retune as vendors change. The review queue should feel like quality control, not random sampling.

## Steps

1. Calibrate confidence bands on labeled invoices.
   Expected: Bands tied to real accuracy.
2. Route high band straight through with validation.
   Expected: Fast path.
3. Recheck medium band fields selectively.
   Expected: Targeted recovery.
4. Present low band to reviewers with highlights.
   Expected: Efficient review.
5. Track hit rate and retune.
   Expected: A queue that stays useful.

## When to use

- Production AP agent operations
- Review queue management
- Accuracy tuning

## When not to use

- Initial threshold design
- Fraud review (different criteria)
- Manual processes

## Compatibility

Needs per-field confidence from the extractor.

## Variant phrasings

### confidence routing AP

### review queue invoice AI

### triage by extraction confidence

## Root cause

Reviewer time is the scarcest resource in AP automation. Confidence routing spends it where the model is actually unsure, which is where humans add value.

## Edge cases

- Overconfident models need the arithmetic validation as backstop
- Bands drift with new vendors; monitor the hit rate
- Reviewers need the why, not just the what: show low-confidence fields

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_TrPeqjeZN40RSVV9CxrbIQ
