## TL;DR
Manual QA reviews 2 percent of tickets and misses everything else. Automated quality measurement uses an LLM judge scoring every answer against a rubric (accuracy, completeness, tone, policy compliance), calibrated against human reviewers until it agrees with them 85 percent of the time. Start with a small rubric and a sampled baseline, calibrate before you trust it, and track the scores weekly. The goal is a quality signal on every ticket, not a replacement for human judgment on the hard ones.

## The query

```text
how to measure support answer quality automatically
```

## Use this when

- You need quality signal on more than a manual sample
- Auditing an AI support agent's answers at scale
- Manual QA is too slow or too expensive
- Quality is slipping and you need data on where

## Not for

- Writing the manual QA rubric itself
- CSAT or customer survey design
- Measuring speed or handle time
- Compliance or legal review

## Steps

### 1. Write a small, concrete rubric

Four to six dimensions max: factual accuracy, completeness (did it answer the actual question), tone match, policy compliance. Each dimension gets a 1 to 5 scale with a one-line description of what 3 means. Vague rubrics produce vague scores.

Expected output: a rubric a new reviewer could apply without training.

### 2. Build a human-scored baseline

Have two human reviewers score 200 to 300 past answers with the rubric. This is your ground truth. It is also where you discover your rubric is ambiguous, fix it now.

Expected output: 200+ answers with human scores and inter-reviewer agreement above 80 percent.

### 3. Calibrate the LLM judge

Prompt the judge with the rubric, the ticket context, and the answer. Score the same 200 answers and compare. Iterate on the prompt until judge-human agreement hits 85 percent. Do not skip this, an uncalibrated judge is a random number generator with confidence.

Expected output: documented agreement rate between the judge and human reviewers.

### 4. Run it on everything, sample the edges

Score every new answer automatically. Humans review a sample: all the lowest scores, plus a random 2 percent. The judge finds the problems, humans confirm them.

Expected output: full-coverage scores with human review concentrated where it matters.

### 5. Track trends, not just scores

Weekly: average score per dimension, score distribution, worst ticket types. A dimension dropping over three weeks is a coaching or training signal, not a one-off.

Expected output: a weekly quality dashboard the team actually looks at.

### 6. Recalibrate quarterly

Products change, policies change, the judge drifts. Re-run the human baseline quarterly and adjust the prompt. Treat the judge like an employee that needs performance reviews.

Expected output: a scheduled recalibration with fresh agreement numbers.

## Variant phrasings

### automated QA for customer support answers

Steps 1 through 4 as the build plan. The calibration in step 3 is the part most teams skip and most regret skipping.

### LLM as judge for support quality

Same approach. Key detail: give the judge the ticket context, not just the answer. Quality is relative to what the customer asked.

### how to audit AI support agent answers at scale

Steps 4 and 5. For AI agents specifically, also track hallucination rate as its own dimension: answers that sound right but cite nonexistent policies or features.

## Why it happens

Manual QA samples too little to catch systematic problems, and by the time a human reviewer finds a pattern it has affected thousands of tickets. Automated scoring flips the ratio: the machine watches everything and humans investigate the interesting bits. Teams that do this find quality problems in days that used to take quarters to surface.

## Edge cases

- The judge disagrees with humans on tone: tone is the hardest dimension to calibrate. Weight it lower, or keep tone as human-only review.
- Multilingual answers: calibrate per language, a judge calibrated on English will misjudge formality norms in other languages.
- Gaming: if agents know the rubric, they will optimize for the rubric. Keep some human review random and rotate rubric emphasis.
- Cost: scoring every ticket with a frontier model adds up. Use a smaller model for the first pass and escalate low-confidence scores to the bigger one.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_HW2oO4-ywj6zmO5fuFZPYw
