## TL;DR
If an AI agent touches customer replies, you need a standing audit, not a one-time test. Sample tickets across topics, grade them blind against a fixed checklist, and trend the scores weekly. Every failure goes back into the agent's instructions or guardrails. The audit is what lets you trust the automation instead of just hoping.

## The query

```text
how to audit an AI support agent's answers
```

## Use this when

- An AI agent drafts or sends customer replies
- Leadership asks for proof the bot is safe
- Customers complain about AI answers
- You changed the agent's instructions and need to verify the effect

## Not for

- Choosing which AI model to use
- Prompt engineering tutorials
- Auditing human agents
- One-time pre-launch testing

## Steps

### 1. Sample smart, not random

Pull a stratified sample: top ticket topics, escalations, low-CSAT tickets, and edge cases like billing and security. Random sampling over-represents easy password resets and misses the risky stuff. Aim for 50 to 100 tickets per audit round.

Expected output: a sample set weighted toward risky and representative tickets.

### 2. Build the checklist

Grade five things: factual accuracy, completeness of the answer, tone, policy compliance, and whether any fact was invented. Write down what a pass looks like for each, with an example. A checklist without examples drifts within a month.

Expected output: a five-dimension checklist with pass and fail examples.

### 3. Grade blind

Reviewers should not know whether the AI or a human wrote the reply. Blind grading kills both the "bots are terrible" bias and the "looks fine to me" laziness. Two graders per ticket, and they discuss disagreements.

Expected output: unbiased scores with inter-grader agreement you can defend.

### 4. Score and trend weekly

Turn the checklist into a percentage and plot it over time. A single audit tells you nothing; the trend tells you whether instruction changes help or hurt. Break out scores by dimension so you know what is actually failing.

Expected output: a weekly trend line per dimension, not just an overall number.

### 5. Feed failures back into the instructions

Every failed ticket gets a fix: an instruction update, a new guardrail, or an escalation rule. Log the failure, the fix, and the re-test. An audit that does not change the agent is theater.

Expected output: a failure log where each entry has a fix and a verified re-test.

## Ready-to-use audit checklist

```text
Ticket ID:
Grader:

1. FACTS: every claim in the reply is true and matches our docs.   pass / fail
2. COMPLETE: the customer's actual question got answered.          pass / fail
3. TONE: warm, professional, no blame, no robotic phrasing.        pass / fail
4. POLICY: no promises we can't keep, no invented timelines.      pass / fail
5. GROUNDED: no invented order numbers, dates, or feature names.  pass / fail

Fail notes (quote the exact sentence that failed):
```

## Variant phrasings

### AI chatbot quality assurance checklist

The checklist above, run monthly instead of weekly. Chatbots drift as content changes, so the cadence matters more than the checklist wording.

### how to test AI customer service responses

Steps 1 through 3 are the test. Step 4 turns a test into an ongoing practice.

### auditing automated support replies

Same five steps. "Automated" covers macros and templates too: blind-grade those the same way.

## Why it happens

AI replies read fluently even when they are wrong, which makes spot-checking useless. Humans skim, see a polite answer, and move on. The confident tone hides invented facts, skipped questions, and policy violations. Only structured, blind, repeated grading catches the pattern, because the failure mode is plausible, not obvious.

## Edge cases

- The AI is right but unhelpful: completeness catches this. A true-but-useless answer still fails.
- Graders disagree constantly: your checklist examples are vague. Rewrite them with real tickets.
- Scores drop after an instruction change: roll back first, diagnose second. The audit trend is your canary.
- Low sample sizes: below 30 tickets the trend is noise. Audit less often rather than with tiny samples.
- The AI handles a language nobody on the team speaks: get a native-speaking reviewer for that slice, or exclude it from auto-send.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_1fbxYly0iC3MTSSJvRFoQQ
