## TL;DR
Test the bot like a grumpy customer, not like its trainer. Run your top 20 real ticket intents through it, then attack it with typos, slang, multi-part questions, and "talk to a human" phrasings. Measure three things: correct-answer rate, fallback rate, and time to handoff. If the bot cannot name what it does not know, it is not ready.

## The query

```text
how to test a support chatbot before launch
```

## Use this when

- A chatbot is about to launch or get a major update
- A bot went live without real testing and complaints are rising
- You are evaluating a bot vendor's claims
- Writing acceptance criteria for a bot project

## Not for

- Building or training the chatbot
- Model fine-tuning
- Marketing the launch
- Post-launch monitoring (different checklist)

## Steps

### 1. Build the test set from real tickets

Pull your top 20 intents by volume from the last 90 days of tickets. For each, write the question three ways: the clean version, the way a frustrated user phrases it, and the way a non-native speaker phrases it. Add 30 adversarial cases: typos, slang, two questions in one message, and topic changes mid-conversation.

Expected output: a test set of 90+ cases grounded in real volume.

### 2. Test phrasing variants, not just intents

The bot will know "reset my password." It will fail on "i cant get in it says my thing expired." Run every variant. Bots fail on phrasing far more than on intent coverage.

Expected output: per-variant pass or fail, not just per-intent.

### 3. Test the fallback and the handoff

Deliberately ask things outside the bot's scope. It must admit it does not know, offer a human, and pass context. A bot that hallucinates an answer to an unknown question fails this test catastrophically.

Expected output: 100 percent clean fallbacks on out-of-scope questions.

### 4. Load test with simultaneous sessions

Run 50 to 100 concurrent conversations. Watch for slowdowns, crossed contexts between users, and rate-limit errors. A bot that is perfect at one conversation and broken at fifty is not launch-ready.

Expected output: latency and error numbers at expected peak load.

### 5. Run a two-week pilot with human review

Launch to 10 percent of traffic with a human reviewing every bot answer before it sends, or reviewing full transcripts daily. Fix the top failure patterns, then widen.

Expected output: reviewed transcripts and a fix list before full launch.

## Template: the test matrix

```text
Intent | Clean phrasing | Frustrated phrasing | Non-native phrasing | Result
password reset | PASS | PASS | FAIL (did not parse "thing expired") | fix alias
refund status | PASS | PASS | PASS | ok
cancel account | PASS | FAIL (offered FAQ, no action) | PASS | fix handoff
... | ... | ... | ... | ...

Out-of-scope probe | Bot response | Clean fallback?
"can you lower my taxes" | "I cant help with that. Want me to connect you?" | yes
```

## Variant phrasings

### chatbot QA checklist

Steps 1 through 3 as the checklist core.

### testing a customer service bot

Full sequence. Emphasize step 5: no launch without a pilot.

### bot gives wrong answers sometimes

Steps 2 and 3. It is a phrasing or fallback gap, not a mystery.

## Why it works

Bot demos use clean phrasings. Real users do not. The gap between demo phrasing and ticket phrasing is where every failed launch lives. Testing with real ticket language, adversarial inputs, and a human-reviewed pilot closes that gap before customers find it.

## Edge cases

- Multilingual: test each supported language separately. Quality is never uniform.
- Sarcasm and jokes: the bot should not take "great, another broken update" as praise.
- Mid-conversation topic changes: users pivot. The bot must follow or hand off.
- The pilot reviewers disagree: define "correct" as resolving the user's ask, not matching a script.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_u96k5t7-x_xSSVoths-d1A
