how to test a support chatbot before launch
A pre-launch test plan for support chatbots: adversarial phrasings, the top ticket intents, fallback behavior, and handoff timing. Use before any bot launch or major bot update, or when auditing a bot that went live without testing. Not for building the bot, training models, or marketing the launch.
TL;DR
Test the bot like a grumpy customer, not like its trainer. Run your top 20 real ticket intents through it, then attack it with typos, slang, multi-part questions, and "talk to a human" phrasings. Measure three things: correct-answer rate, fallback rate, and time to handoff. If the bot cannot name what it does not know, it is not ready.
The query
how to test a support chatbot before launchUse this when
- A chatbot is about to launch or get a major update
- A bot went live without real testing and complaints are rising
- You are evaluating a bot vendor's claims
- Writing acceptance criteria for a bot project
Not for
- Building or training the chatbot
- Model fine-tuning
- Marketing the launch
- Post-launch monitoring (different checklist)
Steps
1. Build the test set from real tickets
Pull your top 20 intents by volume from the last 90 days of tickets. For each, write the question three ways: the clean version, the way a frustrated user phrases it, and the way a non-native speaker phrases it. Add 30 adversarial cases: typos, slang, two questions in one message, and topic changes mid-conversation.
Expected output: a test set of 90+ cases grounded in real volume.
2. Test phrasing variants, not just intents
The bot will know "reset my password." It will fail on "i cant get in it says my thing expired." Run every variant. Bots fail on phrasing far more than on intent coverage.
Expected output: per-variant pass or fail, not just per-intent.
3. Test the fallback and the handoff
Deliberately ask things outside the bot's scope. It must admit it does not know, offer a human, and pass context. A bot that hallucinates an answer to an unknown question fails this test catastrophically.
Expected output: 100 percent clean fallbacks on out-of-scope questions.
4. Load test with simultaneous sessions
Run 50 to 100 concurrent conversations. Watch for slowdowns, crossed contexts between users, and rate-limit errors. A bot that is perfect at one conversation and broken at fifty is not launch-ready.
Expected output: latency and error numbers at expected peak load.
5. Run a two-week pilot with human review
Launch to 10 percent of traffic with a human reviewing every bot answer before it sends, or reviewing full transcripts daily. Fix the top failure patterns, then widen.
Expected output: reviewed transcripts and a fix list before full launch.
Template: the test matrix
Intent | Clean phrasing | Frustrated phrasing | Non-native phrasing | Result
password reset | PASS | PASS | FAIL (did not parse "thing expired") | fix alias
refund status | PASS | PASS | PASS | ok
cancel account | PASS | FAIL (offered FAQ, no action) | PASS | fix handoff
... | ... | ... | ... | ...
Out-of-scope probe | Bot response | Clean fallback?
"can you lower my taxes" | "I cant help with that. Want me to connect you?" | yesVariant phrasings
chatbot QA checklist
Steps 1 through 3 as the checklist core.
testing a customer service bot
Full sequence. Emphasize step 5: no launch without a pilot.
bot gives wrong answers sometimes
Steps 2 and 3. It is a phrasing or fallback gap, not a mystery.
Why it works
Bot demos use clean phrasings. Real users do not. The gap between demo phrasing and ticket phrasing is where every failed launch lives. Testing with real ticket language, adversarial inputs, and a human-reviewed pilot closes that gap before customers find it.
Edge cases
- Multilingual: test each supported language separately. Quality is never uniform.
- Sarcasm and jokes: the bot should not take "great, another broken update" as praise.
- Mid-conversation topic changes: users pivot. The bot must follow or hand off.
- The pilot reviewers disagree: define "correct" as resolving the user's ask, not matching a script.
Provenance
Resolved from the public thread: https://vectle.com/posts/pstu96k5t7-xxSSVoths-d1A
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.