support bot stated a policy that doesn't exist: grounding check
Stop a support bot from inventing policies: ground every policy claim in your real policy docs before it reaches the customer. Use when a bot states policies that don't exist, when customers quote the bot back at your team, or when you're designing grounding for a support agent. Not for wrong product facts, tone problems, or human agent errors.
TL;DR
Bots invent policies because the prompt says 'be helpful' and the model fills gaps with plausible-sounding rules. The fix is grounding: every policy statement must cite a passage from your actual policy docs, and anything without a citation gets blocked or hedged. Uncited policy claims are the failure; the check is the cure.
The query
support bot stated a policy that doesn't exist: grounding checkUse this when
- A support bot states policies that don't exist
- Customers quote the bot's invented policy back at your team
- You're designing grounding for a support agent
Not for
- Bots giving wrong product facts (different grounding target)
- Tone or style problems
- Human agents misstating policy
Steps
1. Collect the invented policies
Pull recent conversations where the bot stated a policy and check each claim against your real docs. Log the exact invented wording. You need real examples to test against, not hypotheticals.
Expected output: a list of invented policy statements with the bot's exact wording.
2. Build a policy source the bot must cite
Put your real policies in a retrievable store: refunds, shipping, warranties, account rules. The bot's instructions should require it to base policy answers on retrieved passages and to say 'I don't have a policy on that' otherwise.
Expected output: a policy corpus the bot retrieves from, wired into its instructions.
3. Add a citation check before sending
After the bot drafts a reply, verify that every policy claim traces to a retrieved passage. Claims without a source get rewritten as 'let me check with the team' instead of sent. This is the grounding check proper.
Expected output: no policy claim reaching a customer without a source passage.
4. Test with adversarial questions
Ask the bot about policies you don't have: 'what's your policy on X' for X that doesn't exist. The correct answer is always a hedge or a handoff, never an invented rule. Run this suite on every prompt change.
Expected output: a red-team suite the bot passes with hedges, not inventions.
Variant phrasings
ai support bot hallucinating policy
Steps 2 and 3: the retrievable corpus plus the citation check.
chatbot making up refund policy
Step 4's adversarial suite catches the invention habit.
Why it happens
Language models are trained to answer, and 'I don't know' is underrepresented in training. Faced with a policy question and no retrieved passage, the model does what it's rewarded for: produces a confident, plausible answer. Grounding works because it changes the task from 'answer' to 'answer from these passages', which the model can actually do.
Edge cases
- Policy docs change. Version the corpus and re-test after every policy update.
- Hedging too much is its own failure. Tune so common policies answer crisply and edge cases hedge.
- The check needs the retrieved passages logged. Without logs you can't audit what grounded what.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_YLcDMjFvMQoZGY1BihNeyQ