support QA rubric for agent answers
A scoring rubric for grading support replies across accuracy, completeness, tone, and process: dimension definitions, what each score looks like, and how to calibrate graders. Use when starting a QA program, coaching underperformers, or proving quality to leadership. Not for automated scoring, CSAT analysis, or performance reviews without coaching.
TL;DR
A QA rubric turns "that reply felt off" into scores you can coach to. Grade four to six dimensions on a 1 to 5 scale, with a written example of what each score looks like. Calibrate your graders on the same tickets or the scores are meaningless. Sample a handful of tickets per agent per month, and always pair a low score with the example that earned it.
The query
support QA rubric for agent answersUse this when
- Starting a QA program from zero
- Coaching an underperforming agent
- Proving answer quality to leadership
- Graders disagree about what "good" means
Not for
- Automated or AI scoring of replies
- CSAT and survey analysis
- Performance reviews without a coaching plan
- Grading chatbot replies, which need their own checklist
Steps
1. Pick four to six dimensions
Accuracy, completeness, tone, personalization, and process compliance cover almost everything. More than six and graders stop reading. Fewer than four and you miss whole failure modes. Name them in plain words your agents use.
Expected output: a short dimension list everyone can recite.
2. Write what each score looks like
For every dimension, describe a 1, a 3, and a 5 with a real ticket example. "Tone: 5" means nothing; "Tone 5: warm, names the customer's frustration, no blame" means everything. The examples are the rubric. Without them, two graders will never agree.
Expected output: anchored descriptions with quoted examples for 1, 3, and 5.
3. Calibrate the graders
Have every grader score the same ten tickets, then argue about the disagreements until the scores converge. Repeat quarterly. Uncalibrated graders produce noise, and noise dressed as data is worse than no QA at all.
Expected output: graders whose scores agree within one point on calibration tickets.
4. Sample fairly
Five to ten tickets per agent per month, pulled randomly across topics and channels. Never let agents or leads hand-pick the sample. Include at least one hard ticket per agent; grading only easy ones flatters everyone.
Expected output: a random, representative sample per agent, every month.
5. Coach from the examples, not the number
Share the score with the two or three ticket excerpts that drove it, and one concrete change for next time. "Your tone scored 2.4" changes nothing. "Here is where you blamed the customer, try this phrasing instead" changes behavior.
Expected output: every QA review ends with examples and one actionable change.
Ready-to-use rubric skeleton
Ticket ID: Agent: Grader:
ACCURACY (facts right, matches docs)
5: every claim correct and current
3: correct but missing a key detail
1: wrong fact or outdated info
COMPLETENESS (the actual question got answered)
5: answered fully, anticipated the follow-up
3: answered, but the follow-up was predictable
1: answered a different question than asked
TONE (warm, professional, no blame)
5: names the frustration, owns the next step
3: polite but distant
1: blames the customer or sounds robotic
PROCESS (tags, links, escalation rules followed)
5: tagged right, KB linked, no missed step
3: answer fine, admin sloppy
1: skipped a required step
Coaching note (one change for next time):Variant phrasings
customer support quality scorecard
Same rubric, formatted as a scorecard for leads. The content does not change; the audience does.
how to grade support agent responses
Steps 2 and 5 are the core. Most teams asking this need the anchored examples, not the process.
QA criteria for helpdesk tickets
The five steps, with process compliance weighted heavier for regulated industries.
Why it happens
Without a rubric, quality feedback is vibes: "be warmer", "try harder". Agents cannot act on vibes, so nothing changes and leads stop giving feedback. The rubric converts taste into criteria, and criteria into coaching. The calibration step exists because every grader's taste differs, and uncalibrated taste presented as scores just creates resentment.
Edge cases
- Agents game the rubric: they will, by hitting the letter of each dimension. Rewrite anchors yearly from fresh tickets.
- One grader is consistently harsh: calibration catches this. Pair them with a lenient grader until they converge.
- High scores but low CSAT: your rubric measures the wrong things. Add the dimension customers actually care about.
- Small teams: one grader is fine, but rotate who grades to avoid a single person's taste becoming law.
- Union or works-council environments: check the rules before scoring individuals. Some places require works-council agreement for individual QA.
Provenance
Resolved from the public thread: https://vectle.com/posts/pst_od0zlZ0hzYqJ3tLrdwKLEQ
Maintainer review
No maintainer verification is recorded for this version.
This records the version a maintainer checked. It does not assert that the version is the latest upstream release.