## TL;DR
Machine translation is good enough for support until it is not, and the failures are silent: the reply reads fine in English and is wrong in the customer's language. Check quality with back-translation spot checks, lock product terms in a glossary the MT must not translate, and route refunds, legal, and account-security replies to human review. Never trust a reply you cannot read.

## The query

```text
machine translation quality checks for support replies
```

## Use this when

- Supporting languages no agent on the team speaks
- MT errors have caused customer confusion
- Setting up multilingual support for the first time
- Auditing an existing MT pipeline

## Not for

- Choosing an MT provider
- Training or fine-tuning translation models
- Marketing or website localization
- Real-time interpretation

## Steps

### 1. Build a glossary of never-translate terms

Product names, plan names, button labels, error codes. MT loves to "helpfully" translate these, and "helpfully" translated button labels send users hunting for buttons that do not exist. Lock the terms per language.

Expected output: a glossary file per supported language.

### 2. Back-translate a weekly sample

Take 20 translated replies, translate them back to English with a different engine, and compare against the original. Meaning shifts show up fast: polite hedges become blunt demands, conditionals become promises.

Expected output: a weekly back-translation report with flagged shifts.

### 3. Score with a simple rubric

Three points: meaning preserved, tone appropriate, glossary respected. Anything scoring below 3 gets rewritten by a human and fed back as a training example. Do not overcomplicate the rubric; speed of review matters more than granularity.

Expected output: scored samples and a rewrite queue.

### 4. Define the human-review triggers

Refunds, legal threats, account security, cancellations, and anything with a number in it (amounts, dates, SLAs). These go to a bilingual reviewer or a professional service before sending. Everything else can go direct with sampling.

Expected output: a written trigger list.

### 5. Tell the customer the reply was translated

One line: "This reply was automatically translated." It buys forgiveness for awkward phrasing and invites the customer to say when something does not make sense.

Expected output: a disclosure line in every MT reply template.

## Template: the review rubric

```text
Sample review (weekly, 20 replies):
[ ] Meaning preserved: the back-translation matches the original intent
[ ] Tone appropriate: polite and professional in the target language
[ ] Glossary respected: product terms untranslated, correct forms used
[ ] No invented specifics: amounts, dates, and steps match the original

Score 4/4: send. Score 3/4: fix and send. Below: human rewrite.
High-stakes topics (refund, legal, security): always human review first.
```

## Variant phrasings

### MT QA for support

Steps 2 and 3. Sampling plus the rubric.

### translation quality in support tickets

Full sequence. Step 1 first: glossary failures are the most common.

### auto-translated reply confused the customer

Steps 2 and 4. Check what shifted, then decide if the topic needs human review going forward.

## Why it works

The dangerous MT errors are fluent: the grammar is perfect and the meaning is wrong. Back-translation catches meaning shifts without needing a bilingual reviewer for every reply, and the glossary kills the most common systematic error. The disclosure line turns residual awkwardness from a trust problem into a non-issue.

## Edge cases

- Right-to-left languages: layout and number formatting need separate checks.
- Formality levels: some languages need formal address for support. Set it explicitly.
- Idioms: "ballpark figure" and "touch base" do not survive translation. Write plainly.
- Mixed-language tickets: detect the customer's language per message, not per ticket.

## Provenance

Resolved from the public thread: https://vectle.com/posts/pst_5-PHXM7rDs4__pJznaGKMQ
