Quick answer — Use automatic translation scores for comparison and triage, structured human review for specific errors, and operational measures for the cost of getting content ready. A single score isn't a release decision.
Vitra TMS brings reviewed translations, terminology, style and context into the work. Measure the finished result as well as the draft.
BLEU, COMET and MQM answer different questions
BLEU is a reference-based metric built around word-sequence overlap, with a brevity penalty. It can help compare systems on a common test set, but it isn't a direct measure of whether one sentence is safe to publish. The original BLEU paper explains the method.
COMET is a family of learned evaluation models. Reference-based models and reference-free quality-estimation models both exist. It's inaccurate to describe every COMET score as similarity to a reference. Check the chosen model's language coverage and evaluation setup in the official COMET documentation.
MQM provides a framework for classifying translation errors and their severity. A reviewer can identify a wrong term, missing meaning or locale-convention error instead of giving an unexplained overall rating. The MQM Council's scoring guidance also makes clear that scoring depends on the evaluation design.
| Method | Useful for | What still needs checking |
|---|---|---|
| BLEU | Comparing outputs against the same references | Meaning, legitimate alternative wording and finished context |
| COMET | Model-based comparison or quality estimation | The exact model, supported languages and important error cases |
| MQM-based review | Naming and classifying observed errors | Reviewer consistency, sampling and acceptance criteria |
| Operational measures | Understanding review work and release performance | Whether comparable work was measured |
Don't compare a BLEU number with a COMET number as if they shared a scale. Keep the tool version, model, language pairs and test set with the result.
Start with the error that matters
Consider a fictional product instruction: “Do not charge the battery below freezing.”
A translation that drops “not” can sound fluent while reversing the instruction. A slightly awkward sentence that preserves the warning is a different kind of problem. Count them separately and let the responsible product reviewer judge their consequence.
Write the acceptance criteria before reviewing the batch. For a campaign, that might mean a correct offer, approved brand terms and no clipped disclaimer. For a support article, it might mean an accurate sequence of steps and working links. The localization QA checklist template is a starting point, not a certificate of quality.
Track four operational measures
Keep the definitions stable from one reporting period to the next.
| Measure | Example definition | Caution |
|---|---|---|
| Review effort | Reviewer minutes per comparable asset | A difficult new format can take longer without worse translation |
| Defects found after release | Confirmed language defects in released assets | Few reports may mean weak reporting, not perfect content |
| Rework | Assets returned for another correction pass | Separate source changes from translation errors |
| Time to release | Approved source to approved localized output | Faster delivery can hide a reduced review scope |
Reviewer edit distance can supplement these measures, but it doesn't prove that adaptive memory is improving. A new reviewer, a revised style guide or a harder source can change the amount of editing. Look at the actual corrections before explaining a trend.
Run a fair before-and-after test
Use a small representative set with the same product context, glossary and reviewer instructions. Include a difficult example rather than choosing only easy material.
Record the first output, the findings and the final approved version. Repeat the exercise after changing one part of the process: the glossary, a style rule or how context is supplied. This is an illustrative evaluation method, not a published Vitra benchmark.
If terminology errors fall but review time rises, inspect why. The reviewer may now be checking the finished layout more carefully. That can be valuable work even though the draft score hasn't changed.
Use the TMS evaluation scorecard to keep the evidence with the decision. For recurring work, translation quality and review should connect automated checks to named reviewers rather than turn a threshold into automatic approval.
If you add an LLM judge, test its findings against qualified reviewers on the same examples. Treat it as a signal to investigate, not an independent certificate of accuracy.
FAQ
What is the difference between BLEU and COMET? BLEU measures overlap with reference translations. COMET uses learned evaluation models; some use references and others estimate quality without one. Specify the model and test set before comparing scores.
Is a high translation quality score enough to publish? No. Use scores to compare outputs or prioritize review, then check the meaning, required terminology and finished content. A score does not establish legal accuracy, brand approval or a usable layout.
Does lower reviewer edit distance prove the memory is learning? Not by itself. Edit distance also changes with source difficulty, reviewer preferences and style decisions. Compare equivalent work and inspect which errors disappeared before attributing the change to memory.
What should a translation quality dashboard show? Show error types and severity alongside review effort, rework and time to release. Record the language, content type, sample size and evaluation method so the numbers can be interpreted.



