You may know this as LQA, linguistic quality assurance, Auto LQA, linguistic QC, language quality review, or MQM, DQF-MQM and J2450 scoring.
Why this became a separate product
Five years ago, checking translation was a step inside a translation job. You translated, someone proofread, and the proofreading was part of the price nobody itemised.
That has changed across the entire industry. Every major language provider now sells quality checking as its own named product, with its own price and in some cases its own patents. The reason is straightforward: when a machine produces most of the words, the scarce thing is no longer the words. It is credible evidence that they are fit for purpose.
The structural argument
An audit performed by the party being audited is not evidence.
That principle is uncontroversial in accounting, in safety engineering and in clinical research, and it applies exactly the same way to language. When the company that produced the translation also grades it, the grade carries the producer's interest in it. When the company that built the engine also certifies the engine's output, the certificate is marketing.
We do not translate the content we review, and we do not sell an engine. That is not a slogan — it is the entire reason the assessment is worth something to the person you hand it to.
What a review contains
- Error identification. Every issue found, located to the segment, with the source and target shown side by side.
- Categorisation. Each error typed against MQM: accuracy, fluency, terminology, style, locale convention, design and markup.
- Severity. Minor, major or critical — because ten cosmetic issues and one flipped negation are not the same finding, and a single score that averages them is worse than useless.
- A quality score. A number you can track across languages, suppliers and time.
- A verdict. Whether, in our assessment, the content is fit for release at the standard agreed — stated plainly, with the reasoning.
- Trend reporting. For ongoing programmes: cumulative reports by language, supplier, content type and error category, so patterns surface before they become incidents.
What you get
A score. A categorised error report. A cumulative view across languages and projects. Not a corrected file — if you want the file corrected, that is post-editing and it is priced differently.
We score against the MQM framework: errors classified by type — accuracy, fluency, terminology, style — and by severity — minor, major, critical. Where you use DQF-MQM or, in automotive, J2450, we work in your framework rather than ours.
What it is for
Three situations.
You need to know whether an engine, a vendor or an internal team is producing output good enough to publish — and you need the answer in a form you can show someone else. The published benchmarks are general-domain and the vendors' own numbers are marketing. Your content is the only test that matters.
Your automated LQA has flagged a set of segments and you need a qualified human decision on them, not another model's opinion.
You are in a regulated sector and the evidence that review happened is itself part of the deliverable.
How we work with automated quality checks
We assume they exist. Buyers increasingly run quality estimation, an automated LQA pass, or both, and route only what falls below threshold to a person. That is a sensible design and we are built to sit at the end of it. Send us the exceptions.
Where it matters, we will also tell you when we think the threshold is set wrong — which is a conversation the party that configured the threshold cannot easily have with you.
Comparing engines and vendors
The same review method, run comparatively: your current engine against alternatives, or one vendor against another, on your own content rather than on a benchmark. No single engine wins everywhere — published comparisons show different engines winning different languages, with gaps between the best and second-best choice large enough to decide whether output is publishable. Most buyers standardised on one engine and have never re-checked. We do not sell an engine, so we have no answer to defend.
Independence
We do not sell a translation engine. We do not sell a platform. On the work we review, we were not the producer. An assessment written by the party that produced the output is a weaker form of evidence, and regulated buyers have started saying so out loud.
We do not independently assess output we produced ourselves. Where we have translated or post-edited a file, the impartial review of it has to come from somebody else.
How we run it
- Reviewers are matched by subject as well as language. A device reviewer knows devices. For safety-critical content, a generalist is not sufficient, and we will not staff it that way.
- Named reviewers with recorded qualifications. Training, credentials, subject expertise, years in domain. If your auditor asks, you have an answer.
- Scoring is calibrated. Reviewers are checked against one another on the same content, so that a score means the same thing in Arabic as in Japanese.
- Findings are versioned. Which source version, which target version, which date.
- We report what we find. Including when the answer is that the content is fine and you are over-spending on review. We would rather say that than bill for the larger programme.
What it costs
Per hour, or per segment on exception work. Not per word — per-word pricing on quality work assumes machine output quality is constant, and it is not.
Sampling programmes are usually priced as a recurring monthly engagement. Release-critical full reviews are priced per project. Ask us and we will give you a number against your real volumes.
Need decisions, not a report?
Need decisions on specific flagged segments rather than a report on a body of work? That is Exception Review.
Delivered, or staffed
Send us the work and we return the deliverable described above. Or place our reviewers into your own process and tooling and run them under your rubric, your schema and your brand. Same people, same qualification records, different contract. Tell us which you want and we will price it that way.
Need the evidence as part of the deliverable? See documentation and evidence.
Common questions
How is this different from proofreading?
Proofreading produces a corrected file. Language quality assurance produces a categorised, scored assessment — what was wrong, what type of error it was, how severe, and what that means for whether the content can be released. You can buy the corrections as well, but the report is the deliverable.
Will you review work produced by another supplier?
Yes. That is most of what this service is for. We assess the output on its merits and we have no relationship with whoever produced it.
Do you sample or check everything?
Both are available. Sampling is right for ongoing supplier monitoring; full review is right for release-critical content. We will tell you which one your content justifies rather than selling you the larger number.
What framework do you score against?
MQM by default, because it is the industry standard and your clients will recognise it. We can work to DQF-MQM, to J2450 for automotive, or to a client's own scorecard if one already exists.