Chinese is where the benchmarks and the real world disagree

On the standard benchmark, machines now beat human translators into Simplified Chinese. On real commercial content, Chinese scores among the lowest of any major language. Both things are true, and the difference is where the work is.

--:--:-- [PST]
88.4 vs 82.1
best machine system against the human translator, English into Simplified Chinese — the machine wins
WMT25
72.2 / 100
Simplified Chinese quality score on real client projects — lowest of ten languages measured
Vendor scoreboard, 5,632 evaluations
S01

We will start with the finding that argues against us

In the 2025 WMT evaluation, English into Simplified Chinese was one of the pairs where the machine clearly beat the human: 88.4 against 82.1. Chinese has vast training data behind it, and Chinese technology companies are among the strongest builders of translation systems in the world.

If someone tells you Chinese and Arabic are equally difficult for machine translation, the best available evidence does not support it. We would rather say that than sell you a review you do not need.

S02

And the finding that complicates it

Benchmarks test general document translation. Real client work is software, marketing, product and regulatory content that needs far more adaptation.

On one published scoreboard covering 5,632 evaluations across 97 real client projects, Simplified Chinese scored 72.2 — the lowest of the ten languages measured, below Arabic, below Korean, below Thai. Traditional Chinese scored 74.4. That source is weaker than the academic benchmark — it is published by a company that sells post-editing, and some per-language samples are small — so treat the number as directional rather than definitive. But it matches what reviewers see: the gap between benchmark performance and shipped-content performance is wider in Chinese than the headline numbers suggest.

S03

What actually goes wrong

Simplified and Traditional are separate deliverables, not a character conversion. They differ in vocabulary, terminology and idiom, and the 2026 benchmark now treats them as separate targets. Converting characters and calling it localised is one of the most common and most visible failures in this language.

No spaces between words. The system has to decide where words begin and end before it can translate anything. Segmentation errors propagate into everything downstream.

No tense marking on verbs. A documented failure has machines rendering "has actively promoted" as "actively promotes" — a change of meaning that is completely invisible to a reviewer who does not read Chinese, and that matters enormously in a legal or regulatory claim.

Classifiers. Chinese requires a measure word between a number and a noun, chosen according to the shape and category of the thing being counted. Errors here are immediately obvious to native readers and mark the text as machine-produced.

S04

Where the stakes are highest: patents

This is the strongest commercial position in Chinese, and it is worth being specific about why.

China is the largest source of international patent filings in the world. Patent translation must be certified accurate, and the consequence of an error is not embarrassment — it is a narrowed claim, an invalidated priority date, or an unenforceable right. Claim language is written to be precise at the level of the individual word, and the features of Chinese described above — segmentation, tense, classifiers, term consistency — bear directly on claim scope.

Patent work also carries a rate premium of roughly 60 to 200% over general translation, for the straightforward reason that the cost of getting it wrong is orders of magnitude higher than the cost of getting it right.

Patents →

S05

What we do

  • We treat Simplified and Traditional as separate jobs, with separate reviewers and separate terminology decisions.
  • We check tense, aspect and modality explicitly, because these are the errors that survive a fluency read and change what a sentence commits you to.
  • We hold terminology across large document sets, which in patent and regulatory families is most of the value.
  • We tell you when you do not need us. For general content in Simplified Chinese, the engines are good. For claim language, regulatory submissions and anything where a term is load-bearing, they are not — and the difference is not visible from the outside.
S06

Send us a Chinese file

Particularly if it is patent, pharmaceutical or regulatory content. We will return it reviewed with terminology and claim-critical language assessed separately.

Send us a Chinese file