What the evidence says
We will start with the part that does not suit us.
In the 2025 evaluation, the best system into Simplified Chinese scored 88.4 against a professional human translation's 82.1. The machine won, clearly, and the human reference ranked outside the top ten systems. Anyone telling you that machines cannot handle Chinese is not reading the evidence.
Now the other dataset. On 5,632 evaluations of real client projects — software, marketing, product content — Simplified Chinese scores 72.2, the lowest of the eighteen languages measured. Traditional Chinese scores 74.4. Both sit below Russian, below Japanese, far below Italian and French.
The same language, measured two ways, comes first and last.
Why the two answers differ
The benchmark tests general document translation. Real client work is interface strings, product names, marketing and support content — where terminology has to hold across thousands of segments, brand voice matters, and the correct output is often an adaptation rather than a translation. Chinese is excellent at the first and weak at the second.
There is also a technical layer worth knowing. Simplified and Traditional are separate deliverables, not a character conversion: they differ in vocabulary, terminology and idiom, and the international evaluation now treats them as separate targets. Chinese has no spaces between words, so the system must decide where words begin before it can translate. Chinese does not mark tense on verbs — a documented failure has machines rendering "has actively promoted" as "actively promotes," which is invisible to any reviewer who does not read Chinese. And classifier words between numbers and nouns are chosen by the shape and category of the thing counted; errors there are immediately obvious to a native reader and completely invisible to everyone else.
One more, and it is the kind of thing that decides an engine: on real content, the engine that wins Simplified Chinese is not the engine that wins most other languages.
What it means for your workflow
If your Chinese engine decision was made on benchmark performance, it was made on the wrong evidence. And a benchmark-derived quality threshold will be set far too high for your actual content.
What we do
We review Chinese on your content type, not on general text. We check tense and aspect, classifiers, and terminology consistency across the whole deliverable. We treat Simplified and Traditional as two jobs. And we will benchmark your engine against alternatives on your material — which, in Chinese, is the change most likely to pay for itself.