Where machines fail, and where they do not

In most pairs the machine now matches a professional translator. These are the ones where it does not — and the reasons are specific.

--:--:-- [PST]
6 of 15
pairs where the human translation placed in the top group in the 2025 evaluation
WMT25
14 of 16
pairs where one frontier model placed in the top group, beating the human outright in ten
WMT25
4 engines
win at least one language across eighteen measured on real client content
Engine index, 2026
S01

Most language pages in this industry say the same thing: translation is hard, culture is subtle, machines miss nuance. It is not very useful, and buyers have stopped reading it.

So here is the uncomfortable version. In the hardest independent evaluation run in 2025 — thirty language pairs, sixty systems, professional annotators marking real errors — the human translation placed in the top group in six pairs out of fifteen. One frontier model was in the top group in fourteen of sixteen and beat the human translation outright in ten.

In most language pairs, for most general content, the machine is now at least as good as a professional translator. Pretending otherwise costs you credibility with anyone who has read the research.

The pairs where it is not are specific, and the reasons are specific. They fall into four groups, and the group tells you what kind of failure to expect and what kind of check catches it.

S02
S03

One thing that runs across all of them

Modern workflows score every segment automatically and send only low-scoring segments to a person. That works when the score is trustworthy.

In the second group above, published research has found automatic metrics ranking the worst system first, and scoring models that cannot tell two different texts apart. A quality gate is only as good as the score it opens on — and in the languages where the risk is highest, the score is weakest.

That is the gap we work in.

S04

Send us a file and a language pair

We will tell you what your engine is getting wrong, and whether the score your pipeline is trusting means anything in that language.

Send us a file to evaluate