Most language pages in this industry say the same thing: translation is hard, culture is subtle, machines miss nuance. It is not very useful, and buyers have stopped reading it.
So here is the uncomfortable version. In the hardest independent evaluation run in 2025 — thirty language pairs, sixty systems, professional annotators marking real errors — the human translation placed in the top group in six pairs out of fifteen. One frontier model was in the top group in fourteen of sixteen and beat the human translation outright in ten.
In most language pairs, for most general content, the machine is now at least as good as a professional translator. Pretending otherwise costs you credibility with anyone who has read the research.
The pairs where it is not are specific, and the reasons are specific. They fall into four groups, and the group tells you what kind of failure to expect and what kind of check catches it.
The four groups
One thing that runs across all of them
Modern workflows score every segment automatically and send only low-scoring segments to a person. That works when the score is trustworthy.
In the second group above, published research has found automatic metrics ranking the worst system first, and scoring models that cannot tell two different texts apart. A quality gate is only as good as the score it opens on — and in the languages where the risk is highest, the score is weakest.
That is the gap we work in.