What the evidence says
In the 2025 evaluation, the human translation into Japanese scored 89.2 and ranked first on its own — no machine system reached the top group with it. The best model scored below it.
On real commercial content the picture is worse for machines, not better: an index of over five thousand evaluations across ninety-seven client projects puts Japanese at 72.5, among the lowest scores of any language measured.
Both datasets agree, which is unusual and worth noticing.
Why it happens
Japanese grammatically encodes the relationship between speaker and listener — who outranks whom, how formal the setting is, how direct it is acceptable to be. English does not carry that information anywhere in the sentence. So the machine has to infer something the source never stated, and it guesses.
A 2026 peer-reviewed study of business emails translated into Japanese found that a plain "translate this" instruction produced only limited adaptation toward Japanese conventions, failing specifically on honorifics, formal address, indirectness in hierarchical contexts, and the appropriate directness between a superior and a subordinate. The study's finding was that supplying context in the instruction fixed it — which is to say the fix is human expertise applied up front, not a better engine.
What it means for your workflow
A Japanese file that is wrong in this way is not wrong in a way a scoring model detects. It is grammatically correct, fluent, and rude. The reader who notices is your customer.
What we do
We review for register and honorific consistency against the audience and the relationship, not just against the source. Where you are producing Japanese at volume, we would rather fix the instruction than the output — so we report the pattern, not only the instance.