What the evidence says
English to Russian produced the largest machine-over-human margin of any pair in the 2025 evaluation: the best system scored 83.4 against a professional human translation's 70.5. That is not a narrow win.
On real commercial content, Russian scores 73.0 — second lowest of the eighteen languages measured, below German, Korean, Arabic and Italian.
Russian is the clearest case in this whole set of a language where the benchmark answer and the production answer point in opposite directions.
Why it happens
Russian is not a low-resource language and it is not culturally opaque to translation systems. Its difficulty in production is grammatical density: six cases, aspect pairs on almost every verb, gender agreement running across the sentence, and free word order used to carry emphasis. Each of those is a place where a small error is grammatical, fluent and wrong — and where an error made in segment forty contradicts a decision made in segment four.
That is a document-level failure mode, and general-document benchmarks are not built to find it.
What it means for your workflow
Russian will pass your quality gate. It scores well on the tests, and modern engines produce confident, fluent Russian. What it will not do reliably is stay consistent across a long technical document or a large string set — and the failure will surface with your reader, not with your scoring model.
What we do
We review Russian for consistency across the full deliverable — terminology, aspect and register — rather than segment by segment. Where you have concluded from benchmark scores that Russian is a safe language to automate, we will test that on your own content before you commit to it.