What the evidence says
In the 2025 evaluation, English to Egyptian Arabic was named one of the two hardest pairs tested. The professional human translation scored 78.5. The best of sixty machine systems, including the leading frontier models, reached 77.0. Not one beat the human.
On real commercial content, a separate index of over five thousand evaluations puts Arabic mid-table at 78.6 — better than Japanese, Russian or Chinese. So this is not a language where machines produce gibberish. It is a language where they produce something fluent that is aimed at the wrong reader.
Why it happens
Four problems compound.
- There is no single Arabic to translate into. Modern Standard Arabic is the written register of news, government and formal documents. Spoken Arabic varies by region — Egyptian, Gulf, Levantine, Maghrebi — enough to impede comprehension between them. "I want" is 'urīdu in the standard, 'āyez in Egyptian, biddī in Levantine. A brief that says "translate into Arabic" has not specified the deliverable, and no engine can make that choice for you.
- The dialects have no standard spelling, so the training data is inconsistent in a way that European-language data is not.
- Machines default to the formal standard when uncertain. That scores well on automatic metrics and is wrong for the audience. In one study the system with the better automatic score rated 1.0 out of 5 for dialect authenticity, against 4.80 for a lower-scoring, dialect-aware system. Your quality score can move in the opposite direction to your quality.
- Short vowels are not written, so one written form can be several words and has to be disambiguated from context far more often than in a fully-spelled language. Add right-to-left layout, which affects page design, interface mirroring and every document containing Latin brand names or numerals — a production cost that sits outside the engine entirely and does not shrink as engines improve.
What it means for your workflow
Automatic scoring is actively unsafe here, because register error is invisible to it. A file can pass every threshold you have set and still be addressed to the wrong country.
What we do
We check register and regional variety before we check anything else, against the audience you name in the brief. We flag standard-Arabic drift in content meant for a specific market. We check right-to-left layout in the rendered file, not the string. And we tell you which of those failures your engine is producing systematically, so it can be fixed at the prompt or the glossary rather than one file at a time.