Arabic is where machine translation still loses

In the 2025 WMT evaluation, not one of sixty machine systems matched a professional human translator into Egyptian Arabic. We work in the gap that leaves.

--:--:-- [PST]
0 of 60
machine systems that beat a professional human translator into Egyptian Arabic
WMT25
78.5 vs 77.0
human translator score against the best machine system (GPT-4.1)
WMT25
1.0 / 5
dialect authenticity rating for the system with the better automated score
Arabic dialect evaluation study
S01

The evidence

The Conference on Machine Translation runs the closest thing this field has to an independent, refereed contest. Its 2025 edition was deliberately made harder — the organisers titled their paper Time to Stop Evaluating on Easy Test Sets — and covered 30 language pairs and 60 systems, with professional annotators marking real errors rather than giving impressionistic scores.

English into Egyptian Arabic was named, jointly with English into Italian, the most challenging pair in the evaluation. The professional human translator scored 78.5. The best machine system, GPT-4.1, scored 77.0. No system tested — not the frontier language models, not the dedicated translation engines — matched the human.

That is the strongest single piece of evidence anywhere in this field for keeping a qualified person in the Arabic workflow.

S02

Why Arabic defeats engines

Four problems compound, and none of them is solved by a larger model.

There are effectively two Arabics, and the machine picks the wrong one. Formal written Arabic — Modern Standard Arabic — is the language of news, government and documents. Spoken Arabic varies by region: Egyptian, Gulf, Levantine, Maghrebi, with differences large enough to impede comprehension between them. "I want" is 'urīdu in the formal standard, 'āyez in Egyptian, biddī in Levantine. A translation "into Arabic" is not one deliverable. It is a choice of register and region, and a machine cannot make that choice unless someone tells it — so it defaults.

Automated quality scores actively mislead here. In one study, the system with the better automated score rated 1.0 out of 5 for dialect authenticity, against 4.80 for the system that scored lower on the metric. If your quality gate is a BLEU or COMET threshold, Arabic is where it will quietly pass content your audience will reject.

The dialects have no standardised spelling, so the training data behind them is inconsistent in a way European-language data is not.

Short vowels are not written. The same written form can be several different words, so the machine must disambiguate from context far more often than in a fully explicit script — and it gets it wrong in exactly the places where precision matters.

Then there is everything outside the engine entirely: right-to-left layout, interface mirroring, and documents containing Latin brand names, numerals and units that break the flow of the line. That is production cost which does not shrink as engines improve.

S03

What we do

  • We establish which Arabic before we start. Region, register, audience. If nobody has decided, we will tell you what the decision costs either way rather than quietly defaulting to MSA.
  • We check dialect authenticity as a named error category, separately from accuracy — because a file can be accurate and still read as foreign.
  • We review layout as well as text. Bidirectional text, embedded Latin strings, numerals, punctuation direction and anything the engine's output does to a page.
  • We use linguists matched to the domain, not generalists with a glossary. For legal and regulatory Arabic that distinction is the whole job.
S04

Where Arabic work is worth most

Arabic crossed with a high-consequence domain is the strongest position in this language. Legal and government content, medical device documentation, pharmaceutical material and financial reporting — where language difficulty raises the error rate and regulation raises the cost of an error.

Published 2026 rate data puts English↔Arabic among the highest-priced major pairs, at roughly one and a half to two times English↔Spanish. The reason cited is thin translator supply. It is a supply constraint, not a hype cycle, and it is not resolving quickly.

S05

Send us an Arabic file

Raw engine output or a finished translation from another supplier. We will return it reviewed, with the errors categorised and dialect handling assessed separately.

Send us an Arabic file