African languages

Where the quality score is not just wrong — it is inverted.

--:--:-- [PST]
Worst ranked best
the COMET metric ranked the worst-performing system highest for Fongbe
Hausa/Fongbe study, 2026
4.0–4.5 vs 1.0–2.2
human ratings out of 5 for Hausa against Fongbe — two African languages, threefold gap
Hausa/Fongbe study, 2026
Below 10 / 100
every system, and the human translation, into Maasai
WMT25
S01

What the evidence says

This is the clearest documented failure of automated quality measurement anywhere in translation, and it is worth setting out precisely.

A 2026 study evaluated four frontier models on Hausa and Fongbe, ten thousand sentences per language, with both automatic metrics and native-speaker judgement.

Hausa came out usable: human ratings of 4.0 to 4.5 out of 5. Fongbe did not: 1.0 to 2.2, with the best system at 2.20 out of 5. Two African languages, a threefold gap in quality — which is itself the first lesson, because "African languages" is not a category that predicts anything.

Then the scoring. For Fongbe, the COMET metric ranked the worst-performing system as the best. BERTScore returned within-language similarity above 0.99 for both languages — meaning the underlying model could not distinguish one text from another and returned high scores regardless of quality. For Hausa, every automatic metric ranked one system first while human evaluators preferred a different one. The authors' recommendation is that human evaluation is mandatory for these languages and that neural metrics should not be trusted until validated on the specific pair.

In the 2025 international evaluation, English to Maasai was scored with a fallback metric because the standard ones were judged unreliable for it. Every system, and the human translation, scored below 10 out of 100.

S02

What it means for your workflow

If your pipeline routes African-language content by automatic score, it is not making a quality decision. On the published evidence it may be making the opposite of one. Content that passes may be unusable; content that is flagged may be fine.

S03

What we do

For these languages we do not review the flagged segments — we sample the whole file, because the flags carry no information. We tell you which of your African languages behave like Hausa and which behave like Fongbe, because the answer determines whether machine translation is viable there at all. And where it is not viable, we say so, rather than selling you post-editing on output that cannot be edited into shape.

S04

Ask us whether machine translation works at all in your African languages

It is a real question with a real answer, and the answer differs sharply between languages that get grouped together.

Talk to us about quality verification