What the evidence says
This is the clearest documented failure of automated quality measurement anywhere in translation, and it is worth setting out precisely.
A 2026 study evaluated four frontier models on Hausa and Fongbe, ten thousand sentences per language, with both automatic metrics and native-speaker judgement.
Hausa came out usable: human ratings of 4.0 to 4.5 out of 5. Fongbe did not: 1.0 to 2.2, with the best system at 2.20 out of 5. Two African languages, a threefold gap in quality — which is itself the first lesson, because "African languages" is not a category that predicts anything.
Then the scoring. For Fongbe, the COMET metric ranked the worst-performing system as the best. BERTScore returned within-language similarity above 0.99 for both languages — meaning the underlying model could not distinguish one text from another and returned high scores regardless of quality. For Hausa, every automatic metric ranked one system first while human evaluators preferred a different one. The authors' recommendation is that human evaluation is mandatory for these languages and that neural metrics should not be trusted until validated on the specific pair.
In the 2025 international evaluation, English to Maasai was scored with a fallback metric because the standard ones were judged unreliable for it. Every system, and the human translation, scored below 10 out of 100.
What it means for your workflow
If your pipeline routes African-language content by automatic score, it is not making a quality decision. On the published evidence it may be making the opposite of one. Content that passes may be unusable; content that is flagged may be fine.
What we do
For these languages we do not review the flagged segments — we sample the whole file, because the flags carry no information. We tell you which of your African languages behave like Hausa and which behave like Fongbe, because the answer determines whether machine translation is viable there at all. And where it is not viable, we say so, rather than selling you post-editing on output that cannot be edited into shape.