What the evidence says
Italian is a well-resourced European language written in the Latin alphabet, with a large translation industry and centuries of parallel text. Every assumption says it should be easy.
In the 2025 evaluation, English to Italian was named — alongside Egyptian Arabic — one of the two most challenging pairs tested. The top system scored around 79. In the same evaluation, English to Serbian scored 94.2, English to Ukrainian 90.3 and English to Czech 88.7.
Italian is the lowest-scoring well-resourced European pair in the hardest independent test available, and it is not close.
Why it happens
Not data scarcity — Italian has plenty. The difficulty is that the acceptable range is narrower. Italian readers and Italian clients are demanding about register, rhythm and formality in a way that shows up immediately as "translated" text. The gap between a sentence that is correct and a sentence that is right is wider here than the training data suggests, and the errors are stylistic rather than semantic — which is exactly the category automatic metrics measure worst.
Note the contrast with real commercial content, where Italian scores 81.0, near the top of the table. Italian is not hard in the sense that Fongbe is hard. It is hard in the sense that it is easy to be almost right.
What it means for your workflow
If your quality model was calibrated on French, Spanish or German, it is miscalibrated for Italian, because it is looking for a class of error that is not the one occurring. Italian output tends to clear thresholds and then fail in review.
What we do
We review Italian for register and naturalness against the intended reader, not against the source sentence. Where you translate Italian at volume, we would recommend a separate quality threshold for it — and we will show you, on your content, why the European default is set wrong.