What the evidence says
In 2026 the main international machine translation conference added a dedicated task covering Chinese into and out of Thai, Vietnamese, Lao, Burmese, Khmer, Indonesian and Malay. The organisers' own description of the problem: these languages are constrained by limited parallel corpora, weak model generalization, and high barriers to edge deployment.
When the research community sets up a task to establish whether translation works in a set of languages, that is the state of the evidence. There is no equivalent of the Arabic or Japanese result here, in either direction, because the measurements have not been made yet.
Indonesian is the partial exception, and it is instructive: on real commercial content it scores 79.3 — comparable to Turkish or Polish — and the engine that wins it is not the engine that wins most other languages. Even where quality is reasonable, the right tool is different.
Why it happens
Two separate problems get grouped together. Indonesian and Malay are relatively well-served, closely related, and mutually confusable — engines routinely blend the two, which is invisible to a reviewer who does not read both. Khmer, Lao and Burmese are genuinely data-poor, use scripts with limited digital corpora, and — like Thai — do not mark word boundaries with spaces.
What it means for your workflow
Do not set one quality policy for "Southeast Asia." Indonesian may safely run at a high automation threshold. Burmese and Khmer, on current evidence, should not run on automatic scoring at all, because there is no published basis for trusting the score.
What we do
We tell you which of these languages your pipeline can treat like a European language and which it cannot, and we set a per-language threshold instead of a regional one. Where a language has no reliable metric, we substitute human sampling at a defined rate — and we write down the rate, so the policy is auditable.