Southeast Asian languages

Vietnamese, Indonesian, Malay, Khmer, Lao, Burmese — the fastest-growing markets with the thinnest evidence.

--:--:-- [PST]
New in 2026
a dedicated task for Chinese with Thai, Vietnamese, Lao, Burmese, Khmer, Indonesian and Malay
WMT26
79.3
Indonesian on real client content — and won by a different engine from most languages
Engine index, 2026
No basis
for trusting an automatic score in Khmer, Lao or Burmese on current published evidence
S01

What the evidence says

In 2026 the main international machine translation conference added a dedicated task covering Chinese into and out of Thai, Vietnamese, Lao, Burmese, Khmer, Indonesian and Malay. The organisers' own description of the problem: these languages are constrained by limited parallel corpora, weak model generalization, and high barriers to edge deployment.

When the research community sets up a task to establish whether translation works in a set of languages, that is the state of the evidence. There is no equivalent of the Arabic or Japanese result here, in either direction, because the measurements have not been made yet.

Indonesian is the partial exception, and it is instructive: on real commercial content it scores 79.3 — comparable to Turkish or Polish — and the engine that wins it is not the engine that wins most other languages. Even where quality is reasonable, the right tool is different.

S02

Why it happens

Two separate problems get grouped together. Indonesian and Malay are relatively well-served, closely related, and mutually confusable — engines routinely blend the two, which is invisible to a reviewer who does not read both. Khmer, Lao and Burmese are genuinely data-poor, use scripts with limited digital corpora, and — like Thai — do not mark word boundaries with spaces.

S03

What it means for your workflow

Do not set one quality policy for "Southeast Asia." Indonesian may safely run at a high automation threshold. Burmese and Khmer, on current evidence, should not run on automatic scoring at all, because there is no published basis for trusting the score.

S04

What we do

We tell you which of these languages your pipeline can treat like a European language and which it cannot, and we set a per-language threshold instead of a regional one. Where a language has no reliable metric, we substitute human sampling at a defined rate — and we write down the rate, so the policy is auditable.

S05

Ask us to set per-language thresholds for your Southeast Asian content

One regional policy is the wrong shape for this group. The languages behave nothing like each other.

Talk to us about quality verification