Fast-growing markets, thin training data

Southeast Asia is where a great deal of commercial content is heading and where machine translation quality is least verifiable from anywhere else in your organisation.

--:--:-- [PST]
S01

Languages we cover

Vietnamese, Indonesian, Malay, Filipino and Tagalog, Khmer, Lao and Burmese. Thai has its own page.

S02

Why these languages are hard for machines

The data is thin, and the marketing hides it. Most of these are low-resource languages in the technical sense: comparatively little digitised parallel text, regardless of speaker numbers. A 2025 academic audit of the standard 200-language benchmark — the basis for most "200 languages supported" claims — found the benchmark's own reference translations fell below its stated quality bar, and that model performance on real-world text was dramatically worse than on the benchmark. The published numbers flatter the engines most in exactly this part of the world.

Vietnamese diacritics are meaning, not decoration. Tone and vowel marks distinguish words entirely. Systems that mishandle, strip or normalise them produce text that is not merely wrong but sometimes unintentionally offensive. Encoding problems in the pipeline compound this before the linguist ever sees it.

Indonesian and Malay are close, related, and not interchangeable. They share a great deal and differ in vocabulary, usage and official terminology, and they serve separate markets with separate regulators. Treating one as a variant of the other is a common and visible failure.

Filipino is code-mixed in practice. Everyday Filipino communication mixes English and Tagalog constantly. Content translated into formal, purely native-vocabulary Filipino reads as stilted to the audience it targets, while content that mixes too freely reads as unserious in official contexts. The correct register is a decision about audience.

Khmer, Lao and Burmese have no spaces between words, like Thai, with the same consequence: the system must segment before it can translate, and segmentation errors produce confident nonsense. Their scripts also have complex rendering and line-breaking behaviour that fails in interfaces and constrained layouts.

Register and formality are grammatically encoded in several of these languages, and English source text does not carry the information the target requires.

S03

What we do

  • We treat each language as its own job, with its own reviewer — Indonesian and Malay are never the same pass.
  • We check encoding and rendering explicitly: diacritics, script shaping, line breaking, and behaviour in constrained layouts.
  • We decide register and code-mixing policy per audience and record it so it holds across a programme.
  • We score segmentation-driven errors as a distinct category in the languages that need it.
  • We are candid about reviewer supply. In Khmer, Lao and Burmese, qualified specialist reviewers are genuinely scarce. We will tell you what we can staff properly rather than take work we cannot review.
S04

Where this work is worth most

Product and software content for the region's very large mobile-first consumer markets. Manufacturing and supply chain documentation, given how much production sits in Vietnam, Indonesia and Malaysia. Maritime crewing documentation — a substantial share of the world's seafarers are recruited from the Philippines and Indonesia, and their operating documentation is safety-critical. Public health and development content, where the audience has no alternative source.

S05

Talk to us about Southeast Asian languages

Send content that has already been machine-translated into one of these languages. Nobody in your organisation can check it, which is exactly why it is worth checking.

Talk to us about Southeast Asian languages