344 million native speakers, and still not enough data
Hindi is a low-resource language for machine translation. So are Bengali, Tamil, Telugu, Marathi and almost every other language of the subcontinent. Low-resource is a statement about digitised parallel text, not about how many people speak it — and buyers get this wrong constantly.
Languages we cover
Hindi, Bengali, Tamil, Telugu, Marathi, Gujarati, Kannada, Malayalam, Punjabi, Odia, Assamese and Urdu, alongside the other languages of the region on request.
The misunderstanding this page exists to correct
"Low-resource" sounds like it means "minor". It does not. It means there is not much digitised parallel text for a system to learn from.
Hindi has 344 million native speakers and is low-resource. Icelandic has fewer than 400,000 speakers and is comparatively well-resourced, because Iceland digitised its language systematically. The category tracks data infrastructure, not demographics.
This matters commercially because buyers assume that a language spoken by hundreds of millions of people must be well served by engines that advertise 200-language coverage. It is not, and the marketing actively obscures it.
The engine coverage claims are worse than they look
A 2025 academic audit re-examined the standard 200-language benchmark that most "we support 200 languages" marketing rests on. It found the benchmark's own reference translations fell below its claimed quality standard — the supposedly correct answers were not correct enough.
The audit also compared benchmark performance against real-world text. In one language, a model scored 13.95 on the benchmark and 2.29 on real-world content. The published gap between languages is narrower than the actual gap, in the direction that flatters the vendor.
If your language coverage decision rests on a supplier's claimed language count, that number is doing much less work than you think.
What actually goes wrong
Script and rendering. Devanagari, Bengali, Tamil, Telugu, Gujarati, Gurmukhi and Malayalam each have their own conjunct forms, matra positioning and shaping behaviour. Text that is linguistically correct renders wrongly in the wrong font, breaks incorrectly in constrained layouts, and is silently mangled by systems that were tested in Latin script.
Register and honorifics. Most of these languages grammatically encode respect and social distance — Hindi's tu, tum and aap being the familiar example. English source text carries none of that information, so the engine invents it.
Code-mixing is the actual register. Urban Indian communication mixes English and the local language constantly, and often writes the local language in Latin script. Content translated into formal, purely native-vocabulary Hindi frequently reads as stilted or bureaucratic to the audience it targets. The right answer depends on the audience, and it is a decision, not a default.
Word order and case. These languages are verb-final with case marking, structurally distant from English. Engines handle short sentences well and long, qualified, conditional sentences badly.
Terminology gaps in technical and medical content. Where standard technical vocabulary does not exist or competes with an English loanword, the correct choice is a judgement about the audience. Engines pick whichever is more frequent in the training data.
What we do
- We decide script, register and code-mixing policy explicitly per audience, and record the decision so it stays consistent across a programme.
- We check rendering as well as text — conjuncts, matras, line breaking, and behaviour in constrained layouts.
- We score honorific and register errors separately from accuracy.
- We use reviewers matched to language and domain. Bengali and Tamil are not interchangeable, and a reviewer who covers "Indian languages" generically covers none of them properly.
- We are honest about supply. Qualified specialist reviewers in some of these languages are genuinely scarce. We will tell you what we can staff and what we cannot rather than accepting work we cannot review properly.
Where this work is worth most
Product and software content, where rendering and constrained layouts produce failures the engine cannot see. Public sector and citizen-facing information, where the audience frequently has no alternative source. Pharmaceutical and clinical content for India's large trial and manufacturing base. Financial and telecommunications content aimed at mass-market consumers.