You may know this as AI training data work, data annotation, data validation, AI data quality, response rating and evaluation, preference ranking, RLHF, transcription adjudication, or content moderation.
What you get
Qualified native speakers judging model output against criteria you define, with the judgement recorded: ratings, rankings, or pass-fail against a rubric, delivered in your schema, with reviewer identity and qualifications attached to every judgement.
Three kinds of work
Why language is the hard part
Automatic metrics are unreliable in exactly the languages where you most need to know.
A 2026 study evaluating four frontier models on Hausa and Fongbe, ten thousand sentences each, found human ratings of 4.0 to 4.5 out of 5 for Hausa and 1.0 to 2.2 for Fongbe. For Fongbe, one standard metric ranked the worst-performing system best. Another returned near-identical scores for every text, because the model underneath could not tell two texts apart. For Hausa, every automatic metric picked one system while human judges preferred a different one. The researchers' own conclusion was that human evaluation is mandatory for these languages.
A separate 2025 audit of the 200-language benchmark most multilingual claims rest on found its own reference translations falling well below the quality standard claimed for them, and showed a model scoring 13.95 on the benchmark and 2.29 on realistic text while a better model scored 4.87 and 13.40. The benchmark ranked the worse model higher.
If your evaluation coverage is thinnest where your metrics are least trustworthy, that is not a measurement problem. It is a staffing problem.
Adversarial testing
We also take red teaming engagements in languages where English-first safety evaluation does not transfer — where the failures are cultural rather than linguistic, and a translated prompt list does not find them. Arabic makes the point: there is no single Arabic to test in, and a suite written in one variety tells you little about the others.
This is a different discipline from linguistic review and we scope it as one, in conversation rather than off a service page. We are not a general AI safety consultancy. What we bring is the part that is hardest to buy — native speakers who can tell you why a response that looks fine is not — and we brief reviewers on what an engagement involves before they accept it.
How we staff it
From a network of linguists we know personally, in named languages, whose subject expertise and qualifications we document for the work they take. Not an open crowd. For work of this kind the difference shows up as consistency between reviewers, which we measure and report rather than assert.
If you are an LSP buying this on behalf of a client
We work as your supplier under your name, in your schema, to your rubric and your throughput. Your client relationship is yours. We are named in your records and nowhere else unless you want us to be, and we do not approach your clients.
What it costs
Per task or per hour, depending on the shape of the work. Volume and turnaround are agreed per language, because our answer is different in Japanese than it is in Fongbe and we would rather tell you that up front.
Delivered, or staffed
Send us the work and we return the deliverable described above. Or place our reviewers into your own process and tooling and run them under your rubric, your schema and your brand. Same people, same documentation, different contract. Tell us which you want and we will price it that way.
Need the evidence as part of the deliverable? See documentation and evidence.