Model output evaluation

Response rating, preference ranking and data validation — against your rubric, not a style guide.

--:--:-- [PST]
S01

You may know this as AI training data work, data annotation, data validation, AI data quality, response rating and evaluation, preference ranking, RLHF, transcription adjudication, or content moderation.

S02

What you get

Qualified native speakers judging model output against criteria you define, with the judgement recorded: ratings, rankings, or pass-fail against a rubric, delivered in your schema, with reviewer identity and qualifications attached to every judgement.

S03

Three kinds of work

01
Vanenberg
Response rating and evaluation
Scoring model output against your quality, safety or helpfulness criteria.
02
Vanenberg
Preference ranking
Ordering candidate responses best to worst — the human input that reinforcement learning from human feedback depends on.
03
Vanenberg
Data validation
Independent checking of annotation or generation work produced by someone else. We did not create the data we are checking, which is the entire point.
S04

Why language is the hard part

Automatic metrics are unreliable in exactly the languages where you most need to know.

A 2026 study evaluating four frontier models on Hausa and Fongbe, ten thousand sentences each, found human ratings of 4.0 to 4.5 out of 5 for Hausa and 1.0 to 2.2 for Fongbe. For Fongbe, one standard metric ranked the worst-performing system best. Another returned near-identical scores for every text, because the model underneath could not tell two texts apart. For Hausa, every automatic metric picked one system while human judges preferred a different one. The researchers' own conclusion was that human evaluation is mandatory for these languages.

A separate 2025 audit of the 200-language benchmark most multilingual claims rest on found its own reference translations falling well below the quality standard claimed for them, and showed a model scoring 13.95 on the benchmark and 2.29 on realistic text while a better model scored 4.87 and 13.40. The benchmark ranked the worse model higher.

If your evaluation coverage is thinnest where your metrics are least trustworthy, that is not a measurement problem. It is a staffing problem.

S05

Adversarial testing

We also take red teaming engagements in languages where English-first safety evaluation does not transfer — where the failures are cultural rather than linguistic, and a translated prompt list does not find them. Arabic makes the point: there is no single Arabic to test in, and a suite written in one variety tells you little about the others.

This is a different discipline from linguistic review and we scope it as one, in conversation rather than off a service page. We are not a general AI safety consultancy. What we bring is the part that is hardest to buy — native speakers who can tell you why a response that looks fine is not — and we brief reviewers on what an engagement involves before they accept it.

S06

How we staff it

From a network of linguists we know personally, in named languages, whose subject expertise and qualifications we document for the work they take. Not an open crowd. For work of this kind the difference shows up as consistency between reviewers, which we measure and report rather than assert.

S07

If you are an LSP buying this on behalf of a client

We work as your supplier under your name, in your schema, to your rubric and your throughput. Your client relationship is yours. We are named in your records and nowhere else unless you want us to be, and we do not approach your clients.

S08

What it costs

Per task or per hour, depending on the shape of the work. Volume and turnaround are agreed per language, because our answer is different in Japanese than it is in Fongbe and we would rather tell you that up front.

S09

Delivered, or staffed

Send us the work and we return the deliverable described above. Or place our reviewers into your own process and tooling and run them under your rubric, your schema and your brand. Same people, same documentation, different contract. Tell us which you want and we will price it that way.

Need the evidence as part of the deliverable? See documentation and evidence.

S10

Tell us what you need evaluated, and in which languages

Send the rubric and the language list. We will tell you where we can staff it properly and where we cannot.

Talk to us about quality verification