Model output evaluation

Response rating, preference ranking and data validation — against your rubric, not a style guide.

--:--:-- [PST]
S01

What you get

Qualified native speakers judging model output against criteria you define, with the judgement recorded: ratings, rankings, or pass-fail against a rubric, delivered in your schema, with reviewer identity and qualifications attached to every judgement.

S02

Three kinds of work

01
Vanenberg
Response rating and evaluation
Scoring model output against your quality, safety or helpfulness criteria.
02
Vanenberg
Preference ranking
Ordering candidate responses best to worst — the human input that reinforcement learning from human feedback depends on.
03
Vanenberg
Data validation
Independent checking of annotation or generation work produced by someone else. We did not create the data we are checking, which is the entire point.
S03

Why language is the hard part

Automatic metrics are unreliable in exactly the languages where you most need to know.

A 2026 study evaluating four frontier models on Hausa and Fongbe, ten thousand sentences each, found human ratings of 4.0 to 4.5 out of 5 for Hausa and 1.0 to 2.2 for Fongbe. For Fongbe, one standard metric ranked the worst-performing system best. Another returned near-identical scores for every text, because the model underneath could not tell two texts apart. For Hausa, every automatic metric picked one system while human judges preferred a different one. The researchers' own conclusion was that human evaluation is mandatory for these languages.

A separate 2025 audit of the 200-language benchmark most multilingual claims rest on found its own reference translations falling well below the quality standard claimed for them, and showed a model scoring 13.95 on the benchmark and 2.29 on realistic text while a better model scored 4.87 and 13.40. The benchmark ranked the worse model higher.

If your evaluation coverage is thinnest where your metrics are least trustworthy, that is not a measurement problem. It is a staffing problem.

S04

Adversarial testing

We also take red teaming engagements in languages where English-first safety evaluation does not transfer — where the failures are cultural rather than linguistic, and a translated prompt list does not find them. Arabic makes the point: there is no single Arabic to test in, and a suite written in one variety tells you little about the others.

This is a different discipline from linguistic review and we scope it as one, in conversation rather than off a service page. We are not a general AI safety consultancy. What we bring is the part that is hardest to buy — native speakers who can tell you why a response that looks fine is not — and we brief reviewers on what an engagement involves before they accept it.

S05

How we staff it

From a network of linguists we know personally, whose subject expertise and qualifications are on record, in named languages. Not an open crowd. For work of this kind the difference shows up as consistency between reviewers, which we measure and report rather than assert.

S06

Reliability

We calibrate reviewers against a shared rubric before production work begins and report inter-rater agreement with the delivery. Where agreement is low, that is a finding about the rubric as often as about the reviewers, and we will say so.

S07

What it costs

Per task or per hour, depending on the shape of the work. Volume and turnaround are agreed per language, because our answer is different in Japanese than it is in Fongbe and we would rather tell you that up front.

S08

Tell us what you need evaluated, and in which languages

Send the rubric and the language list. We will tell you where we can staff it properly and where we cannot.

Talk to us about quality verification