Reading an Eval Report whether the number was honestly assembled › Why one grader's verdict cannot be checked 12 steps, no labs
09 — part 3, whether the number was honestly assembledsourced

Why one grader's verdict cannot be checked

Something looked at every single answer and ruled on it, and you want to know what that was and whether anybody ever checked its rulings.

Something looked at every answer and ruled on it. That is not a detail of the method; it is where the result came from. And it is missing from almost every summary that circulates, which means most people making decisions on these numbers do not know who or what produced them.

Naming it is half of this step. The question is not hard: what decided that each of these answers was right? There are only a few possible answers — a person, several people, a rule that could be written down, or another model — and which one it was changes what the number can support.

something ruled on every answer. Did anything check it? the same answers one grader ruled on all of them a second grader ruled on them too they agreed they did not how often they said the same thing this is the check one opinion, held with confidence and nothing to hold it against no rate exists here. not a low one. None. first half: two graders, and the check they make possible. Second half: the second grader is removed. the thing that vanishes is not the quality of the grading. It is your ability to say anything about it.
An agreement is not a property of a grader. It is a property of a pair, so removing one end does not make the check worse — it removes the check.
What moves: answers flow to two graders whose verdicts accumulate in an agreed bin and a disagreed bin. Halfway through the cycle the second grader is removed and the check disappears entirely rather than degrading.

The other half is whether anything ever checked the grader. A grader nothing has been compared against is an instrument nobody calibratedagreement methodologyunverified, and the form of the check is the same whoever is doing the grading: have two of them rule on the same answers, and count how often they say the same thing.

what graded itthe check that makes it usablethe bad answer
one persona second person on a sample of the same answers, and the agreement counted“our domain expert reviewed them” — one expert, no second opinion, and no way to know if a different expert would have agreed
a rulethe rule written out, so you can see what it counts as righta rule nobody can state. If it cannot be written down it is not a rule
another modelits verdicts compared against a person's on a sample“we used a model to grade it”, full stop. This is now the most common answer and the least often checked

Ask what the grader does when it cannot tell. A grader forced to return right or wrong on an answer it has no basis to judge will return one of them anyway, and that verdict is noise entering your number as signal. A grader allowed to say I cannot tell gives you a smaller, honest result instead of a complete invented one.

→ For one result you actually rely on, find out what graded it. If nobody can tell you within a day, you have your answer — and you now know something about every decision that number has been used for.

Every term this page uses

denominator
The count a rate was worked out over. Ninety-four percent of what, and of how many.
baseline
What the same measurement says about doing it the old way, or the trivial way, or not at all.
held-out
Cases kept away from whoever built the thing, so passing them means something.
contamination
The thing being tested has already seen the test cases, so the score measures memory rather than ability.
grader
Whoever or whatever looked at each answer and decided it was right or wrong.
agreement
How often two graders looking at the same answers said the same thing.
abstention
A grader allowed to return a third verdict: I cannot tell.
stale
Measured on a version that no longer exists.
exhibit
A report made up for this volume, for you to judge. Its numbers are not claims about the world.

Taken as already known, and so not defined here: eval, test, model, agent, metric. That list is a claim about who is reading, and it is printed so it can be argued with.