Why one grader's verdict cannot be checked
Something looked at every single answer and ruled on it, and you want to know what that was and whether anybody ever checked its rulings.
Something looked at every answer and ruled on it. That is not a detail of the method; it is where the result came from. And it is missing from almost every summary that circulates, which means most people making decisions on these numbers do not know who or what produced them.
Naming it is half of this step. The question is not hard: what decided that each of these answers was right? There are only a few possible answers — a person, several people, a rule that could be written down, or another model — and which one it was changes what the number can support.
The other half is whether anything ever checked the grader. A grader nothing has been compared against is an instrument nobody calibratedagreement methodologyunverified, and the form of the check is the same whoever is doing the grading: have two of them rule on the same answers, and count how often they say the same thing.
| what graded it | the check that makes it usable | the bad answer |
|---|---|---|
| one person | a second person on a sample of the same answers, and the agreement counted | “our domain expert reviewed them” — one expert, no second opinion, and no way to know if a different expert would have agreed |
| a rule | the rule written out, so you can see what it counts as right | a rule nobody can state. If it cannot be written down it is not a rule |
| another model | its verdicts compared against a person's on a sample | “we used a model to grade it”, full stop. This is now the most common answer and the least often checked |
Ask what the grader does when it cannot tell. A grader forced to return right or wrong on an answer it has no basis to judge will return one of them anyway, and that verdict is noise entering your number as signal. A grader allowed to say I cannot tell gives you a smaller, honest result instead of a complete invented one.