The questions, on one page
For someone who has read the volume and wants it in front of them.
The three to carry
04
Why a result with nothing beside it cannot be read
What does the same measurement say about the old way, the cheap way, or doing nothing?
bad answer: A discussion rather than a number. If the gap has not been measured on the same cases, it does not exist yet.
07
The runs that left before the counting started
What happened to the runs that errored, timed out or came back empty?
bad answer: “It did not come up.” Also a suspiciously round count: real runs are ragged.
10
Why three bad outputs beat any share
Can I see three failures?
bad answer: “I will send some afterward” — nobody has looked. “They are mostly edge cases” — a defense formed before a test.
The other five, in the order the volume asks them
05
What a share hides when its count is missing
How many cases was this worked out over?
bad answer: A very large count that arrived very quickly. Expensive sets are slow to build, so a fast big one usually means the right answers were generated rather than determined.
06
How a true number can describe something that no longer exists
When was this run, and on what exactly?
bad answer: “Recently.” Or “nothing significant has changed” — a judgment about the thing in question, made by somebody with an interest in it.
08
When a high score is measuring memory
Who wrote these cases, and could the thing have seen them already?
bad answer: A denial. Ask who could find out and how long it would take them — that version has an answer, and “nobody, and we could not” is informative.
09
Why one grader's verdict cannot be checked
What decided each answer was right, and what checked that?
bad answer: “We used a model to grade it”, full stop. Now the most common answer and the least often checked.
11
Telling a real test from a case already made
What result would have led you to recommend against this?
bad answer: Silence. Or a reframing to say the decision was already made — sometimes legitimate, and a different document than the one you were handed.
Why those three and not the others
| the criterion | why it ranks |
|---|
| how often it is the fatal one | a question usually answered well is not worth one of three slots |
| what it costs to ask | a question answerable in the room beats a deeper one that needs a week |
| whether the answer changes what you do | some answers change the decision; some only widen your doubt |
| whether it covers a distinct failure | three questions probing the same thing leave three other ways to be wrong |
The most frequently violated question in the volume — how many cases was this worked out over — is not one of the three. Its answer is a number that on its own rarely changes what you do, and question seven is the deeper version of it.
The four parts of the machine
| the part | what can go wrong with it |
|---|
| the cases | the wrong ones, too few of them, or ones the thing had already seen |
| the run | a version that no longer exists, and runs that never finished |
| the grading | nobody knows what ruled, and nothing ever checked it |
| the number | nothing, usually. This is the part that is almost always correct |
Doubting the arithmetic is the wrong instinct. Almost every bad result is a correct sum over the wrong cases, or a correct sum whose losses happened earlier in the machine.
What none of this catches
A report can answer all eight questions and still measure the wrong thing. The eight test whether a measurement is trustworthy; they cannot test whether it is relevant, and a trustworthy measurement of a quantity nobody needed is the most expensive kind of report there is, because it survives scrutiny. Ask what the number is for before asking whether it is sound.
Every term the volume uses
- denominator
- The count a rate was worked out over. Ninety-four percent of what, and of how many.
- baseline
- What the same measurement says about doing it the old way, or the trivial way, or not at all.
- held-out
- Cases kept away from whoever built the thing, so passing them means something.
- contamination
- The thing being tested has already seen the test cases, so the score measures memory rather than ability.
- grader
- Whoever or whatever looked at each answer and decided it was right or wrong.
- agreement
- How often two graders looking at the same answers said the same thing.
- abstention
- A grader allowed to return a third verdict: I cannot tell.
- stale
- Measured on a version that no longer exists.
- exhibit
- A report made up for this volume, for you to judge. Its numbers are not claims about the world.
Taken as already known, and so not defined: eval, test, model, agent, metric. That list is a claim about who is reading, and it is printed so it can be argued with.
Sources, and their state
None of these has been opened. Every entry below was written down from memory in a session with no access to the works themselves. The volume's reasoning does not rest on any of them — every question is argued from first principles — but do not repeat a citation from this page until somebody has checked it.
agreement methodology unverified
the standard treatment of measuring whether two graders agree
used here for: a grader with no measured agreement against another grader is uncalibrated
Campbell, 1979 unverified
Campbell's law, on the corruption of social indicators used for decisions
used here for: the more an indicator is used to decide, the more it is gamed
contamination studies, 2023 onward unverified
work measuring test cases appearing in training data
used here for: published scores fall, sometimes sharply, when contamination is controlled for
Goodhart, 1975 unverified
the original statement of the law now quoted as 'when a measure becomes a target'
used here for: a number that becomes a target stops measuring what it measured