Reading an Eval ReportReference

The questions, on one page

For someone who has read the volume and wants it in front of them.

The three to carry

04
Why a result with nothing beside it cannot be read
What does the same measurement say about the old way, the cheap way, or doing nothing?
bad answer: A discussion rather than a number. If the gap has not been measured on the same cases, it does not exist yet.
07
The runs that left before the counting started
What happened to the runs that errored, timed out or came back empty?
bad answer: “It did not come up.” Also a suspiciously round count: real runs are ragged.
10
Why three bad outputs beat any share
Can I see three failures?
bad answer: “I will send some afterward” — nobody has looked. “They are mostly edge cases” — a defense formed before a test.

The other five, in the order the volume asks them

05
What a share hides when its count is missing
How many cases was this worked out over?
bad answer: A very large count that arrived very quickly. Expensive sets are slow to build, so a fast big one usually means the right answers were generated rather than determined.
06
How a true number can describe something that no longer exists
When was this run, and on what exactly?
bad answer: “Recently.” Or “nothing significant has changed” — a judgment about the thing in question, made by somebody with an interest in it.
08
When a high score is measuring memory
Who wrote these cases, and could the thing have seen them already?
bad answer: A denial. Ask who could find out and how long it would take them — that version has an answer, and “nobody, and we could not” is informative.
09
Why one grader's verdict cannot be checked
What decided each answer was right, and what checked that?
bad answer: “We used a model to grade it”, full stop. Now the most common answer and the least often checked.
11
Telling a real test from a case already made
What result would have led you to recommend against this?
bad answer: Silence. Or a reframing to say the decision was already made — sometimes legitimate, and a different document than the one you were handed.

Why those three and not the others

the criterionwhy it ranks
how often it is the fatal onea question usually answered well is not worth one of three slots
what it costs to aska question answerable in the room beats a deeper one that needs a week
whether the answer changes what you dosome answers change the decision; some only widen your doubt
whether it covers a distinct failurethree questions probing the same thing leave three other ways to be wrong

The most frequently violated question in the volume — how many cases was this worked out over — is not one of the three. Its answer is a number that on its own rarely changes what you do, and question seven is the deeper version of it.

The four parts of the machine

the partwhat can go wrong with it
the casesthe wrong ones, too few of them, or ones the thing had already seen
the runa version that no longer exists, and runs that never finished
the gradingnobody knows what ruled, and nothing ever checked it
the numbernothing, usually. This is the part that is almost always correct

Doubting the arithmetic is the wrong instinct. Almost every bad result is a correct sum over the wrong cases, or a correct sum whose losses happened earlier in the machine.

What none of this catches

A report can answer all eight questions and still measure the wrong thing. The eight test whether a measurement is trustworthy; they cannot test whether it is relevant, and a trustworthy measurement of a quantity nobody needed is the most expensive kind of report there is, because it survives scrutiny. Ask what the number is for before asking whether it is sound.

Every term the volume uses

denominator
The count a rate was worked out over. Ninety-four percent of what, and of how many.
baseline
What the same measurement says about doing it the old way, or the trivial way, or not at all.
held-out
Cases kept away from whoever built the thing, so passing them means something.
contamination
The thing being tested has already seen the test cases, so the score measures memory rather than ability.
grader
Whoever or whatever looked at each answer and decided it was right or wrong.
agreement
How often two graders looking at the same answers said the same thing.
abstention
A grader allowed to return a third verdict: I cannot tell.
stale
Measured on a version that no longer exists.
exhibit
A report made up for this volume, for you to judge. Its numbers are not claims about the world.

Taken as already known, and so not defined: eval, test, model, agent, metric. That list is a claim about who is reading, and it is printed so it can be argued with.

Sources, and their state

agreement methodology unverified
the standard treatment of measuring whether two graders agree
used here for: a grader with no measured agreement against another grader is uncalibrated
Campbell, 1979 unverified
Campbell's law, on the corruption of social indicators used for decisions
used here for: the more an indicator is used to decide, the more it is gamed
contamination studies, 2023 onward unverified
work measuring test cases appearing in training data
used here for: published scores fall, sometimes sharply, when contamination is controlled for
Goodhart, 1975 unverified
the original statement of the law now quoted as 'when a measure becomes a target'
used here for: a number that becomes a target stops measuring what it measured