Reading an Eval Report whether the number was honestly assembled › When a high score is measuring memory 12 steps, no labs
08 — part 3, whether the number was honestly assembledsourced

When a high score is measuring memory

Ask who wrote the cases and whether the thing being tested had already seen them, because both answers can make the score meaningless.

Two different things can be wrong with where the cases came from, they have different tells, and only one of them is anybody's fault.

who wrote these cases, and had the thing already seen them? written by the builders they know what it can do written by somebody else kept away from them the cases the score is over what the thing learned from before you ever tested it if the circle closes, the score measures memory the answerable version: who could tell me, and how long would it take them?
Two authors for the same cases, and the loop that makes a score meaningless. If the circle closes, the number measures memory rather than ability.
What moves: cases arrive from two sources, connect to what the thing learned from, and a dashed loop runs back from there into the cases, closing a circle.

The cases were written by the people who built the thing

Not dishonesty. It is close to unavoidable, because the people who built it are the people who know what it is supposed to do, and they are therefore also the people who know, without deciding to, which situations it handles. A set of cases written by a builder is a set of cases shaped like the thing that was built.

This is what held-out means and why anybody bothers: cases kept away from whoever built the thing, so that passing them says something. You are not asking for purity. You are asking whether anyone outside the team contributed cases, and if the answer is nobody, you now know what the number means and can price it accordingly.

The thing had already seen the cases

The second problem, and the more serious one, because it does not degrade a score — it invalidates it. If the cases were already part of what the system learned from, then answering them correctly demonstrates memory and nothing else. Published scores have been found to fall, sometimes sharply, once this is controlled forcontamination studies, 2023 onwardunverified.

Do not ask “is it contaminated”. It invites a denial, and the person answering usually does not know. Ask instead: who could tell me whether these cases were in the training data, and how long would it take them? That version has an answer, and the answer is informative even when it is “nobody, and we could not find out”.

Notice that both problems point at the same part of the machine, the first of the four, and that neither is visible anywhere in the number. A result built on cases the system had memorized looks exactly like a result built on cases it handled.

Every term this page uses

denominator
The count a rate was worked out over. Ninety-four percent of what, and of how many.
baseline
What the same measurement says about doing it the old way, or the trivial way, or not at all.
held-out
Cases kept away from whoever built the thing, so passing them means something.
contamination
The thing being tested has already seen the test cases, so the score measures memory rather than ability.
grader
Whoever or whatever looked at each answer and decided it was right or wrong.
agreement
How often two graders looking at the same answers said the same thing.
abstention
A grader allowed to return a third verdict: I cannot tell.
stale
Measured on a version that no longer exists.
exhibit
A report made up for this volume, for you to judge. Its numbers are not claims about the world.

Taken as already known, and so not defined here: eval, test, model, agent, metric. That list is a claim about who is reading, and it is printed so it can be argued with.