Reading an Eval Report whether the number was honestly assembled › The runs that left before the counting started 12 steps, no labs
07 — part 3, whether the number was honestly assembledreasoning

The runs that left before the counting started

Ask what left before the counting started, because whatever fell out on the way does not appear anywhere in the answer.

This is the second of the three, it is the least glamorous question in the volume, and in my experience it is the one most often fatal. It is also the one nobody thinks to ask, because it is a question about an absence and absences do not appear on slides.

Runs fail. They error, they time out, they come back empty, they hit a rate limit, they crash on an input nobody expected. Every one of those is an outcome, and every one of them has to go somewhere. There are exactly two places it can go: counted as a failure, or removed before the counting started.

the same machine, with the floor gone under one part a situation you tested another one you tested one nobody thought of one that only happens at month end two of these never enter the machine at all. the cases what did get in the run some errored the counter climbing, and correct the total it reports the runs that fell out before they were counted no gauge reads this pile
The counter is climbing and it is not lying. It simply has no way to report the pile underneath it, because the number and the remainder are in different places and only one of them has a gauge.
What moves: two real situations never enter the machine at all, runs fall through a gap in the floor into a growing pile, and the counter climbs cleanly with no indication that anything was lost.

The effect on the share has no upper bound, which is what makes this different from the other seven questions. A small count merely widens your doubt. Dropped failures move the figure in one direction, by as much as you like, while every individual calculation along the way remains correct.

Two kinds of absence, one question

what is missinghow it happenswhat to ask
runs that did not finishan error is not a verdict, so the row has nothing to score and quietly leaveshow many runs did not produce an answer, and where are they in this number?
situations with no case at allnobody wrote a case for the thing that only happens at quarter endwhich of the situations we actually care about has no case here?

The tell is a suspiciously round count. Real runs produce ragged numbers. A set of cases that started at some round figure and reported a result over exactly that same round figure means nothing errored, nothing timed out, and nothing was excluded — which happens, but is worth one question.

The good answer is specific and slightly embarrassing: a handful errored, here is why, we scored them as failures. Someone who has looked will tell you this without being pushed. Someone who has not will tell you it did not come up.

Every term this page uses

denominator
The count a rate was worked out over. Ninety-four percent of what, and of how many.
baseline
What the same measurement says about doing it the old way, or the trivial way, or not at all.
held-out
Cases kept away from whoever built the thing, so passing them means something.
contamination
The thing being tested has already seen the test cases, so the score measures memory rather than ability.
grader
Whoever or whatever looked at each answer and decided it was right or wrong.
agreement
How often two graders looking at the same answers said the same thing.
abstention
A grader allowed to return a third verdict: I cannot tell.
stale
Measured on a version that no longer exists.
exhibit
A report made up for this volume, for you to judge. Its numbers are not claims about the world.

Taken as already known, and so not defined here: eval, test, model, agent, metric. That list is a claim about who is reading, and it is printed so it can be argued with.