Reading an Eval Report what a result is, and the three questions to carry › Three questions that do most of the work 12 steps, no labs
02 — part 1, what a result is, and the three questions to carryreasoning

Three questions that do most of the work

Three questions do most of the work, and this page is all three of them, before anything has been explained.

There are eight questions in this volume and nobody remembers eight in a meeting. So here are the three to carry, before any of them has been explained and before you have any reason to trust the ordering.

  1. Compared to what? What does the same measurement say about doing this the old way, or the obvious cheap way, or not at all? A result with nothing beside it cannot be read in any direction.
  2. What is not in it? Which runs errored, timed out or came back empty, and were they counted as failures or dropped before the counting started? And which real situations have no case at all?
  3. Show me three failures. Not a share of failures. Three actual bad outputs, on screen. If nobody can produce them, nobody looked.
many fair questions. Three you can carry every one of these is worth asking how often it is the fatal one what it costs to ask whether the answer changes what you do compared to what what is not in it show me three failures the rest are the deeper pass, not the discard pile one asks whether you can read it one whether it is honest one whether anyone looked
Eight fair questions, and the three that survive the criteria. The other five are the deeper pass, not the discard pile.
What moves: candidate questions flow through a sieve labeled with the criteria, and three emerge, each annotated with the distinct thing it tests.

Why these three

Four things decided it, and they are printed here so you can disagree with the ordering rather than take it.

the criterionwhy it ranks
how often it is the fatal onea question usually answered well is not worth one of three slots
what it costs to aska question answerable in the room beats a deeper one needing a week
whether the answer changes what you dosome answers change the decision; some only widen your doubt
whether it covers a distinct failurethree questions probing the same thing leave three other ways to be wrong

The three cover different ground on purpose: can I read it, is it honest, did anyone look. Three questions all about the test cases would have left the grading and the decision entirely unexamined.

The most frequently violated question in this volume is not one of the three. “What is the denominator” — how many cases was this worked out over — is missing from almost every result that circulates. It did not make the list because the answer is a number that, on its own, rarely changes what you decide, and because the second question is the deeper version of it. That trade is the clearest case of the criteria doing real work, and if you think it is the wrong call, the criteria are above and the argument is yours to have.

The hardest cut was what result would have changed the recommendation. It is free to ask and it is the only question that catches a measurement that was really a case already made. It came fourth because it is confrontational, and many readers will not ask it — which makes it a worse thing to carry than a question they will actually use.

Every term this page uses

denominator
The count a rate was worked out over. Ninety-four percent of what, and of how many.
baseline
What the same measurement says about doing it the old way, or the trivial way, or not at all.
held-out
Cases kept away from whoever built the thing, so passing them means something.
contamination
The thing being tested has already seen the test cases, so the score measures memory rather than ability.
grader
Whoever or whatever looked at each answer and decided it was right or wrong.
agreement
How often two graders looking at the same answers said the same thing.
abstention
A grader allowed to return a third verdict: I cannot tell.
stale
Measured on a version that no longer exists.
exhibit
A report made up for this volume, for you to judge. Its numbers are not claims about the world.

Taken as already known, and so not defined here: eval, test, model, agent, metric. That list is a claim about who is reading, and it is printed so it can be argued with.