Reading an Eval Report whether it supports the decision attached to it › Why three bad outputs beat any share 12 steps, no labs
10 — part 4, whether it supports the decision attached to itreasoning

Why three bad outputs beat any share

Ask to see three things that went wrong, because a report with no examples of failure is a report nobody actually read.

The last of the three, and the one that will change your decision most often. It is also the only question on the list whose value does not depend on getting an answer: a refusal tells you as much as a reply.

A share tells you how much went wrong. It cannot tell you what went wrong, and what went wrong is what you need, because it decides whether this is a small fix, a design problem, or a reason not to do this at all. Three concrete bad outputs settle questions no figure can reach.

the share grows and tells you nothing new the number, climbing one fact, repeated it invented a source and cited it precisely a different cause it dropped the last row every single time a different cause it answered a different question, very well a different cause three specimens, three separate problems, three different fixes none of which the share could have told you
The bar is the most prominent thing in the frame and the least informative. Three specimens resolve into three separate problems with three different fixes.
What moves: the share climbs steadily while three cards arrive one at a time, each naming a distinct failure with a distinct cause.

Look at what the drawing separates. The same measurement produced both sides of it. The bar was on the slide; the cards were available to anyone who opened the results.

Why a refusal is an answer

Failures are the easiest thing in the world to produce if anyone has looked at the output, because they are sitting right there in the results, and they are the first thing a curious person reads. So there are only a few reasons you cannot have three.

what you hearwhat it means
“I can send you some after the meeting”nobody has looked, and the report was assembled from a summary
“they are mostly edge cases”somebody has looked, has already formed a defense, and has not tested it. Ask for the three anyway — edge cases are a claim about frequency, and you have the question for that
“here, and this one is interesting because…”this person has read their own results. You can trust the rest of the report much further than you could a minute ago

The shape of the failures matters more than their number. Three failures that are all the same failure is good news badly presented: it is one bug. Three failures with nothing in common is much worse than the same share spread across one cause, and no summary figure distinguishes those two situations.

Every term this page uses

denominator
The count a rate was worked out over. Ninety-four percent of what, and of how many.
baseline
What the same measurement says about doing it the old way, or the trivial way, or not at all.
held-out
Cases kept away from whoever built the thing, so passing them means something.
contamination
The thing being tested has already seen the test cases, so the score measures memory rather than ability.
grader
Whoever or whatever looked at each answer and decided it was right or wrong.
agreement
How often two graders looking at the same answers said the same thing.
abstention
A grader allowed to return a third verdict: I cannot tell.
stale
Measured on a version that no longer exists.
exhibit
A report made up for this volume, for you to judge. Its numbers are not claims about the world.

Taken as already known, and so not defined here: eval, test, model, agent, metric. That list is a claim about who is reading, and it is printed so it can be argued with.