Three questions that do most of the work
Three questions do most of the work, and this page is all three of them, before anything has been explained.
There are eight questions in this volume and nobody remembers eight in a meeting. So here are the three to carry, before any of them has been explained and before you have any reason to trust the ordering.
- Compared to what? What does the same measurement say about doing this the old way, or the obvious cheap way, or not at all? A result with nothing beside it cannot be read in any direction.
- What is not in it? Which runs errored, timed out or came back empty, and were they counted as failures or dropped before the counting started? And which real situations have no case at all?
- Show me three failures. Not a share of failures. Three actual bad outputs, on screen. If nobody can produce them, nobody looked.
Why these three
Four things decided it, and they are printed here so you can disagree with the ordering rather than take it.
| the criterion | why it ranks |
|---|---|
| how often it is the fatal one | a question usually answered well is not worth one of three slots |
| what it costs to ask | a question answerable in the room beats a deeper one needing a week |
| whether the answer changes what you do | some answers change the decision; some only widen your doubt |
| whether it covers a distinct failure | three questions probing the same thing leave three other ways to be wrong |
The three cover different ground on purpose: can I read it, is it honest, did anyone look. Three questions all about the test cases would have left the grading and the decision entirely unexamined.
The most frequently violated question in this volume is not one of the three. “What is the denominator” — how many cases was this worked out over — is missing from almost every result that circulates. It did not make the list because the answer is a number that, on its own, rarely changes what you decide, and because the second question is the deeper version of it. That trade is the clearest case of the criteria doing real work, and if you think it is the wrong call, the criteria are above and the argument is yours to have.
The hardest cut was what result would have changed the recommendation. It is free to ask and it is the only question that catches a measurement that was really a case already made. It came fourth because it is confrontational, and many readers will not ask it — which makes it a worse thing to carry than a question they will actually use.