The report that answers everything and still misleads
Here are four results and one of them holds up, and finding out which one teaches you what these questions cannot catch.
Four results. One of them holds up. Mark each one before reading on, and write down the single question you would ask about it — the questions are on every page of this volume and you are meant to look back at them.
Note the phrase “all completed”.
The marking
One, two and three each fail on something structural. Four does not.
| report | what it fails | the one question to ask |
|---|---|---|
| One | everything. It is a headline with no result behind it | compared to what? |
| Two | the runs that did not finish, and an unchecked grader | what happened to the questions it could not answer? |
| Three | where the cases came from — a public collection the system may well have trained on | could it have seen these issues already? |
| Four | nothing on this list | — |
Report three is the one most people pass, because it is visibly careful: it excludes runs and says so, it has a reference point, it is dated, and its grading is sensible. Being careful in seven ways does not help if the cases were memorized.
What the eight questions do not catch
Report four answers all eight. It is honest, well-built and re-runnable, and the people who made it did everything this volume asks. It is also measuring the wrong thing.
It measures whether the drafted reply is factually correct. Nobody in the business cares whether replies are factually correct. They care whether the customer's problem got resolved, and a reply can be entirely accurate and resolve nothing — correct, polite, on topic, and leaving the person exactly where they started.
No structural question can catch this, and a ninth would not help. Everything about report four is in order. The defect is a gap between what was measured and what anybody wanted, and closing it needs someone who knows the business well enough to notice — which is you, and is the part nobody can hand off to a checklist.
So the eight questions have one prerequisite
- Ask what this number is for before asking whether it is sound. The eight questions test whether a measurement is trustworthy. They cannot test whether it is relevant, and a trustworthy measurement of the wrong quantity is the most expensive kind of report there is, because it survives scrutiny.
- They are still worth having. Three of the four reports here fail on structure, and that is the usual proportion. Most bad results are bad in ways the questions catch.
- They are no defense against someone who knows them. A report built to pass this list will pass it. At that point you are back to the last question in part four: what would have changed the recommendation.
Now go back to the first page of this volume. The card there has not changed.