Reading an Eval Report whether it supports the decision attached to it › The report that answers everything and still misleads 12 steps, no labs
12 — part 4, whether it supports the decision attached to itexhibit

The report that answers everything and still misleads

Here are four results and one of them holds up, and finding out which one teaches you what these questions cannot catch.

Four results. One of them holds up. Mark each one before reading on, and write down the single question you would ask about it — the questions are on every page of this volume and you are meant to look back at them.

One — classification qualityconstructed for this volume
headline96% correct
casesnot stated
compared againstnot stated
date and versionnot stated
graded bythe team, by inspection
Two — document question answeringconstructed for this volume
headline89% correct
cases2,000 questions, all completed
compared againstthe previous build, on the same questions
date and versionlast Tuesday, build named in the appendix
graded bya model, prompted to return correct or incorrect

Note the phrase “all completed”.

Three — code repairconstructed for this volume
headline71% of issues resolved
cases340 issues, 22 excluded as unrunnable
compared againstan unassisted engineer, sampled, on the same issues
date and versiondated, versioned, re-runnable
graded bythe project's own tests, plus a reviewer on disagreements
cases drawn froma well-known public collection of issues
Four — support reply draftingconstructed for this volume
headline82% factually correct
cases614 conversations, 9 errored and scored as failures
compared againstthe current human-drafted replies, on the same conversations
date and versiondated, versioned, re-runnable
graded bytwo reviewers independently, agreement measured and reported
cases drawn froma set assembled by a team that did not build the agent
failurestwelve shown in full, grouped by cause
stated in advancebelow the current human rate, we would not ship

The marking

One, two and three each fail on something structural. Four does not.

reportwhat it failsthe one question to ask
Oneeverything. It is a headline with no result behind itcompared to what?
Twothe runs that did not finish, and an unchecked graderwhat happened to the questions it could not answer?
Threewhere the cases came from — a public collection the system may well have trained oncould it have seen these issues already?
Fournothing on this list—

Report three is the one most people pass, because it is visibly careful: it excludes runs and says so, it has a reference point, it is dated, and its grading is sensible. Being careful in seven ways does not help if the cases were memorized.

What the eight questions do not catch

Report four answers all eight. It is honest, well-built and re-runnable, and the people who made it did everything this volume asks. It is also measuring the wrong thing.

It measures whether the drafted reply is factually correct. Nobody in the business cares whether replies are factually correct. They care whether the customer's problem got resolved, and a reply can be entirely accurate and resolve nothing — correct, polite, on topic, and leaving the person exactly where they started.

the fourth one passes all eight and is still wrong the eight questions each one a filter caught at the first caught in the middle caught late and still wrong three are caught. one clears every gate. a careful report, competently made, measuring the wrong thing entirely
Three reports are caught. The fourth clears every gate and is still wrong, which is the limit of any checklist including this one.
What moves: four reports drop through eight filters. Three are stopped at different filters; the fourth passes through all of them and lands in a box labeled and still wrong.

No structural question can catch this, and a ninth would not help. Everything about report four is in order. The defect is a gap between what was measured and what anybody wanted, and closing it needs someone who knows the business well enough to notice — which is you, and is the part nobody can hand off to a checklist.

So the eight questions have one prerequisite

  • Ask what this number is for before asking whether it is sound. The eight questions test whether a measurement is trustworthy. They cannot test whether it is relevant, and a trustworthy measurement of the wrong quantity is the most expensive kind of report there is, because it survives scrutiny.
  • They are still worth having. Three of the four reports here fail on structure, and that is the usual proportion. Most bad results are bad in ways the questions catch.
  • They are no defense against someone who knows them. A report built to pass this list will pass it. At that point you are back to the last question in part four: what would have changed the recommendation.

Now go back to the first page of this volume. The card there has not changed.

→ Re-read the card from the first step. It is the same card, and the two conclusions it invited — that the agent works, and that what remains is a business decision — should now look like what they are: neither of them available from anything printed on it.

Every term this page uses

denominator
The count a rate was worked out over. Ninety-four percent of what, and of how many.
baseline
What the same measurement says about doing it the old way, or the trivial way, or not at all.
held-out
Cases kept away from whoever built the thing, so passing them means something.
contamination
The thing being tested has already seen the test cases, so the score measures memory rather than ability.
grader
Whoever or whatever looked at each answer and decided it was right or wrong.
agreement
How often two graders looking at the same answers said the same thing.
abstention
A grader allowed to return a third verdict: I cannot tell.
stale
Measured on a version that no longer exists.
exhibit
A report made up for this volume, for you to judge. Its numbers are not claims about the world.

Taken as already known, and so not defined here: eval, test, model, agent, metric. That list is a claim about who is reading, and it is printed so it can be argued with.