The Judgment Gap Introduction 12 steps, no labs

Why evaluation, not production, became your constraint

One number that tells you how wide the gap is, why every figure you report improves while it opens, three reasons more reviewers will not close it — and what does.

Producing work got cheap and checking it did not.

The cost of producing a candidate answer fell sharply. The cost of judging whether it is any good did not, because judging needs knowing the domain well enough to recognize a wrong answer — and that knowledge was always the expensive part. So the amount produced grows away from the amount anybody actually looked at.

This volume is twelve short steps on how wide that gap is on your own team, why every number you report improves while it opens, why the obvious remedy does not work, and what does. It assumes you have read nothing else in this series.

work produced work examined the remainder nothing breaks. The level simply rises wide, and getting wider narrow, and fixed no gauge reads this the difference does not spill and does not complain.
The spine of the volume: nothing breaks and nothing spills. The level simply rises, and no gauge reads it.
What moves: work flows through a wide produced pipe and a narrow examined pipe, and the difference collects in a reservoir below that rises steadily.

The number this turns on

Work produced in a period, divided by work somebody actually looked at. The division is easy; the difficulty is that examined has a comfortable reading and a useful one, and the comfortable reading gives you a figure that means nothing. Step three is about getting that right, and it comes before the evidence so the rest reads against your own number.

The obvious remedy fails on three counts you can check yourself. The checking capacity you would need grows with production. The checkers have to know the domain, which is the scarcest thing you have. And nothing would tell you afterward whether it worked, because what it would improve is the quantity with no gauge on it.

What does work gets bigger as the problem does. Making less of the work need judgment at all — four moves, each a real trade, listed with what it costs — and staffing judgment as a career rather than a chore, which is the half that needs a director specifically.

What you can do afterward

Compute one number about your own team and know what it means. Say which of your improving reported figures you can no longer read at face value. Have the sum ready before somebody proposes more reviewers. And start one of the four moves without asking anybody.

The ratio, the three reasons and the four moves are on one page you can print.

Every term this page uses

produced-to-examined
How much work your team produced, divided by how much of it somebody actually looked at.
examined
Looked at by somebody with enough context to have caught a problem. Not the same as having passed through a review step.
remainder
Work that was produced and not examined. It does not go anywhere — it merges into what everything after it is built on.
judgment
Deciding whether a piece of work is good, which needs knowing the domain well enough to recognize a wrong answer.
inspection
Checking finished work at the end, rather than changing how the work is made.
checkable
Able to be judged without understanding the whole of it. The property that decides how much needs judgment at all.
constraint
The thing that limits output. Adding more of anything else does not help until it moves.
rework
Work that had to be revisited after somebody called it done.
vigilance
Staying alert for a rare problem across a long stretch of time.

Taken as already known, and so not defined here: agent, model, review, throughput, headcount. That list is a claim about who is reading, and it is printed so it can be argued with.