The Judgment Gap how the gap between producing and checking opens › Evaluation, not production, is now your constraint 12 steps, no labs
01 — part 1, how the gap between producing and checking opensreasoning

Evaluation, not production, is now your constraint

Producing work got cheap and checking it did not, so checking is what you are short of — and you hired for the other one.

Here is the whole argument before any of it is defended. If you read nothing else, this is the thing to disagree with.

The cost of producing candidate work fell sharply. The cost of judging whether a piece of work is good did not fall, because judging needs knowing the domain well enough to recognize a wrong answer, and that knowledge was always the expensive part. So the amount produced grows away from the amount anybody actually looked at.

The remainder does not announce itself. It merges, it ships, and it becomes what the next round is built on. Throughput and time-to-ship improve during exactly the interval when this is happening, so your reporting says the opposite of the truth. And the obvious remedy fails on three counts you can check yourself: the checking capacity you would need grows with production, the checkers have to know the domain, and nothing would tell you afterward whether it helped. What works is making less of the work need judgment at all, and staffing judgment as a career rather than a chore.

work produced work examined the remainder nothing breaks. The level simply rises wide, and getting wider narrow, and fixed no gauge reads this the difference does not spill and does not complain.
Nothing breaks and nothing spills. The level simply rises, and no gauge reads it.
What moves: work flows through a wide produced pipe and a narrow examined pipe, and the difference collects in a reservoir below that rises steadily.

What this volume is not

It is not an argument for producing less, and it is not a claim that the work being produced is bad. Most of it is fine. The claim is narrower: the share of it that anybody has looked at is falling, you have no instrument that reads that share, and the remedy everybody reaches for first does not work.

It is also not about your reviewers, who are doing their jobs. Everything here is about a ratio between two quantities, one of which grew and one of which did not.

The uncomfortable property. Every number you report is correct and every one is going the right way. That is not a paradox to be resolved — it is the same event seen from the end that has instruments on it, and part two is about exactly that.

Every term this page uses

produced-to-examined
How much work your team produced, divided by how much of it somebody actually looked at.
examined
Looked at by somebody with enough context to have caught a problem. Not the same as having passed through a review step.
remainder
Work that was produced and not examined. It does not go anywhere — it merges into what everything after it is built on.
judgment
Deciding whether a piece of work is good, which needs knowing the domain well enough to recognize a wrong answer.
inspection
Checking finished work at the end, rather than changing how the work is made.
checkable
Able to be judged without understanding the whole of it. The property that decides how much needs judgment at all.
constraint
The thing that limits output. Adding more of anything else does not help until it moves.
rework
Work that had to be revisited after somebody called it done.
vigilance
Staying alert for a rare problem across a long stretch of time.

Taken as already known, and so not defined here: agent, model, review, throughput, headcount. That list is a claim about who is reading, and it is printed so it can be argued with.