The Judgment Gap three reasons more reviewers will not close it › You would not be able to tell whether it worked 12 steps, no labs
09 — part 3, three reasons more reviewers will not close itsourced

You would not be able to tell whether it worked

Add people to check things and then ask what would show you they were catching anything. Nothing would.

The third reason, and it is the one that survives any disagreement about the other two. Suppose you added the checkers and could afford them. What would tell you it had worked?

one of these added checkers. Nothing on either panel says which as things are throughput time to ship queue depth the remainder with checkers added throughput time to ship queue depth checkers added the remainder the levels differ. Nothing measures the level. so the remedy is expensive, and cannot be evaluated even afterward.
Two organizations with identical panels. The reservoir levels differ and nothing measures the level, so the remedy cannot be evaluated even afterward.
What moves: two rigs side by side fill their instrument panels identically while their reservoirs settle at different levels.

Work through what would move. Throughput would go down slightly, because more work is waiting on people. Time-to-ship would go up. Queue depth would go up. Every number on your panel would get worse, and the thing that got better — the remainder — is the quantity with no gauge on it.

what you would seewhat it means
throughput downthe intervention is being felt. Not that it is working
time to ship upthe same. This is the cost, not the benefit
reviewer load upyou spent the money
fewer of the failures from step sixthe actual benefit, arriving quarters later, indistinguishable from luck, and impossible to attribute

So the position you would be in is: a visible cost, an invisible benefit, and no way to tell whether you bought anything. That is a bad position to be in even when the intervention is correct, and it is why this remedy gets tried, quietly abandoned after two quarters, and remembered as not having worked.

There is a fourth reason, and this volume does not lean on it. There is research suggesting the checkers themselves degrade when most of what they see is correctvigilance and automation-complacency researchunverified. If that holds up it makes everything on this page worse — but it is a single unverified citation, and the three reasons above are arithmetic and mechanism you can check yourself. Treat it as a reason to take part four seriously rather than as evidence for anything here.

→ For one review step your team already runs, say what would tell you it is working. Not what it caught — what would show you it is catching a useful share of what there is. If the answer is that nothing would, you have found the shape of this problem inside something you already own and pay for.

Every term this page uses

produced-to-examined
How much work your team produced, divided by how much of it somebody actually looked at.
examined
Looked at by somebody with enough context to have caught a problem. Not the same as having passed through a review step.
remainder
Work that was produced and not examined. It does not go anywhere — it merges into what everything after it is built on.
judgment
Deciding whether a piece of work is good, which needs knowing the domain well enough to recognize a wrong answer.
inspection
Checking finished work at the end, rather than changing how the work is made.
checkable
Able to be judged without understanding the whole of it. The property that decides how much needs judgment at all.
constraint
The thing that limits output. Adding more of anything else does not help until it moves.
rework
Work that had to be revisited after somebody called it done.
vigilance
Staying alert for a rare problem across a long stretch of time.

Taken as already known, and so not defined here: agent, model, review, throughput, headcount. That list is a claim about who is reading, and it is printed so it can be argued with.