The Judgment Gap how the gap between producing and checking opens › Why producing got cheap and judging did not 12 steps, no labs
02 — part 1, how the gap between producing and checking opensreasoning

Why producing got cheap and judging did not

Making something and knowing whether it is any good are different activities, and only one of them got cheaper.

Two activities that were always spoken about as one thing. Making a candidate answer, and knowing whether it is any good. Only the first got cheaper, and the reason is worth a page because it explains why nothing will make the second follow.

work produced work examined they were never joined, so nothing had to break producing a candidate answer got much cheaper knowing whether it is any good did not people assumed a link like this there was never one judging needs knowing the domain well enough to recognize a wrong answer — and that was always the expensive ingredient. nothing automated it, so nothing made it cheaper. two costs that moved together for years, for no reason anybody had written down.
They were never joined, so nothing had to break. The link people assumed was there never existed.
What moves: the produced pipe widens across the cycle while the examined pipe stays the same, and a dashed link between them appears briefly before being labeled as never having existed.

Judging a piece of work requires knowing the domain well enough to recognize a wrong answer. Not to produce the right one — to recognize a wrong one, which is cheaper and is still the expensive ingredient. Nothing about the last few years made that knowledge cheaper to acquire or faster to spread.

the activitywhat it needswhat happened to its cost
producing a candidate answera description of what is wantedcollapsed, and unevenly, which is a separate problem
judging whether it is any goodknowing the domain well enough to recognize a wrong answerunchanged. Nothing automated the knowing

Why the two used to move together

Because for a long time the only way to produce a lot of work was to have a lot of people who knew the domain, and those people could also judge it. Production capacity and judgment capacity were carried by the same bodies, so they scaled together automatically and nobody had to notice that they were separate things.

They have come apart. Production is now available without the knowledge behind it, and judgment still is not. That is the entire mechanism, and everything else in this volume is a consequence of it.

Which is why the constraint moved. The thing limiting what your organization can safely ship is no longer how much it can make. It is how much it can assess. If you staffed for the first, you staffed for the wrong one, and not through any error — the two were indistinguishable until recently.

Every term this page uses

produced-to-examined
How much work your team produced, divided by how much of it somebody actually looked at.
examined
Looked at by somebody with enough context to have caught a problem. Not the same as having passed through a review step.
remainder
Work that was produced and not examined. It does not go anywhere — it merges into what everything after it is built on.
judgment
Deciding whether a piece of work is good, which needs knowing the domain well enough to recognize a wrong answer.
inspection
Checking finished work at the end, rather than changing how the work is made.
checkable
Able to be judged without understanding the whole of it. The property that decides how much needs judgment at all.
constraint
The thing that limits output. Adding more of anything else does not help until it moves.
rework
Work that had to be revisited after somebody called it done.
vigilance
Staying alert for a rare problem across a long stretch of time.

Taken as already known, and so not defined here: agent, model, review, throughput, headcount. That list is a claim about who is reading, and it is printed so it can be argued with.