Evaluation, not production, is now your constraint
Producing work got cheap and checking it did not, so checking is what you are short of — and you hired for the other one.
Here is the whole argument before any of it is defended. If you read nothing else, this is the thing to disagree with.
The cost of producing candidate work fell sharply. The cost of judging whether a piece of work is good did not fall, because judging needs knowing the domain well enough to recognize a wrong answer, and that knowledge was always the expensive part. So the amount produced grows away from the amount anybody actually looked at.
The remainder does not announce itself. It merges, it ships, and it becomes what the next round is built on. Throughput and time-to-ship improve during exactly the interval when this is happening, so your reporting says the opposite of the truth. And the obvious remedy fails on three counts you can check yourself: the checking capacity you would need grows with production, the checkers have to know the domain, and nothing would tell you afterward whether it helped. What works is making less of the work need judgment at all, and staffing judgment as a career rather than a chore.
What this volume is not
It is not an argument for producing less, and it is not a claim that the work being produced is bad. Most of it is fine. The claim is narrower: the share of it that anybody has looked at is falling, you have no instrument that reads that share, and the remedy everybody reaches for first does not work.
It is also not about your reviewers, who are doing their jobs. Everything here is about a ratio between two quantities, one of which grew and one of which did not.
The uncomfortable property. Every number you report is correct and every one is going the right way. That is not a paradox to be resolved — it is the same event seen from the end that has instruments on it, and part two is about exactly that.