AI engineering · 40 of 42
Nothing changed, and it got worse
Scroll
Nothing changed, and it got worse
Your prompt is the same, the model version is pinned, the code has not moved — and quality is falling. What changed is everything around it: new products, new phrasing from users, new failure modes, documents rewritten under the index.
It is gradual, which is what makes it dangerous. There is no incident, no alert and no commit to blame, so it gets noticed as a vague sense that things used to be better.
The defenses are unglamorous: run the golden set on a schedule rather than only on change, watch the score as a trend, sample real traffic periodically, and keep adding cases. A metric you only look at during deploys cannot show you a slope.
Operations