Live demo
4 dimensions scored in. 3 out.
The scorecard from week one, computed on a real agent. Then the same four dimensions at handover.
Weakest dimension on arrival: escalation calibration. The blend is half mean and half minimum, so it cannot be averaged away.
A threshold rule would have paged 88 times over this window. After FDR correction, 5 are real.
The current threshold delivers 5.9% errors against a stated budget of 2%. It was picked by hand and has no guarantee attached.
The suite can see a 6.7 point regression. You said you care about 3.
A suite that passes 61% of the time on unchanged code cannot tell you whether a change broke something.
Verdict on arrival: do not ship, with roughly 8 days of work to close the gaps. That estimate is what the fixed price is built on, and you see it before committing to anything.