Live demo

4 dimensions scored in. 3 out.

The scorecard from week one, computed on a real agent. Then the same four dimensions at handover.

Week one
2.5/10
At handover
9.3/10

Weakest dimension on arrival: escalation calibration. The blend is half mean and half minimum, so it cannot be averaged away.

goodDrift detection9.4
88 → 5 alerts

A threshold rule would have paged 88 times over this window. After FDR correction, 5 are real.

criticalEscalation calibration0.0
5.9% against a 2% budget

The current threshold delivers 5.9% errors against a stated budget of 2%. It was picked by hand and has no guarantee attached.

warningEval sensitivity4.5
6.7pp detectable

The suite can see a 6.7 point regression. You said you care about 3.

warningTest determinism6.1
61% clean-run pass rate

A suite that passes 61% of the time on unchanged code cannot tell you whether a change broke something.

Verdict on arrival: do not ship, with roughly 8 days of work to close the gaps. That estimate is what the fixed price is built on, and you see it before committing to anything.