Live demo

3,000 judged outputs; human labels in. 300 out.

A support agent graded by a lenient LLM judge. 300 of the outputs were also graded by a person. Watch the correction turn the judged pass rate into an estimate of the true one, and what it does to a prompt comparison the judge called a wash.

Judge-reported pass rate
70.4%
Corrected pass rate
62.4%
95% interval 48.7% to 73.1%
Judge true-pass / false-pass
86.7% / 43.3%
from 300 human labels
Separation
0.43
above the 0.2 floor, correction usable
Two prompt variants, same judge, same calibration
VariantJudgedJudge saysCorrected95% interval
prompt v13,00070.4%62.4%48.7% to 73.1%
prompt v23,00069.4%60.2%47.1% to 71.3%
v2 minus v1Judged -0.9pp, corrected -2.2pp, interval -7.9pp to +3.2pp

v2 reads -2.2pp after correction, -0.9pp judged. The interval includes zero: this eval cannot certify v2 as an improvement, and the judged number was compressing a loss toward parity.

A judge with separation 0.43 reports every true difference multiplied by 0.43. That is why two variants that differ by a few points read as identical.

Reported today
70.4%
Corrected
62.4%

The judge passes 43.3% of outputs a human fails. The reported number is a linear function of the true one, and the line has been known since 1978.

Diagnostic: the correction assumes the labelled slice is drawn like the judged set
QueueShare of judged setShare of labelled sliceJudge pass rate, setJudge pass rate, slice
Account access20.0%20.0%69.3%66.7%
Shipping20.0%20.0%73.8%68.3%
Returns20.0%20.0%68.3%61.7%
Integrations20.0%20.0%68.3%73.3%
Billing20.0%20.0%72.0%76.7%

If one queue is judged much more leniently than the rest and the labelled slice under-samples it, the judge's error rates measured on the slice are not the error rates on the set, and the corrected number is off. Here the queues are judged alike, so a single pair of error rates is defensible. The interval is wide because 300 labels pin the judge's false-pass rate to a few points either way; halving it takes about four times the labels.

Agent
support-answer v1
Outputs judged
3,000
Human-labelled
300
Compared against
prompt v2