Live demo
3,000 judged outputs; human labels in. 300 out.
A support agent graded by a lenient LLM judge. 300 of the outputs were also graded by a person. Watch the correction turn the judged pass rate into an estimate of the true one, and what it does to a prompt comparison the judge called a wash.
| Variant | Judged | Judge says | Corrected | 95% interval |
|---|---|---|---|---|
| prompt v1 | 3,000 | 70.4% | 62.4% | 48.7% to 73.1% |
| prompt v2 | 3,000 | 69.4% | 60.2% | 47.1% to 71.3% |
v2 reads -2.2pp after correction, -0.9pp judged. The interval includes zero: this eval cannot certify v2 as an improvement, and the judged number was compressing a loss toward parity.
A judge with separation 0.43 reports every true difference multiplied by 0.43. That is why two variants that differ by a few points read as identical.
The judge passes 43.3% of outputs a human fails. The reported number is a linear function of the true one, and the line has been known since 1978.
| Queue | Share of judged set | Share of labelled slice | Judge pass rate, set | Judge pass rate, slice |
|---|---|---|---|---|
| Account access | 20.0% | 20.0% | 69.3% | 66.7% |
| Shipping | 20.0% | 20.0% | 73.8% | 68.3% |
| Returns | 20.0% | 20.0% | 68.3% | 61.7% |
| Integrations | 20.0% | 20.0% | 68.3% | 73.3% |
| Billing | 20.0% | 20.0% | 72.0% | 76.7% |
If one queue is judged much more leniently than the rest and the labelled slice under-samples it, the judge's error rates measured on the slice are not the error rates on the set, and the corrected number is off. Here the queues are judged alike, so a single pair of error rates is defensible. The interval is wide because 300 labels pin the judge's false-pass rate to a few points either way; halving it takes about four times the labels.