The eval report
The number that goes in the launch review
Not the judge's opinion. The corrected pass rate with an interval, the judge's measured error rates, and a verdict on the variant comparison that says what the evidence can and cannot support.
gpt-class rubric judge, lenient: 3,000 outputs, 300 human labels
The judge passes 86.7% of outputs a human passes and 43.3% of outputs a human fails, measured on 300 labelled outputs. Separation 0.43, above the floor for a usable correction.
Reported pass rate 70.4%. Corrected pass rate 62.4%, interval 48.7% to 73.1%, from a bootstrap over both the labelled slice and the judged set.
Prompt v2 against v1: judged -0.9pp, corrected -2.2pp, interval -7.9pp to +3.2pp.
Verdict: v2 reads -2.2pp after correction, -0.9pp judged. The interval includes zero: this eval cannot certify v2 as an improvement, and the judged number was compressing a loss toward parity.