The eval report

The number that goes in the launch review

Not the judge's opinion. The corrected pass rate with an interval, the judge's measured error rates, and a verdict on the variant comparison that says what the evidence can and cannot support.

Eval report, corrected

gpt-class rubric judge, lenient: 3,000 outputs, 300 human labels

Judge reported
70.4%
Corrected
62.4%
95% interval
48.7% to 73.1%
Judge TPR / FPR
86.7% / 43.3%

The judge passes 86.7% of outputs a human passes and 43.3% of outputs a human fails, measured on 300 labelled outputs. Separation 0.43, above the floor for a usable correction.

Reported pass rate 70.4%. Corrected pass rate 62.4%, interval 48.7% to 73.1%, from a bootstrap over both the labelled slice and the judged set.

Prompt v2 against v1: judged -0.9pp, corrected -2.2pp, interval -7.9pp to +3.2pp.

Verdict: v2 reads -2.2pp after correction, -0.9pp judged. The interval includes zero: this eval cannot certify v2 as an improvement, and the judged number was compressing a loss toward parity.