Live demo
1,890 paired cases in. 3 out.
One eval run, v41 against v42, cut into 12 customer slices. The aggregate moved +0.5pp and the gate the team runs today passed it. Three slices did not.
| Slice | Cases | Baseline | Candidate | Delta | 95% CI | p | q | |
|---|---|---|---|---|---|---|---|---|
| en-free | 220 | 72.9% | 81.6% | +8.7pp | +3.4pp to +14.1pp | 0.001 | 0.004 | improved |
| en-pro | 210 | 80.4% | 86.7% | +6.4pp | +2.5pp to +10.4pp | 0.001 | 0.004 | improved |
| en-enterprise | 200 | 75.9% | 76.6% | +0.7pp | -2.0pp to +3.4pp | 0.572 | 0.762 | flat |
| es-free | 160 | 70.6% | 72.0% | +1.5pp | -0.4pp to +3.9pp | 0.236 | 0.404 | flat |
| es-pro | 150 | 71.0% | 72.3% | +1.3pp | -2.0pp to +4.8pp | 0.449 | 0.674 | flat |
| fr-pro | 140 | 71.7% | 72.1% | +0.4pp | -3.1pp to +3.9pp | 0.760 | 0.870 | flat |
| pt-free | 130 | 72.4% | 73.1% | +0.7pp | -4.0pp to +5.3pp | 0.811 | 0.870 | flat |
| it-pro | 120 | 76.1% | 72.5% | -3.6pp | -7.3pp to -0.7pp | 0.029 | 0.058 | suppressed |
| nl-pro | 120 | 78.5% | 78.0% | -0.5pp | -4.4pp to +3.2pp | 0.870 | 0.870 | flat |
| de-enterprise | 140 | 78.4% | 70.6% | -7.9pp | -12.4pp to -3.9pp | <0.001 | 0.001 | regressed |
| ja-enterprise | 130 | 72.2% | 66.7% | -5.5pp | -9.1pp to -2.4pp | <0.001 | 0.001 | regressed |
| ko-enterprise | 170 | 70.1% | 65.4% | -4.7pp | -8.2pp to -1.6pp | 0.004 | 0.009 | regressed |
1 slice cleared p < 0.05 on its own and did not survive correction: it-pro (p=0.029, q=0.058). Checking 12 slices uncorrected fails a candidate that changed nothing 46.0% of the time, which is why teams stop checking slices.
Assumptions, stated. The test treats cases within a slice as exchangeable and independent; a suite that templates one prompt into many variants breaks that and the p-values come out optimistic. 14.9% of cases here carry a rubric fraction rather than a 0/1, and the smallest slice has 120 cases. Both go through the same paired test; neither is approximated as normal.