Live demo

1,890 paired cases in. 3 out.

One eval run, v41 against v42, cut into 12 customer slices. The aggregate moved +0.5pp and the gate the team runs today passed it. Three slices did not.

Aggregate delta
+0.5pp
95% -0.7pp to +1.6pp
Aggregate gate
pass
0 regressions seen
Uncorrected per slice
4
slices at p < 0.05, no correction
After correction
3
slices at q < 0.05
Per slice, paired bootstrap on per-case differences, BH across the twelve
SliceCasesBaselineCandidateDelta95% CIpq
en-free22072.9%81.6%+8.7pp+3.4pp to +14.1pp0.0010.004improved
en-pro21080.4%86.7%+6.4pp+2.5pp to +10.4pp0.0010.004improved
en-enterprise20075.9%76.6%+0.7pp-2.0pp to +3.4pp0.5720.762flat
es-free16070.6%72.0%+1.5pp-0.4pp to +3.9pp0.2360.404flat
es-pro15071.0%72.3%+1.3pp-2.0pp to +4.8pp0.4490.674flat
fr-pro14071.7%72.1%+0.4pp-3.1pp to +3.9pp0.7600.870flat
pt-free13072.4%73.1%+0.7pp-4.0pp to +5.3pp0.8110.870flat
it-pro12076.1%72.5%-3.6pp-7.3pp to -0.7pp0.0290.058suppressed
nl-pro12078.5%78.0%-0.5pp-4.4pp to +3.2pp0.8700.870flat
de-enterprise14078.4%70.6%-7.9pp-12.4pp to -3.9pp<0.0010.001regressed
ja-enterprise13072.2%66.7%-5.5pp-9.1pp to -2.4pp<0.0010.001regressed
ko-enterprise17070.1%65.4%-4.7pp-8.2pp to -1.6pp0.0040.009regressed
Uncorrected per-slice check
4
After BH correction
3

1 slice cleared p < 0.05 on its own and did not survive correction: it-pro (p=0.029, q=0.058). Checking 12 slices uncorrected fails a candidate that changed nothing 46.0% of the time, which is why teams stop checking slices.

Assumptions, stated. The test treats cases within a slice as exchangeable and independent; a suite that templates one prompt into many variants breaks that and the p-values come out optimistic. 14.9% of cases here carry a rubric fraction rather than a 0/1, and the smallest slice has 120 cases. Both go through the same paired test; neither is approximated as normal.