Live demo

62 cases today in. 4 out.

The suite a Series A team actually had, scored against five checks. Two weeks later, the same scorecard.

What they had
4.0/10
What was delivered
8.0/10

62 cases became 840, sized from the effect they said they cared about rather than from anecdote.

failCan it see the effect you care about6.7pp

Smallest regression this suite can detect is 6.7 points; you said you care about 3.

warnHow much movement is the judge0%

No repeated judgements, so judge noise is unmeasured — which is itself the finding.

failEvery slice big enough to report3 thin

multilingual (9), ambiguous (9), adversarial (10) are under 30 cases and cannot support a per-slice claim.

passNot saturated or floored90%

Pass rate sits in the range where changes are detectable.

warnNo duplicate cases4

4 duplicate cases are double-weighting whatever they test.

Required cases per slice is a formula, not a judgement call. The reason nearly every suite is the wrong size is that nobody runs it.