The sizing

How many cases each slice actually needs

The output of the free scorecard: a target, computed from the effect you said you cared about.

eval-build size --effect 3ppfail
current suite: 62 cases across 4 slices
pass rate 90.3% · duplicates 4

  FAIL  Can it see the effect you care about   6.7pp
  WARN  How much movement is the judge         0%
  FAIL  Every slice big enough to report       3 thin
  PASS  Not saturated or floored               90%
  WARN  No duplicate cases                     4

required per slice for a 3pp effect at 80% power: 305
thin slices: multilingual(9), ambiguous(9), adversarial(10)

score 4.0/10