Lesson two, running live

200 A/A experiments in. 72 out.

Two hundred experiments where the treatment is identical to the control. Every “win” below is false by construction.

Checked every morning
36.0%
Always-valid bound
1.0%

Nominal false-positive rate is 5.0%. Peeking daily turns that into 36.0% without anyone changing a single line of analysis code.

The six lessons
#The mistakeThe fix
1Your eval cannot see what you are looking forPower and minimum detectable effect
2Checking every morning is not freeAlways-valid sequential testing
3Twelve slices is twelve testsBenjamini-Hochberg
4Your confidence score is not a probabilityConformal prediction
5The top of your leaderboard is noiseEmpirical Bayes shrinkage
6A percentage throws away when it happenedSurvival analysis