Live demo
3 claims in. 3 out.
Three claims from a Series A deck. Two do not survive checking, one does — and the difference is recoverable from the summary alone.
Stopped after 14 looks, so the real false-positive rate is 51%, not 5%
Checking a running experiment 14 times and stopping when it crosses significance inflates alpha from 0.05 to roughly 0.51. The reported p of 0.296 is not the probability it reads as.
Corrected: Re-run with an always-valid sequential bound, or hold to a fixed horizon.
Only 19% power to see the effect being claimed
At n=940 per arm this test would miss an effect of this size 81% of the time. A win here is as likely to be noise as signal.
41% of the cohort has not been around long enough to churn
Counting them as retained makes retention look like 92%. Restricted to customers actually observed for the full period, it is 87%.
Corrected: 87% on the observed cohort
The second experiment is left alone deliberately. A diligence tool that finds something wrong with every claim is not doing diligence, it is producing anxiety.