Live demo

This experiment has no effect in it at all.

Both arms were generated from the same 12% conversion rate. There is nothing to find. Watch what happens when you monitor it the way everyone actually monitors a running test — by looking at it.

A/A experiment · true effect = 0Naive test declares a winnerSequential test does not
Naive stops at
n = 120
first crossing of p < 0.05
Apparent lift there
+7.50 pts
entirely noise
Sequential p-value
0.41
best it ever reached
Effect at full sample
+2.50 pts
where it actually settles
naive p-value (fixed horizon) always-valid p-value (mSPRT)log scale · lower is more significant
10.1α = .050.010.001naive stops here0 observations600 observations

From inside the experiment there is no way to tell this apart from a real win. The treatment is up two and a half points and the p-value is under 0.05. A team with a stopping rule of “stop when significant” ships it, books the lift, and never finds out. The always-valid p-value never gets close, because it is priced for the fact that you are looking repeatedly.

Not one unlucky experiment

200 A/A tests, monitored continuously

Two hundred experiments with no effect in any of them, each one watched every ten observations, exactly as a growth team watches a dashboard. Every declared winner below is false by construction.

Naive rule
36.0%
72 of 200 A/A tests declared a winner
Sequential rule
1.0%
2 of 200, near the nominal 5%
Nominal rate
5.0%
what everyone assumes they have
Inflation factor
36.0×
false wins vs the corrected test
Naive, monitored continuously36.0%
Sequential, monitored continuously1.0%
Nominal rate everyone assumes5.0%

The nominal 5% is only valid if you look exactly once, at a sample size fixed before you started. Peek continuously and you are taking many shots at the same threshold — and you stop precisely on the shot that crosses it. The rate above is not a modelling artefact; it is what the stopping rule does.

The other half

A real lift still gets called, and called early

A correction that never fires is not a correction, it is a brake. This experiment has a genuine 2.4 point lift, and the sequential test stops on it — legitimately, because the always-valid p-value is licensed for exactly this stopping rule.

True effect
+2.40 pts
injected ground truth
Sequential stops at
n = 1560
valid early stop
Measured effect
+2.80 pts
at full sample
Sample saved
61%
vs running to the fixed horizon
naive p-value (fixed horizon) always-valid p-value (mSPRT)log scale · lower is more significant
10.1α = .050.010.001naive stops here0 observations4,000 observations

This is the trade the method actually offers. You do not give up early stopping — you give up pretending that a fixed-horizon p-value survives being checked every morning. Real effects still get called, often sooner, and the ones you call are the ones that hold.