Live demo

200 A/A canaries in. 6 out.

200 canaries where the upgrade changes nothing, each watched for 14 days. The morning t-test halts 44 of them. The sequential test halts 6, which is what alpha 0.05 promised. Then the same test on two upgrades that do change something.

Halted by daily peeking
22.0%
44 of 200 neutral upgrades
Halted by sequential
3.0%
promised at most 5.0%
Regression caught at
30,000
requests, 1,500 on the canary
Improvement
promote
after 4,000 canary requests
Two upgrades with a real effect, checked after every batch
RolloutTruthDecisionCanary requestsEstimateAlways-valid pDaily test crossed at
prompt-v42prompt-v42, truly 3 points worseroll back1,500-4.7pt0.012500
prompt-v43prompt-v43, truly 2 points betterpromote4,000+2.4pt0.0301,750
Evidence on prompt-v42, batch by batch
Canary requestsRequests seenEstimateDaily pAlways-valid p
2505,000+2.8pt0.3691.000
50010,000-5.0pt0.0330.395
75015,000-2.9pt0.1250.805
1,00020,000-4.5pt0.0060.118
1,25025,000-4.2pt0.0050.090
1,50030,000-4.7pt0.0000.012
Neutral upgrades halted, daily t-test
22.0%
Neutral upgrades halted, sequential
3.0%

Same traffic, same 14 looks. The daily test crossed on prompt-v42 earlier than the sequential one did, and it also crosses on 22.0% of nothing. A test that fires early on noise is not a fast test.

AssumptionCanary and control traffic must be exchangeable

The guarantee holds when requests are split at random. A sticky-by-customer split, or a canary that only sees one region, breaks it. Each batch here pairs one control request with each canary request and leaves the rest of the control arm unused, which is conservative: the 5.0% slice is the bottleneck, not the control.