The deliverable

The rollout verdict

Two readouts on the same rollout, from the same data, differing only in which comparison group was used. One is an answer. The other is a refusal, and it is the one that is hard to buy anywhere else.

Self-serve Team tier, rolled out to North America first
Live Jun 29, 2026 · Trial-to-paid conversion rate · US-West, US-East, Canada
Estimate issueddifference-in-differences
Verdict

The rollout moved trial-to-paid conversion rate by +0.41 pp, from 3.23% to an underlying 3.64%. The dashboard shows +1.45 pp.

The control regions rose +1.04 pp over the same window without getting the change, so 72% of the before/after number was movement the rollout did not cause. The effect is real and it is about a quarter of what the chart implies.

Numbers
Treated regions, before3.23%90 days
Treated regions, after4.68%60 days
Control regions, same window3.12% → 4.16%+1.04 pp
Before/after read+1.45 ppconfounded
Difference-in-differences+0.41 ppreported
Checks
Parallel pre-trends0.982threshold 0.80
Placebo rollout on the pre-period+0.02 ppvs +0.41 pp real
Agreement across control regions taken alone+0.40 pp to +0.42 pp4 regions
What this verdict does not decide

It compares regions that got the change to regions that did not, over one window. It cannot tell you whether the same change lands the same way in the control regions when it eventually ships there, and it does not separate the tier itself from the launch messaging around it. It also assumes the rollout had no effect on the control regions, which stops being true the moment the two groups share customers.

Same rollout, control group: Japan
Trial-to-paid conversion rate
No estimate issuedparallel-trends check failed
Verdict

Japan moved with the treated regions at 0.18 before the rollout, against a threshold of 0.80. It is not a control group and no effect is reported against it.

Separate marketing calendar. Almost none of its week-to-week movement is shared with the treated regions.

The number you would have gotten
+1.16 ppagainst a true effect of +0.42 pp, so 2.8 times too large

Shown once, so the size of the mistake is visible, then withheld. It is not a noisy estimate of the right quantity. Subtracting a group that was already moving differently produces a different quantity in the same units, and the error does not shrink with more data.

What to do instead

Use the 4 regions that passed. If none had passed, the honest options are a staggered rollout that creates its own control, a synthetic control built from a weighted blend rather than a single group, or accepting that this change is not measurable and saying so before the readout gets quoted.

Same rollout, control group: Brazil
Trial-to-paid conversion rate
No estimate issuedparallel-trends check failed
Verdict

Brazil moved with the treated regions at 0.43 before the rollout, against a threshold of 0.80. It is not a control group and no effect is reported against it.

In structural decline since a local competitor cut prices. Comparable level, opposite trend.

The number you would have gotten
+1.34 ppagainst a true effect of +0.42 pp, so 3.2 times too large

Shown once, so the size of the mistake is visible, then withheld. It is not a noisy estimate of the right quantity. Subtracting a group that was already moving differently produces a different quantity in the same units, and the error does not shrink with more data.

What to do instead

Use the 4 regions that passed. If none had passed, the honest options are a staggered rollout that creates its own control, a synthetic control built from a weighted blend rather than a single group, or accepting that this change is not measurable and saying so before the readout gets quoted.

Why the refusal is the product

Any tool can subtract two averages. The reason teams get this wrong is not that the arithmetic is hard, it is that the control group gets picked by whoever is in the room and nobody checks whether it was ever comparable. Both refused groups above look completely reasonable on a chart of levels, and both produce a number close to the naive read while feeling like rigour. Printing the check, the threshold, and the number the bad control would have given is the only way to make that visible before it becomes a strategy.