The deliverable
The rollout verdict
Two readouts on the same rollout, from the same data, differing only in which comparison group was used. One is an answer. The other is a refusal, and it is the one that is hard to buy anywhere else.
The rollout moved trial-to-paid conversion rate by +0.41 pp, from 3.23% to an underlying 3.64%. The dashboard shows +1.45 pp.
The control regions rose +1.04 pp over the same window without getting the change, so 72% of the before/after number was movement the rollout did not cause. The effect is real and it is about a quarter of what the chart implies.
| Treated regions, before | 3.23% | 90 days |
| Treated regions, after | 4.68% | 60 days |
| Control regions, same window | 3.12% → 4.16% | +1.04 pp |
| Before/after read | +1.45 pp | confounded |
| Difference-in-differences | +0.41 pp | reported |
It compares regions that got the change to regions that did not, over one window. It cannot tell you whether the same change lands the same way in the control regions when it eventually ships there, and it does not separate the tier itself from the launch messaging around it. It also assumes the rollout had no effect on the control regions, which stops being true the moment the two groups share customers.
Japan moved with the treated regions at 0.18 before the rollout, against a threshold of 0.80. It is not a control group and no effect is reported against it.
Separate marketing calendar. Almost none of its week-to-week movement is shared with the treated regions.
Shown once, so the size of the mistake is visible, then withheld. It is not a noisy estimate of the right quantity. Subtracting a group that was already moving differently produces a different quantity in the same units, and the error does not shrink with more data.
Use the 4 regions that passed. If none had passed, the honest options are a staggered rollout that creates its own control, a synthetic control built from a weighted blend rather than a single group, or accepting that this change is not measurable and saying so before the readout gets quoted.
Brazil moved with the treated regions at 0.43 before the rollout, against a threshold of 0.80. It is not a control group and no effect is reported against it.
In structural decline since a local competitor cut prices. Comparable level, opposite trend.
Shown once, so the size of the mistake is visible, then withheld. It is not a noisy estimate of the right quantity. Subtracting a group that was already moving differently produces a different quantity in the same units, and the error does not shrink with more data.
Use the 4 regions that passed. If none had passed, the honest options are a staggered rollout that creates its own control, a synthetic control built from a weighted blend rather than a single group, or accepting that this change is not measurable and saying so before the readout gets quoted.
Any tool can subtract two averages. The reason teams get this wrong is not that the arithmetic is hard, it is that the control group gets picked by whoever is in the room and nobody checks whether it was ever comparable. Both refused groups above look completely reasonable on a chart of levels, and both produce a number close to the naive read while feeling like rigour. Printing the check, the threshold, and the number the bad control would have given is the only way to make that visible before it becomes a strategy.