Live demo
The rollout worked. It worked less than a third as well as the chart says.
Self-serve Team tier, rolled out to North America first, live Jun 29. Trial-to-paid conversion rate rose sharply in the regions that got it. It also rose in the regions that did not, because the season turned in the same week. This data has a known effect injected into it, so the right answer is available to check every number against.
The generator added exactly +0.42 pp to the treated regions and +0.60 pp of seasonal lift to every region including the controls. A before/after read collects both and reports +1.45 pp, which is 3.4 times the truth. Subtracting the control group recovers +0.41 pp. Nobody had to know the season was coming.
What the two groups did
Both lines went up. Only the distance between them is yours.
Treated regions against the pooled control, 90 days before the rollout and 60 after. The two tracked each other closely beforehand, which is the assumption difference-in-differences needs, and it is checked rather than assumed.
The gate
Every candidate control is scored before any number comes out
Difference-in-differences is only as good as the claim that the two groups were moving together before the change. That claim is checkable. The threshold is 0.80 pre-period correlation, stated so you can argue with it, and a candidate below it does not get an estimate at all.
| Candidate control | Pre-trend correlation | Verdict | Estimate it produces |
|---|---|---|---|
UK & Ireland Same product, same calendar, no rollout | 0.95 | Usable | +0.40 pp |
DACH Same product, same calendar, no rollout | 0.96 | Usable | +0.42 pp |
Nordics Same product, same calendar, no rollout | 0.95 | Usable | +0.40 pp |
Australia & NZ Same product, same calendar, no rollout | 0.97 | Usable | +0.40 pp |
Japan Separate marketing calendar. Almost none of its week-to-week movement is shared with the treated regions. | 0.18 | Below 0.80 | +1.16 pp (2.8× truth, not reported) |
Brazil In structural decline since a local competitor cut prices. Comparable level, opposite trend. | 0.43 | Below 0.80 | +1.34 pp (3.2× truth, not reported) |
Levels are removed, so the only thing left is shape. The controls that pass ride the same weekly demand as the treated regions. The two that fail do not, and no amount of arithmetic downstream repairs that.
The refusal
What a bad control would have told you
Both of these look like reasonable comparison groups. Same product, same metric, same window, similar levels. Both produce a number that is several times the truth, and nothing in a dashboard would have flagged either one.
Separate marketing calendar. Almost none of its week-to-week movement is shared with the treated regions.
This is not a noisy estimate of the right thing. Subtracting a group that was already moving differently produces a different quantity wearing the same units, and it is confidently wrong in a direction you cannot predict from the number alone.
In structural decline since a local competitor cut prices. Comparable level, opposite trend.
This is not a noisy estimate of the right thing. Subtracting a group that was already moving differently produces a different quantity wearing the same units, and it is confidently wrong in a direction you cannot predict from the number alone.
The second check
Run the same estimator on a week nothing shipped
The estimator is refitted on the pre-period alone with a fake rollout date halfway through it. Nothing shipped then, so it should return roughly zero. If it returned a large number, the design would be reading structure rather than the change, and the real estimate would be worth nothing.
4 control regions, each taken on its own, land between +0.40 pp and +0.42 pp. They were not fit together and they agree. That agreement, the placebo above, and the pre-trend check are the whole case for the number. None of them is a p-value, and all of them are things you can look at.