Live demo

The rollout worked. It worked less than a third as well as the chart says.

Self-serve Team tier, rolled out to North America first, live Jun 29. Trial-to-paid conversion rate rose sharply in the regions that got it. It also rose in the regions that did not, because the season turned in the same week. This data has a known effect injected into it, so the right answer is available to check every number against.

Trial-to-paid conversion rate, treated regionscontrol group passed at 0.983.23%4.68%
Before/after says+1.45 pp
+0.41 pp
+1.04 pp that was not the rollout
the changea season the untreated regions got too
Before/after
+1.45 pp
what the chart shows
Difference-in-differences
+0.41 pp
control moved +1.04 pp on its own
Truth injected
+0.42 pp
ground truth in this data
Confounded portion
72%
of the naive number was never yours

The generator added exactly +0.42 pp to the treated regions and +0.60 pp of seasonal lift to every region including the controls. A before/after read collects both and reports +1.45 pp, which is 3.4 times the truth. Subtracting the control group recovers +0.41 pp. Nobody had to know the season was coming.

What the two groups did

Both lines went up. Only the distance between them is yours.

Treated regions against the pooled control, 90 days before the rollout and 60 after. The two tracked each other closely beforehand, which is the assumption difference-in-differences needs, and it is checked rather than assumed.

2.60%3.22%3.85%4.47%5.09%rollout · Jun 29Apr 6Aug 27
Treated regions, pooledControl regions, pooled7-day trailing mean

The gate

Every candidate control is scored before any number comes out

Difference-in-differences is only as good as the claim that the two groups were moving together before the change. That claim is checkable. The threshold is 0.80 pre-period correlation, stated so you can argue with it, and a candidate below it does not get an estimate at all.

Candidate controlPre-trend correlationVerdictEstimate it produces
UK & Ireland
Same product, same calendar, no rollout
0.95
Usable+0.40 pp
DACH
Same product, same calendar, no rollout
0.96
Usable+0.42 pp
Nordics
Same product, same calendar, no rollout
0.95
Usable+0.40 pp
Australia & NZ
Same product, same calendar, no rollout
0.97
Usable+0.40 pp
Japan
Separate marketing calendar. Almost none of its week-to-week movement is shared with the treated regions.
0.18
Below 0.80+1.16 pp (2.8× truth, not reported)
Brazil
In structural decline since a local competitor cut prices. Comparable level, opposite trend.
0.43
Below 0.80+1.34 pp (3.2× truth, not reported)
Pre-period movement, each region against its own average

Levels are removed, so the only thing left is shape. The controls that pass ride the same weekly demand as the treated regions. The two that fail do not, and no amount of arithmetic downstream repairs that.

Japan · failsBrazil · fails4 controls · passtreatedApr 6rollout day

The refusal

What a bad control would have told you

Both of these look like reasonable comparison groups. Same product, same metric, same window, similar levels. Both produce a number that is several times the truth, and nothing in a dashboard would have flagged either one.

JapanNo estimate issued

Separate marketing calendar. Almost none of its week-to-week movement is shared with the treated regions.

Pre-trend correlation
0.18
threshold is 0.80
The number it would give
+1.16 pp
2.8× the true +0.42 pp

This is not a noisy estimate of the right thing. Subtracting a group that was already moving differently produces a different quantity wearing the same units, and it is confidently wrong in a direction you cannot predict from the number alone.

BrazilNo estimate issued

In structural decline since a local competitor cut prices. Comparable level, opposite trend.

Pre-trend correlation
0.43
threshold is 0.80
The number it would give
+1.34 pp
3.2× the true +0.42 pp

This is not a noisy estimate of the right thing. Subtracting a group that was already moving differently produces a different quantity wearing the same units, and it is confidently wrong in a direction you cannot predict from the number alone.

The second check

Run the same estimator on a week nothing shipped

The estimator is refitted on the pre-period alone with a fake rollout date halfway through it. Nothing shipped then, so it should return roughly zero. If it returned a large number, the design would be reading structure rather than the change, and the real estimate would be worth nothing.

Placebo rollout, day 45
+0.02 pp
on a window where nothing changed
Real rollout
+0.41 pp
19× the placebo
Truth injected
+0.42 pp
ground truth in this data

4 control regions, each taken on its own, land between +0.40 pp and +0.42 pp. They were not fit together and they agree. That agreement, the placebo above, and the pre-trend check are the whole case for the number. None of them is a p-value, and all of them are things you can look at.