Geo holdout incrementality for ad spend.

Most of what the platform took credit for would have happened anyway.

Attributed conversions are not incremental conversions, and before/after is confounded by whatever season you launched into. Lift holds out geographies, builds a counterfactual from the markets you did not touch, and reports what the campaign actually caused.

No spam. One email when it is ready to try.

Geo holdout · 20 held-out markets28 days · $640K of mediasynthetic control, pre-period fit 0.97%
Platform attributed
$2.15M
4.4× the truth
Before/after
$1.88M
3.8× the truth
Synthetic control
$448K
0.91× the truth
Truth injected
$490K
known, because we put it there

A geo holdout with a known injected incremental effect. Three numbers describe the same four weeks, and only one of them is right.

The problem

Nobody being paid for the ads has a reason to tell you which ones worked.

Every platform reports the conversions it decided to claim, using its own window, its own identity graph, and its own credit rules. A large share of those buyers were going to buy regardless, and nothing in the report separates the two. The usual fallback is worse: spend went up, revenue went up, therefore it worked. That comparison silently includes seasonality, price changes, other channels, and the fact that media plans are built around the season in the first place.

4.4×
How much larger the platform-attributed number is than the real incremental effect in the scenario on this site, where the incremental effect is injected and therefore known exactly.
Arithmetic from this page’s own data, not a market claim
3.8×
Before/after on the identical data. The campaign launched into a rising season, and the season ends up inside the number.
Same data, computed in your browser
Double-counted
Two platforms claiming the same conversion is the ordinary case, because each runs its own attribution over its own view of the user. Adding up platform-reported conversions overcounts by construction.

The insight

You cannot randomise users any more. You can still randomise cities.

Geography is the last unit of randomisation that survives platform walls, signal loss, and cross-device buying. Hold out a set of markets, then build a weighted blend of them that tracks your treated markets before the campaign starts. After launch, the gap between what your markets did and what the blend did is the effect. The pre-period fit is the whole credibility argument: if the blend does not track before treatment, there is no counterfactual and the number means nothing. Saying so out loud is the product. A measurement vendor that always returns a confident number is not measuring anything.

Method

Synthetic control, following Abadie, Diamond and Hainmueller (2010). Donor weights fit by projected gradient descent constrained to the simplex, so the counterfactual is a convex combination of real markets and cannot extrapolate past them. Credibility reported four ways: pre-period RMSE against the treated level, the post/pre gap ratio, an in-time placebo on a window where nothing ran, and a rank test that refits every donor market as if it had been treated.

How it works

Four steps, no data science team

01
Pick the holdout before you spend

Split markets so the held-out set covers the size, growth and seasonal profile of the treated set. The donor pool is the experiment design; getting it right beforehand is cheaper than arguing about it afterwards.

02
Fit the counterfactual on history you already have

Weekly or daily revenue by market, from your own warehouse. No pixels, no platform API, no identity resolution. The weights are fit only on the pre-period, so nothing about the campaign can leak into the fit.

03
Read the gap, and the fit that earns it

Incremental revenue, incremental ROAS, and the pre-period tracking error side by side. If the fit is poor the readout says the test is inconclusive and why, instead of dressing up a number.

04
Take the report to the budget meeting

One page with the platform number, the before/after number, and the incremental number, plus the placebo tests that back the third one. It is written to survive someone who wants the answer to be the first number.

Who it is for

The person who has to defend the media budget

Teams spending enough on paid media that a wrong ROAS number moves real money, and who have already noticed that their platform-reported revenue exceeds their actual revenue.

Pricing

Single test
$4,000/test
One campaign, one readout
  • Holdout design and power check
  • Synthetic control readout
  • Placebo tests in time and in space
  • Report written for a finance review
Most common
Always on
$3,500/mo
Continuous geo testing across channels
  • Rolling holdout rotation
  • Per-channel incrementality
  • Warehouse sync, no pixel
  • Budget reallocation from measured iROAS
Enterprise
$12,000/mo
Multi-brand, multi-region
  • Self-hosted
  • Custom donor pools
  • SSO and audit log
  • Review with a statistician

Competition

What exists, and what it does not do

WhoWhat they doThe gap
Platform lift studiesConversion lift and ghost-ad tests run inside Meta or Google.The platform designs the test, holds the data, and reports the result on the spend it is selling. Even when the method is sound, it cannot be compared across platforms and it cannot be audited by you.
Measured, Haus, INCRMNTALIncrementality vendors that also run geo and holdout experiments.The category is right and these are the real competitors. The wedge is that the estimator, the fit diagnostics and the placebo tests are visible rather than a black box, and that a bad fit produces a stated refusal instead of a number.
Media mix modellingRegression on aggregate spend and revenue, now open-sourced by Meta and Google.Answers a different question at a coarser resolution, needs years of history, and has no experiment behind it. MMM is where geo test results should be used as priors, not a substitute for running one.
The platform dashboardWhat most teams actually use, because it is free and already there.Overstates by construction, and the overstatement is invisible unless you run a holdout.
How this fails

Geo tests are expensive in a way an A/B test is not: holding out markets means turning off spend that might have worked, and the smallest effect a geo design can resolve is large. If a customer runs one test, gets a number they do not like, and concludes the method is wrong rather than the spend, they churn and tell people. The honest version of this business only works with buyers who already suspect their attribution is inflated, which is a smaller market than every advertiser.

Market

A line item inside paid media budgets, which is where the money already is

A team spending $2M a year on paid media is deciding how to place that money largely on platform-reported numbers. Measurement at $3,500 a month is under 3% of that spend, and the decision it changes is worth multiples of it. Two thousand mid-market advertisers at the always-on tier is $84M ARR, sold to the same buyer who already pays agencies for worse answers.

Request a holdout design

No spam. One email when it is ready to try.

Or just go look at the demo first →