Quantitative Methods
Holdouts and Geo Experiments When You Cannot Randomize Users
When you cannot assign individual users (television, radio, brand, anything with spillover), you randomize markets or hold out a slice of customers. Both work, and both are far blunter instruments than anyone expects.
- Advanced
- 12 min total
- 14 chapters
What decision this helps you make: How to measure a channel you cannot randomize at the user level: how many markets, how long, how large a holdout, and whether the design can resolve the effect you are arguing about at all.
- Related calculator: Newsvendor Order Quantity Calculator
What this topic is
A geo experiment randomizes geographic markets rather than people: some markets get the campaign, some do not, and the difference between them estimates the effect. A holdout excludes a random slice of customers from a treatment (a campaign, a CRM programme, a discount) and uses them as a control. Both are answers to the same problem: the treatment cannot be delivered to an individual in isolation, or its effects leak between individuals, so the unit of randomization has to be something larger than a person.
Why it matters
Most marketing spend is measured by attribution models, which assign credit to touchpoints using rules, not randomization. Those models cannot tell you what would have happened without the spend, which is the only question that matters when you are deciding whether to keep spending. A holdout or a geo test can, but only if it is designed with enough units, enough time, and enough variance reduction to detect the size of effect you are arguing about, and the honest arithmetic is sobering.
Who should learn it
Marketing leaders, growth teams, and finance partners deciding whether a channel earns its budget, plus anyone measuring a treatment that reaches households, stores, regions, or whole markets rather than individuals.
What you will understand
- Why a market-level test with forty markets is powered for effects several times larger than most campaigns produce
- How pre-period adjustment turns an unusable geo design into a workable one, with the arithmetic
- Why a 5% holdout costs you more than twice as much precision as a balanced split of the same population
- What assumptions a geo test rests on (no spillover between markets, no coincident local shocks), and how each one fails
Prerequisites
Common misconception
"We ran a geo test and found no significant lift, so the channel does not work." A forty-market test over eight weeks typically cannot detect anything under about a 6% revenue lift, and most channels do not claim to move total revenue by 6%. A null from that design is entirely compatible with the channel delivering a healthy return. The variance across markets is enormous, the number of units is tiny, and the resulting instrument is coarse, which is a fact about the test, not about the channel.