Quantitative Methods
Synthetic Control and the Constructed Counterfactual
When you changed something in exactly one market and there is no comparable market to hold up against it, build the comparison out of a weighted blend of everywhere else — then find out how often a market that got nothing produces the same result.
- Expert
- 13 min total
- 14 chapters
What decision this helps you make: Whether a single-market test — a campaign, a price change, a service launch in one city — produced a real effect large enough to justify the national rollout, or a gap indistinguishable from what untreated markets produce by themselves.
- Related case study: A DTC Brand That Grew Into a Cash Crunch
- Related data & research: How to Run an Experiment That Actually Decides Something
What this topic is
Synthetic control builds a counterfactual for one treated unit out of a weighted average of untreated units. The weights are chosen so that the blend tracks the treated unit closely over a long pre-treatment period. After treatment, the gap between the treated unit and its synthetic version is the estimated effect. Inference comes not from a formula but from permutation: run the same procedure pretending each untreated unit was treated, and see how unusual the real gap is against that distribution.
Why it matters
A great many business tests are run on exactly one unit. One metro gets the brand campaign. One country gets the new pricing. One warehouse gets the new picking system. There is no control group with sample size, no randomization, and usually no single comparable market — Phoenix is not Tucson. Synthetic control is the standard method for this shape of problem, and it is the method most likely to produce a persuasive chart attached to a result that would not survive a permutation test.
Who should learn it
Marketing teams running geo tests, operators piloting in a single site or region, pricing teams rolling out by country, and anyone about to fund a national programme on the strength of one market's performance.
What you will understand
- How donor weights are chosen and why they are constrained to be non-negative and sum to one
- What pre-treatment fit quality tells you, and when a bad fit means you must abandon the design
- How permutation inference works, and why your donor pool sets a floor on the smallest p-value you can achieve
- Which conditions must hold — no anticipation, no spillover to donors, no other shock to the treated unit
Prerequisites
Common misconception
"The synthetic version tracks our market perfectly for two years and then diverges after the launch, so the campaign worked." A close pre-treatment fit is a necessary condition, not evidence of an effect. The question is how large the post-treatment gap is relative to the gaps that appear when you run the identical procedure on markets that received nothing at all. In the worked example the campaign market shows a 4.7% lift and a beautiful chart — and three of thirty-eight untreated markets produce a larger relative divergence over the same weeks, giving a permutation p-value near 0.10. The chart is real. The inference it invites is not.