Quantitative Methods

Switchback Tests and Interference Between Treated Units

In a marketplace, the treated arm wins partly by taking supply from the control arm. Switchback designs randomize time instead of units to measure what the change does to the whole market rather than what it takes from the other half.

  • Expert
  • 12 min total
  • 14 chapters

What decision this helps you make: Whether your last marketplace or network A/B result measured a real improvement or a transfer between arms — and whether to pay roughly eight times the precision for an unbiased answer.

What this topic is

Interference means one unit's treatment changes another unit's outcome. It violates the assumption underneath every standard experiment, and it is the normal condition in marketplaces, social products, ad auctions, and anywhere capacity is shared. A switchback experiment responds by randomizing time: the whole market runs on treatment for a block, then control for a block, then treatment again, with the assignment drawn at random for each block. Every participant in a block sees the same system, so nobody is competing against a differently treated counterpart.

Why it matters

A standard A/B test in a marketplace does not merely have noisy results — it has systematically wrong ones, and the direction of the error is usually flattering. If treated orders get preferential access to a shared courier pool, they get faster and control orders get slower, and the difference between them exaggerates what the change would deliver if rolled out to everyone. Companies have rolled out algorithm changes on the strength of measured improvements that were largely a transfer from the control group, and then watched the market-level metric fail to move.

Who should learn it

Anyone running experiments in a two-sided marketplace, a delivery or logistics network, an advertising system with shared budgets, a social product, or any business where inventory or capacity is pooled across customers.

What you will understand

  • What the no-interference assumption is, where it fails commercially, and which way the bias runs
  • How a switchback design removes cross-arm interference and what it costs in statistical precision
  • Why the unit of analysis in a switchback is the time block, not the order — and how badly wrong the standard errors are otherwise
  • Why block length barely affects precision but decides how much carryover contaminates your estimate

Prerequisites

Common misconception

"We randomized properly, so the estimate is unbiased." Randomization guarantees the two groups are comparable. It does not guarantee that each group's outcome would be the same in a world where everyone got the same treatment — and that is the quantity you need before rolling out. In a shared-capacity system the two arms interact, so the measured difference contains both the effect of the change and whatever the treated arm took from the control arm. Randomization does nothing about the second part.