Quantitative Methods

Statistical Power and the Minimum Detectable Effect

Find out what the sample you actually have could ever have detected — and why an underpowered test that does come back significant is more dangerous than one that does not.

  • Advanced
  • 11 min total
  • 14 chapters

What decision this helps you make: Whether a "no significant difference" result licenses you to ship the cheaper, simpler, or lazier option — and what minimum detectable effect you commit to before spending the traffic.

What this topic is

Power is the probability that a test detects an effect of a given size, assuming that effect is real. The minimum detectable effect is the same idea inverted: given the sample you have, the smallest true effect the test would find most of the time. Every experiment has both, whether or not anyone computed them, and they are fixed by the design rather than discovered in the results.

Why it matters

Power decides what your evidence can mean. A low-powered test that finds nothing tells you almost nothing, and a low-powered test that finds something tells you something inflated — the estimate that clears the significance bar in a small sample has to be large, so the ones you get to see are systematically bigger than the truth. That is why a year of two-percent wins produces a flat revenue line, and why so many "proven" improvements evaporate when they are rolled out.

Who should learn it

Anyone who reads experiment results and has to decide whether to act on them, and anyone deciding how much traffic, budget, or calendar to commit to finding out.

What you will understand

  • How to compute the minimum detectable effect from the sample you have, in one line of arithmetic
  • Why a significant result from a low-powered test overstates the true effect, and by roughly how much
  • Why post-hoc power calculated from your own result is not informative and should never be reported
  • How to test for equivalence when the decision you want to make is "these are close enough"

Prerequisites

Common misconception

"We were underpowered, but we found a significant effect anyway, so the power problem does not apply." It applies more. Low power does not just make effects hard to find; it distorts the ones you find. To clear the threshold on a small sample, the observed difference has to be large, so conditional on being significant, the estimate is biased upward — at around 17% power, by a factor of about two and a half. The result being significant is precisely what should make you suspicious of its size.