Quantitative Methods
Sampling Error and How Wrong a Small Sample Can Be
Learn how much a number can move for no reason at all, so you stop promoting the manager whose store simply had a small quarter.
- Intermediate
- 11 min total
- 13 chapters
What decision this helps you make: Whether a gap between two stores, reps, cohorts, or variants is large enough to act on — and what minimum event count you require before you rank anything.
- Related calculator: Newsvendor Order Quantity Calculator
What this topic is
Sampling error is the difference between what you measured and what is true, caused by nothing except the fact that you measured part of something rather than all of it. It has a size you can calculate before you collect a single row, and that size depends almost entirely on how many events you observe — not on how large the population is, and not on how carefully you gathered the data.
Why it matters
Almost every operating number a business looks at is a sample: this month's conversion rate, this quarter's close rate by rep, the pilot results from four stores, the 60 responses to the customer survey. Ranked lists built from those numbers put small units at both the top and the bottom, and businesses then spend real money copying the top and fixing the bottom. Both actions are frequently responses to arithmetic rather than to performance.
Who should learn it
Operators who compare units against each other — multi-location owners, sales leaders ranking reps, marketers reading campaign tests, and anyone about to present a leaderboard.
What you will understand
- How to compute the margin of error on a rate before you interpret the rate
- Why the number of events, not the number of rows, sets your precision
- Why the best and worst performers in any ranking tend to be the smallest units
- What minimum event count to require before a difference goes on a slide
Prerequisites
Common misconception
"We have 40,000 sessions a month, so our data is big enough." The 40,000 is not what sets your precision — the roughly 840 conversions inside it are. And the moment you split those sessions by channel, device, region, and week to find the interesting story, each cell holds a few dozen events and carries a margin of error wider than any difference you are about to report. Sample size is not a property of the dataset. It is a property of the specific comparison you are making, and it collapses every time you add a filter.