Experiments & Data

A/B Test Sample Size Calculator

The number that decides whether a test is worth running at all. Most teams pick a two-week window and read whatever they see; this works the other way round — you say what size of improvement would actually change a decision, and the maths tells you how many people it takes to see it, and how many weeks of your traffic that is. Very often the honest answer is that you cannot detect the lift you are hoping for, and knowing that before you start is worth more than the test.

Inputs

  • Baseline conversion rate
  • Minimum detectable lift (relative) — The smallest improvement worth detecting — 15% here means 4.0% → 4.6%, not 4% → 19%.
  • Statistical power — Your chance of spotting the lift if it is really there. 80% is the convention.
  • Significance level — Your tolerance for calling a win that is not real. 5% is the convention.
  • Weekly traffic into the test — Total visitors across both arms per week.

How to use this calculator

  1. Enter the conversion rate you get today on the page or flow you are about to change. Use your own analytics for the same audience and the same window. A rate borrowed from a benchmark article will size the wrong test.
  2. Set the minimum detectable lift as a relative improvement: 15% means taking a 4.0% rate to 4.6%, not to 19%. Choose it by asking what improvement would be big enough to change what you do. If a 2% lift would not change anything, do not ask the test to find one, because you would be paying weeks of traffic for a number you will ignore.
  3. Leave power at 80% and significance at 5% unless you have a reason. Power is your chance of catching a lift that is genuinely there; significance is your tolerance for celebrating one that is not. Raising either costs sample, and the cost climbs steeply.
  4. Enter the weekly traffic that will actually enter the test, total across both arms, not your whole site. If only visitors who reach the checkout see the change, that is the number.
  5. Read the weeks figure, not the sample size. That is the real decision. Anything past about eight weeks will not survive contact with a roadmap, and the last row tells you what four weeks of your traffic could realistically detect instead.
  6. Fix your stopping point now, before you start: one look, at the total sample shown. Checking every morning and stopping the day it turns green is the single most common way a test lies to a team.

What each term means

Baseline conversion rate
The rate you convert at today, the thing the variant has to beat. Everything on this page scales off it.
Minimum detectable lift (MDE)
The smallest improvement the test is built to see, expressed relative to the baseline. Smaller lifts are not undetectable, they are just far more expensive to detect: halving the lift roughly quadruples the sample.
Statistical power
Your chance of finding the lift if it is really there. At 80% power, one real winner in five looks like nothing and gets thrown away.
Significance level (α)
How often you are willing to call a win that was only chance. 5% means one in twenty null results looks like a winner.
Sample per arm
Visitors needed in the control and in the variant each, not between them. Double it for the total.
Peeking
Looking at results repeatedly and stopping when they look good. It quietly turns a 5% false-positive rate into something closer to 20–30%, and no calculator can correct for it after the fact.

Educational disclaimer: Outputs are simplified educational estimates built from the numbers you enter — they are not financial, legal, tax, or investment advice, and real decisions deserve verified figures and qualified professionals.

Quantitative Methods

More Experiments & Data tools