Case Study
The Test That Was Winning on Tuesday
What happened
A store running about 3,000 checkout visits a day tested a new checkout, expecting roughly 90 orders a day at $70 each. On day six someone opened the dashboard, saw the new version 9% ahead with a p-value of 0.04, and shipped it, which penciled out to about $567 a day. Over the next three months orders per visit did not move, up or down, and analysts looked for the money in weekly and then monthly reports without finding it. They reran it properly with the end date fixed in advance, and afterwards worked out that the original test, at its size, could never have reliably detected a lift as small as the one it had reported.
Anonymized composite: a mid-sized direct-to-consumer store that ships a checkout redesign six days into a fourteen-day test. The behavior is documented, not invented: Berman, Pekelis, Scott and Van den Bulte studied 2,101 commercial experiments run on the Optimizely platform in 2014 and found that about 73% of experimenters stopped the moment a positive result crossed 90% confidence. The false-alarm rates in the chart are simulated for this case study by the method stated in the chart note. All dollar figures are illustrative only.
- Illustrative composite — not a real company
- E-commerce
- Direct-to-consumer store
- Moderate risk
- Turnaround
- Beginner
The case, start to finish
Day 6 was not a decision point. It only became one because nobody had written down when the test would end.
The store, the plan, and the moment
This is an anonymized composite of a mid-sized direct-to-consumer store, with illustrative dollar figures. The behavior it describes is documented: in a study of 2,101 commercial experiments run on the Optimizely platform in 2014, about 73% of experimenters stopped the moment a positive result crossed 90% confidence.
The store gets about 3,000 checkout visits a day. Three in a hundred finish an order, at a $70 average order value, so roughly 90 orders and $6,300 a day. A designer builds a new checkout, and the team plans a fourteen-day test: 1,500 visits a day into each version, 21,000 per version at the end.
On day 6 somebody opens the dashboard. The new checkout is 9% ahead and the tool reports p = 0.04. Nine percent of 90 orders is 8.1 more orders a day, about $567, roughly $207,000 a year, and eight more days of waiting looks like leaving $4,500 on the table. They ship it and end the test. Nothing about that reasoning is lazy, which is exactly why the mistake is so common.
What a p-value is actually promising
A p-value answers one narrow question: supposing the two versions really are identical, how often would chance alone hand you a gap this large or larger? Five percent is the answer for one look, at a moment fixed before you started.
Look every day for two weeks and you are not taking one shot at a 5% risk. You are taking fourteen shots and stopping the instant one lands. Simulated for this case, that turns a 5% false-alarm rate into about 22%; watch a live dashboard hourly and it reaches about 46%. None of that math is new or contested. It is simply invisible from inside a tool that recalculates in real time and shows you a green number the moment there is one, which is why optional stopping is the default behavior rather than a deliberate choice.
The other half, which nobody checked at all
There was a second problem underneath the first, and it was fatal on its own. Months later somebody finally computed what the original test could ever have detected. At 21,000 visits per version, the minimum detectable effect was about 15.5%. The 9% they shipped on was below the floor of what that test was capable of seeing.
That is worth restating, because it inverts the intuition. A 9% reading on an underpowered test is not a small real win, it is an ordinary-sized fluctuation. And detecting a genuine 5% lift on a 3% baseline would have needed roughly 208,000 visits per version, about twenty weeks of traffic. Working that out before launch costs nothing and takes minutes.
The money that was never there
Through weeks three to twelve, orders per visit did not move. Not down, not up. Analysts hunted for $567 a day in weekly reports and then in monthly ones and could not find it. Nobody concluded the test had been wrong; everybody concluded the effect had been masked by traffic mix, a bad month, the weather.
The rerun in month four fixed the rules first: end date written down before launch, dashboard locked, forty-four days, about 66,000 visits per version. The new checkout came out 2% behind, in a range running from roughly 8% worse to 4% better. No measurable difference, and certainly not a 9% gain.
Nobody gets promoted for that result, and it was the most valuable thing the team learned all year. The real cost of the false winner was never the checkout page, which was fine. It was that for nearly half a year everybody believed checkout was solved and stopped looking. In the study behind this case, the expected cost of a false discovery was estimated at a loss of about 1.95% in lift, which sat at the 76th percentile of the lifts actually observed. The bad page is cheap. The good page you stopped searching for is not.
The fix is a decision made on day 0
Write the sample size and end date into the ticket before launch, and hide mid-test significance from the dashboard by default. The second half matters more than it sounds: it removes the temptation instead of asking busy people to resist it, which is the only version of this that survives a quarter with a deadline in it.
There is a legitimate alternative, with one condition. Sequential methods exist precisely so a test can be watched continuously and stopped whenever the answer is clear. But you choose that framework before you look, not after a number you like has appeared, because adopting it afterward is fitting a rule to a result: the original error in better clothes. Those methods also buy their freedom with statistical power, so an underpowered test becomes more underpowered rather than less.
The team's version of the lesson was blunt and worth copying: run fewer tests, and believe more of them.
Timeline
- Week 0 — the plan The store gets about 3,000 checkout visits a day. Three in a hundred finish an order, at an average order value of $70, so that is roughly 90 orders and $6,300 a day. A designer has a new checkout. The team plans a fourteen-day test: 1,500 visits a day into each version, 21,000 per version at the end.
- Day 6 — the peek Someone opens the dashboard. The new checkout is 9% ahead and the tool reports p = 0.04. Nine percent of 90 orders is 8.1 more orders a day, about $567 a day, roughly $207,000 a year. Waiting eight more days now looks like leaving $4,500 on the table. They ship it and end the test.
- Weeks 3–12 — the missing money Orders per visit do not move. Not down, not up. Analysts look for the $567 a day in weekly reports, then in monthly ones, and cannot find it. Nobody concludes the test was wrong; everyone concludes the effect was masked by something: traffic mix, a bad month, the weather.
- Month 4 — the honest rerun They run it again, with the end date written down before launch and the dashboard locked. Forty-four days, about 66,000 visits per version. Result: the new checkout is 2% behind, in a range running from roughly 8% worse to 4% better. Translation: no measurable difference, and certainly not a 9% gain.
- Month 5 — the number that stings Someone computes what the original test could ever have detected. At 21,000 visits per version, the smallest lift it could reliably find was about 15.5%. The 9% they shipped on was below the floor of what that test was able to see. It was never a small win. It was noise the size of a win.
- Year 2 — the rule Sample size and end date are chosen before launch and written into the test ticket. Mid-test significance is hidden from the dashboard by default. The team runs fewer tests and believes more of them, and the checkout work restarts, because for nearly half a year everyone had believed checkout was solved.
You're in the owner's chair
Day 6 of a 14-day test. The new checkout is 9% ahead, the tool says p = 0.04, and Black Friday is three weeks out. Eight more days of the old version looks like $4,500 of lost orders. What do you do?
- Switch to a sequential method built for continuous monitoring, so you can stop as soon as the result is genuinely clear
- Run to day 14 exactly as planned, then decide once
- Ship it — you have significance, and waiting has a price you can calculate
Chance of declaring a winner when the two versions are identical
- Look once, at the planned end (day 14): 5%
- Look once a day for 14 days: 22%
- Watch a live dashboard, hourly: 46%
Simulated for this case study: 60,000 simulated A/A tests in which both versions convert at exactly 3.0%, 1,500 visits per version per day for 14 days, judged by a two-sided z-test at the 5% level, counting a run as a "winner" if it ever crossed that threshold at any check. Raw results 4.9%, 22.2% and 45.7%. The first bar is the number the tool is reporting. The third is the number you are actually running when the dashboard is open.
Business model
An online store with one funnel and no salespeople. Conversion rate is the whole business: traffic costs the same whether visitors buy or not, so a percentage point at checkout is worth more than almost anything else the team can do. That is exactly why the pressure to find a winner is what it is.
Revenue model
Illustrative only: about 3,000 checkout visits a day at a 3.0% conversion rate and $70 average order is roughly 90 orders and $6,300 a day. A genuine 9% lift would be worth about $207,000 a year. That number is why nobody waited eight days, and it is also why nobody checked whether the test was capable of measuring it.
Cost structure
The direct cost of running the test is nearly zero. The real costs are the ones that do not appear on an invoice: eight days of engineering patience, and the far larger cost of what a false winner does next: the team stops looking. Berman and colleagues estimate the expected cost of a false discovery at a loss of about 1.95% in lift, which in their data sits at the 76th percentile of all the lifts actually observed. The bad page is cheap. The good page you stopped searching for is not.
Strategic challenge
A p-value answers exactly one question: if the two versions were truly identical, how often would chance alone produce a gap at least this big? Five percent is the answer for one look, at a moment fixed before you started. Look every day for two weeks and you are no longer taking one shot at a 5% risk. You are taking fourteen, and stopping the instant one of them lands. Simulated below, that turns a 5% false-alarm rate into 22%. Watch a live dashboard hourly and it reaches 46%. Nothing about the math is unknown; it is just invisible from inside a tool that updates in real time.
Key decision
Whether the stopping rule is chosen before the data or after it. Everything else, including the design, the traffic and the statistics package, is downstream of that single question. A test with the end date written on the ticket is an experiment. A test that ends when someone likes the number is a search for a number someone likes.
What worked
The rerun, and only because the rules were fixed first. Forty-four days is a long time to hold a decision open, and the answer it produced was "no difference", an outcome no one gets promoted for, and the most valuable thing the team learned all year, because it freed the checkout back up for real work. The second win was cultural: hiding mid-test significance removed the temptation rather than asking people to resist it, which is the only version of this fix that survives a busy quarter.
What failed
Three things at once, and they compounded. The test was underpowered: 21,000 visits per version could only reliably detect a 15.5% lift, so a 9% reading was inside the noise before anyone peeked. It was then stopped early, on a threshold checked repeatedly, which inflates the false-alarm rate from 5% to about 22%. And the result was never verified after shipping, so the missing $567 a day was explained away for nearly half a year instead of investigated. Any one of these alone is survivable. Together they produced a confident, expensive, entirely fictional win.
Risk factors
Dashboards that display significance continuously, which makes peeking the default rather than a decision; low base conversion rates, where the sample sizes needed for small effects are far larger than intuition suggests; seasonal traffic, which makes a six-day window unrepresentative on top of everything else; incentive structures that reward shipping winners rather than resolving questions; and the quiet one, which is that a false winner ends the search, so the cost is the improvement you never went looking for.
Lesson summary
Decide when the test ends before you start it, and write it down. A p-value assumes one look at one moment fixed in advance; check daily and the 5% chance of a false alarm becomes about 22%, and watch a live dashboard and it becomes about 46%. Then check the other half, which is whether the test is even large enough to see the effect you are hoping for. Here it was not: 21,000 visits per version could only detect a 15.5% lift, and detecting a 5% lift would have taken about 208,000 per version, or roughly twenty weeks. The honest answer is often "this test cannot answer that question," and knowing it before you run costs nothing.
Key data
- 3.0% Baseline checkout conversion (illustrative)
- 21,000 visits per version, 14 days Planned test size
- 15.5% Smallest lift that test could reliably detect
- 9%, at p = 0.04 Lift claimed on day 6
- ~$207,000 Annual revenue that claimed lift implied
- ~208,000 (about 20 weeks) Visits per version needed to detect a 5% lift
- ~75% of 2,101 experiments Effects that were truly null in the Optimizely study
- 40% when stopped at 90% confidence, vs. 33% Share of “winners” that were false, in that study
Sources & basis
The business in this story is a stand-in, not a company you can look up. This case is an illustrative composite: the operator, the people and most of the dollar figures represent a pattern rather than reporting one firm's history. What the list below cites is the other half, the documented industry data and public reporting the composite was assembled from, including any real company whose published figures the case draws on by name. The mechanism and the arithmetic are real even where the business is not.
- Ron Berman, Leonid Pekelis, Aisling Scott and Christophe Van den Bulte, "p-Hacking and False Discovery in A/B Testing" (December 11, 2018) — 2,101 commercial experiments run on the Optimizely platform in 2014; about 73% of experimenters stop just as a positive effect reaches 90% confidence; approximately 75% of effects are truly null; improper optional stopping raises the false discovery rate from 33% to 40%; expected cost of a false discovery estimated at a loss of 1.95% in lift, the 76th percentile of observed lifts View source ↗
- The Berman et al. paper is also the version of record on SSRN, abstract id 3204791, as printed in the footer of the PDF above.
- Ramesh Johari, Leo Pekelis and David J. Walsh, "Always Valid Inference: Bringing Sequential Analysis to A/B Testing" (arXiv:1512.04922) — the statement that standard frequentist inferences are unreliable when sample sizes are chosen by continuous monitoring, and the always-valid p-values and confidence intervals proposed as the remedy View source ↗
- Chart figures are a simulation written for this case study, not a published result: 60,000 simulated A/A tests, 3.0% true conversion in both arms, 1,500 visits per arm per day for 14 days, two-sided z-test at alpha = 0.05, recording whether the threshold was ever crossed under one look, 14 daily looks, and 336 hourly looks. Reproducible in a few lines of Python.
- Sample-size and detectable-effect figures use the standard two-proportion formula at 80% power and alpha = 0.05 two-sided, on a 3.0% baseline: 21,000 per arm detects 15.5% relative lift; 5% relative lift needs about 208,000 per arm.