I sized a checkout test for a client and the answer was thirty months. Most shops cannot test their checkout at all.
A 5% lift on checkout completion needs about 15,000 sessions per variant. At a thousand checkout starts a month that is two and a half years, which is longer than the checkout will exist.
- Detecting a 5% relative lift on a 29.78% completion rate needs 15,014 checkout starts per variant.
- At 1,000 checkout starts a month that test runs for 30 months. At 20,000 it runs for six weeks.
- Baymard puts average documented cart abandonment at 70.22%, but 42% of those people were only browsing.
- The largest fixable reason is extra costs appearing late, which accounts for roughly 23% of all abandoners.
- Below about 20,000 checkout starts a month, apply the documented fixes and stop pretending the test was a test.
A client asked me last month to plan a checkout test. Good business, sensible people, a checkout that genuinely needed work. I opened the sample size sheet while they were still talking, and by the time they had finished describing the redesign I had a number on my screen that I did not want to read out loud, because it was the kind of number that ends a project rather than starting one.
Thirty months.
I read it out anyway. The room went quiet. Their head of ecommerce said “that cannot be right, we run tests all the time”, and he was correct on the second half of that sentence.
Here is the useful part first, so nobody has to read three thousand words for it. You should size a checkout test before you scope it, never after. Always convert the sample into a date using your own traffic, because the sample alone means nothing. And do not run a test you cannot finish, because an unfinished test does not give you a smaller answer, it gives you a wrong one.
What follows is how I got to thirty months, and the thing I had been quietly getting wrong for about two years.
The number nobody sizes against
Baymard Institute keeps a meta analysis of cart abandonment across 50 published studies. The average documented rate is 70.22%, last updated on 22 September 2025. The individual studies in it range from 55% at the low end, from Forrester in 2010, up to 84.27% from SaleCycle in 2020.
So roughly three in ten carts complete. Call it 29.78%. That is the number to size against.
That is the baseline every checkout test is fighting, and it is a comfortable one. Rates near 30% are statistically friendly, because the variance of a proportion peaks at 50% and falls away slowly on either side. A 3% conversion rate is far harder to test than a 30% one.
None of which helps, because of how the arithmetic actually works.
Where thirty months comes from
Sample size scales with the inverse square of the effect you want to catch. Halve the effect and the sample roughly quadruples. That single sentence is the whole article, and I still had to watch it happen in a spreadsheet before I believed how brutal it is.
At a 29.78% baseline, two sided, 95% confidence and 80% power, here is what the sample looks like and how long each business would spend collecting it:
| Relative lift | Absolute change | Sample per variant | 1,000 starts a month | 5,000 | 20,000 | 100,000 |
|---|---|---|---|---|---|---|
| 3% | 0.89 points | 41,478 | 83 months | 17 months | 4.1 months | 25 days |
| 5% | 1.49 points | 15,014 | 30 months | 6 months | 1.5 months | 9 days |
| 10% | 2.98 points | 3,803 | 7.6 months | 1.5 months | 11 days | 2 days |
| 20% | 5.96 points | 974 | 1.9 months | 12 days | 3 days | Under a day |
My client does a bit over a thousand checkout starts a month. They wanted to know whether a rebuilt checkout was better. Better by how much was the question nobody had asked. We put a realistic 5% on it. The answer came back as thirty months.
Thirty months is not a long test. It is not a test at all. Over that window the checkout gets rebuilt twice, the traffic mix drifts, the payment provider changes something without telling anyone, and in month four somebody runs a promotion that moves the baseline further than the treatment ever could. You would not be measuring your change. You would be measuring two and a half years of everything else, with your change somewhere inside it.
What I had been getting wrong
For about two years I gave the same advice to small shops that I gave to large ones. Test the checkout, one change at a time, read it in six weeks. It sounded responsible. It is the advice you will find in every article written on this subject, including several I have recommended.
It was wrong for anyone under roughly 20,000 checkout starts a month. I did not work that out from reading. I worked it out because a founder asked me a simple question I could not answer. She said “how long, and how sure will we be at the end”. I had the second half of that answer ready and had never once, in two years of recommending checkout tests to people, computed the first half in front of a client.
Twenty thousand a month is where a 5% test lands inside six weeks. That is my line now. Below it the honest answer to “should we test this” is usually no, and saying no is a strange thing to sell, which is probably why so few people in this trade say it.
An aside that has nothing to do with checkout. This arithmetic is also why the tests that get reported as wins at conferences almost all come from enormous sites, and why the effects they report are almost all tiny. Big traffic makes small effects visible. Everyone else is looking at the same small effects through a dirtier window and calling them noise, which is exactly what they are at that sample size. A speaker once told me after a talk that his 0.4% lift had taken eleven weeks at four million sessions. Nobody in the audience had asked. Anyway. Back to the checkout.
The prize is smaller than 70% too
There is a second thing that sizing exposes. The 70.22% headline makes checkout look like the richest ground on the site, and it is not quite that rich, because most of those people were never buying anything.
In Baymard’s breakdown, 42% of abandonments come from people who were only browsing. That is not a leak in your funnel, it is the ordinary behaviour of a species that uses a shopping cart as a wish list, a price comparison tool and a bookmark, and no checkout redesign in the history of the web has ever converted somebody who came to look at a jacket they cannot afford. That is the internet. It is not a problem to be solved.
The remaining 58% left for a reason, and the reasons are ranked. Extra costs being too high sits at 40% of that group. Delivery being too slow is 20%. Not trusting the site with a card is 19%. A forced account is 18%. A long or complicated checkout is 17%. People pick more than one reason, so these do not add to a hundred. The ranking matters more than the percentages do.
Multiply that back out and the picture changes shape completely. Hidden costs account for roughly 23% of every abandoner on the site. A forced account is about 10%. The checkout flow itself, the thing we all keep redesigning and testing and arguing about in meetings, is also about 10%.
The biggest single fixable problem in most checkouts is a price the customer did not see coming. It is not a layout problem. No amount of button testing will touch it. Nobody tests the thing at the top of that list, and I include myself in that, because for years I was testing the thing at position five.
Method, so you can rerun it
Baseline completion of 29.78% is the complement of Baymard’s documented 70.22% average. Sample per variant uses the two proportion formula, two sided, alpha 0.05 and power 0.80, with pooled variance in the null term. Duration assumes traffic splits evenly between two variants and that every checkout start is eligible. It ignores novelty effects, seasonality and the fact that almost nobody runs a clean 50-50 for the whole window, all of which make the real numbers worse rather than better. My first version of this table used sessions rather than checkout starts and was wrong by a factor of about six, which is a good reminder to write down which population you are counting before you count it.
Count the right population, or none of this holds
One more thing, and it is the one that has bitten me hardest.
Every number in that table counts checkout starts. Not sessions, not carts, not users. A checkout start is a session that reaches the first checkout step.
The distinction is not pedantry. My first version of this table used sessions and produced numbers about six times too optimistic, which I only noticed because a colleague looked at it and said “your denominators do not match your rates”. She was right. It took me most of an afternoon to redo it properly, and I have not trusted a table of mine since without checking which population each column is counting.
Sessions include everyone who bounced off a category page. Carts include everyone who saved something for later. Neither group can complete a checkout, so putting them in the denominator dilutes the effect and inflates your confidence at the same time.
Count it at the same point in both variants. Write the population down before you count it. Then check the counts every week, because a sample ratio that drifts is the most common way a checkout test quietly stops meaning anything at all.
What to do below the line
If you are under 20,000 checkout starts a month, here is what I now say.
Ship the documented fixes without testing them. Show the full price early, including shipping and tax. Offer guest checkout. Remove the fields nobody reads. These are not hypotheses. They are the top of a ranked list of reasons people actually gave for leaving, and you are not going to out-measure a survey of that size with a thousand sessions a month.
Watch a long window, honestly labelled. Twelve weeks before and twelve weeks after, with the promotions and the seasonality written down next to the numbers. It is not a controlled experiment and you must not call it one. It is still better than nothing. Given that a genuine experiment would take you thirty months, a twelve week window with the confounders written down honestly beside the numbers is the most rigorous thing available to you, and pretending otherwise helps nobody. It will catch a disaster. That is most of what you need.
Test the big things only. A 20% lift needs 974 per variant, which lands inside two months at a thousand starts a month. If a change is not plausibly worth 20%, it is not worth your only test slot. You get about four of those slots a year.
Fix the errors first. Seventeen percent of the fixable group left because the site broke. That one is a bug, not a hypothesis. It needs a tracker and somebody who checks the payment step on a real phone, on a bad connection, on a Sunday evening when the traffic actually arrives.
Spend the rest of the budget upstream. More checkout starts still work. If completion is stuck at 30% and you cannot move it measurably, the only lever left is the size of the number you are taking 30% of, and that lever does not need statistical significance to pull. The same arithmetic that makes a checkout test impossible at low volume is what makes traffic worth more than optimisation there. That trade is uncomfortable to say out loud when optimisation is the thing you sell.
What I have not settled
The twenty thousand line is a single number standing in for a messy reality, and I know it. It assumes a 5% target, which some businesses would find far too modest and others would find fantastical. It assumes you get one clean test at a time, and most teams are running three overlapping changes and a redesign.
I keep thinking about what happens between 5,000 and 20,000 starts a month, which is where a large share of real shops actually sit. Six months for one answer is absurd and it is not infinite either. There is probably a sensible middle path involving sequential designs, a much larger minimum effect and a written rule about when to give up. I have not built it yet.
The other thing I still do not know is how many teams have run a checkout test that could never have detected anything, got a flat result, and concluded that checkout does not matter. My guess is that it is a lot of them, and that this is the quiet cost of the advice I was giving.
Questions people ask about this
What counts as a checkout start?
A session that reaches the first checkout step, not a session that adds to cart and leaves the product page. Count it at the same point in both variants, because a mismatch here quietly breaks the comparison before any statistics are involved.
Can I test on cart sessions instead, to get more traffic?
You can, and your effect size gets smaller because you have diluted the population with people who were never going to reach payment. Bigger sample, smaller signal, and usually a longer test rather than a shorter one.
So small shops should never test?
They should not run underpowered tests and then act on the result. They can still ship the documented fixes, watch a long before and after window, and call it what it is, which is a change rather than an experiment.
Why does a 5% lift need so much more traffic than a 20% lift?
Sample size scales with the inverse square of the effect. Halving the effect you want to detect roughly quadruples the sample. A 3% lift needs 41,478 per variant, a 20% lift needs 974.
Is 29.78% a realistic completion rate for my store?
It is the complement of Baymard's documented 70.22% average, so it is a starting point rather than your number. Put your own rate into the calculator, because the sample requirement moves with it.
- Baymard Institute. Cart abandonment rate statistics, a meta analysis of 50 studies, average 70.22%, updated 22 September 2025. Checked 12 August 2026.
- Kohavi, R. and Longbotham, R. Online Controlled Experiments and A/B Testing. Encyclopedia of Machine Learning and Data Mining, 2015. Checked 12 August 2026.
Size the test first. The tool on our homepage gives you the sample per variant and how long your traffic needs to produce it, and it tells you when the answer is that the test cannot work.
Open the sample size tool → See what we do