How long an ecom A/B test needs to run is set by 2 numbers you know before the test exists: weekly orders through the tested experience, and the smallest lift you'd actually act on. Everything else cancels out of the math.. conversion rate, traffic, even which metric moved.. and the whole calculation compresses to weeks ≈ 31 ÷ (weekly orders × lift²).
Most teams never run that number. Tests get a default 2 weeks, the order counts get read like a scoreboard, and week-1 Slack fills with screenshots of the variant "winning". So before the formula, here is what a typical 2-week read actually contains:
The reason intuition fails here is that order counts are noisy and the noise shrinks slowly, with the square root of volume. A 5-order gap feels like signal because in every other part of the business 5 orders is real money, but as a statistical read it's indistinguishable from shuffling the same customers into different buckets.
Run an A/A test once
The cheapest way to make that lesson stick org-wide isn't a stats lecture, it's an A/A test: split traffic 50/50 between two identical versions of a page and report the results like a normal test. Since the true difference is 0% by construction, everything the report shows is noise and tooling, and it buys 3 things at once. It calibrates everyone on noise, because the gap between identical pages at your volume becomes the bar every future gap gets compared against. It shows metric speed: high-sample rates settle toward 0% within days while order-based metrics swing for weeks, which is why upper-funnel metrics carry early reads. And it validates the setup: the one line that should never hold a gap is the user split itself, so a 50/50 allocation sitting outside its noise band is an instrumentation problem to fix before any real test is trusted.
Here is what one looks like:
| Orders per side | Typical CVR "difference" | 1 test in 20 shows over |
|---|---|---|
| 100 | ~10% | ~28% |
| 200 | ~7% | ~20% |
| 500 | ~4% | ~12% |
| 1,500 | ~2.5% | ~7% |
| 5,000 | ~1.3% | ~4% |
I would suggest running one sitewide for 4-6 weeks once a year or after any tooling change. It costs nothing.. both arms are the same site.. and it retires the early-read culture on its own: after everyone has watched identical pages "beat" each other by 10% in week 1, day-3 winner screenshots stop happening. It also converts "what is noise" from an argument into a screenshot everyone remembers.
From there, the only version of the run-length question with a useful answer is the one asked before launch: given our volume, how long until a lift we care about would be readable at all? That's the table. Find your row:
| Weekly orders (in test) | Detect a 20% lift | 10% lift | 5% lift |
|---|---|---|---|
| 250 | 3.4 wk | 13 wk | 50 wk |
| 500 | 2 wk | 6.4 wk | 25 wk |
| 1,000 | 2 wk | 3.2 wk | 13 wk |
| 2,500 | 2 wk | 2 wk | 5 wk |
| 5,000 | 2 wk | 2 wk | 2.5 wk |
| 10,000 | 2 wk | 2 wk | 2 wk |
What the red rows actually mean
If you're doing 250 orders a week through the tested page, a 5% conversion test needs about a year, and no ecom site holds still for a year.. cookies churn, promos shift, the site around the test keeps changing. So the honest conclusion isn't "we can't test", it's that your testing program has a minimum interesting effect size. At low volume, the only tests worth running are big swings: full page rebuilds, offer structure, template changes.. things that could plausibly move results 10-20%. Button colors and copy tweaks are tests for brands doing 5,000+ weekly orders, and running them below that produces results nobody can stand behind either way.
The other discipline the table buys you is a fixed stop date. Checking daily and stopping the moment significance appears roughly triples your false-winner rate, because noise crosses the threshold repeatedly on the way to nowhere. Compute the length, hold the length, and if the length is unacceptable, change the test rather than the calendar.
And it's worth saying plainly: statistics exists to improve a decision. If no decision changes with the result, or the readable version of the test costs a quarter of roadmap, I would suggest not running it and spending the traffic on a bigger question.
The most common way this goes wrong in practice is the test that won't die. We recently watched a paid-traffic landing page test kept alive for weeks.. first waiting for last-click orders to reach significance, then extended again to wait for attribution-modeled orders to build sample.. when both were always going to be too small to read. Meanwhile the upper-funnel metrics were already statistically significant and clearly positive: product-page view rate and add-to-cart up, purchases not moving the wrong way. That, plus the direction the team already wanted the experience to go, was the decision. The extra weeks of waiting bought nothing except delay on the next test.
Reading an inconclusive result
Even sized correctly, plenty of tests end without significance. That's information, not failure, and it has a reading procedure:
Hypothesis confidence. Why did you believe it before the test? A mechanism you still believe survives an inconclusive read; a coin-flip idea doesn't.
Behavioral inputs. Did add-to-cart, PDP click-through or units per transaction move in the supporting direction? Inputs are higher-volume than orders, so they read earlier and steady a noisy topline.
Cost of being wrong. Shipping a false copy winner costs nothing. Shipping a false discount winner trains customers to wait for codes.. asymmetric costs get the strict read.
Strong hypothesis, supporting inputs and a cheap reversal: shipping on an inconclusive read is reasonable, and say so honestly. Weak hypothesis or costly reversal: kill it and spend the traffic on a bigger swing. The one non-option is calling it a win.
What to do with it
Run an A/A test once (and after any tooling change): 4-6 weeks sitewide, reported like a real test. It validates allocation, and it calibrates the whole team on what noise looks like at your volume.
Before building any test, run the formula: 31 ÷ (weekly orders × lift²). If the answer is longer than ~a quarter, redesign the test around a bigger swing rather than extending the calendar.
At low volume, test fewer, bolder variants. One strong alternative against control beats 4 tweaks splitting the same traffic, since each added arm stretches every arm's timeline.
Pre-commit the decision each outcome triggers.. ship, kill, iterate.. before data arrives, and hold the stop date. If you can't name a decision the test would change, that's the answer about whether to run it.
Read inconclusive results with the 3 checks above, and be strictest where a false winner has a training cost. A discount test shipped on noise doesn't just misread the past, it teaches customers to wait for codes.
Happy to walk through the math for your volume and test queue.
Common questions
What is an A/A test and why run one?
An A/A test splits traffic 50/50 between two identical versions of a page, so the true difference is 0% by construction and everything the report shows is noise and tooling. It answers 3 questions no A/B test can: whether your allocation is actually random (more than ~1% uneven on a 50/50 split means a setup issue), how big the gaps between identical experiences are at your volume, and how long each metric takes to settle.. top-of-funnel rates typically converge within 1-2 weeks while conversion metrics swing for 4+ even sitewide. The result doubles as an internal calibration artifact: every future "the variant is up 8%" gets compared against what identical pages showed, which is usually the end of early winner declarations.
Can't I just run the test longer to detect a smaller lift?
In principle yes, and the formula tells you exactly how long. In practice, past ~10-12 weeks the test starts measuring a moving target: cookies churn so returning visitors get re-assigned, seasonality and promo mix shift who's arriving, and the site around the tested element keeps changing. A test that needs 6 months isn't a test, it's a queue blocking bolder ideas that could have read in 3 weeks. The workable range for most ecom tests is 2-8 weeks, which is why the effect size, not the calendar, is the lever worth adjusting.
Do I always need 95% significance?
No. Significance is a decision threshold, not a law of nature, and the right threshold depends on the cost of being wrong. A reversible, zero-cost change (copy, layout) can reasonably ship at 80-90% confidence, or even on a well-reasoned inconclusive result. An asymmetric-cost change (discounts, pricing, anything that trains customer behavior) deserves the strict read. What's not negotiable is choosing the threshold before looking at results, because a threshold chosen afterward always ends up being whatever the data happened to clear.
Should I test revenue per visitor instead of conversion rate?
Usually not as the primary metric. Revenue per visitor carries all of AOV's variance on top of conversion's, so it typically needs several times the sample.. and a single large order landing in one arm can swing the read for weeks. The decomposition works better: conversion rate as the primary metric, with AOV and units per transaction tracked as supporting reads. If the hypothesis is genuinely about basket size (bundles, thresholds), test on AOV directly with outliers capped and expect a longer run than the table shows.
See how we design experimentation programs
See how it works →Sample sizes use a two-proportion test at 95% confidence (two-sided) and 80% power with a 50/50 split, run in whole weekly cycles with a 2-week floor. The weeks formula is ~31 ÷ (weekly orders × relative lift²); baseline conversion rate cancels out, so it holds across CVRs up to ~10%. Identical-page gaps are the sampling distribution of the difference between two equal conversion rates (SE of the relative gap ≈ √(2/n) at n orders per side; "typical" is the median absolute gap, "1 in 20" the 95th percentile). The 75-vs-70 example is illustrative.