Insights · Experimentation

How long to run an ecom A/B test

Sam Gil·Sep 2026·6 min
Key takeaways
1weeks ≈ 31 ÷ (weekly orders × lift²).. conversion rate and traffic cancel out, so run length is knowable before the test is built
2Two identical pages typically show a ~7% CVR gap at 200 orders per side; run an A/A test once and the whole team learns what noise looks like
3If your row says a quarter or more, the fix is testing bigger changes, not running longer
4Non-significant results are inconclusive, not losses; decide on hypothesis confidence, behavioral inputs and the cost of being wrong

How long an ecom A/B test needs to run is set by 2 numbers you know before the test exists: weekly orders through the tested experience, and the smallest lift you'd actually act on. Everything else cancels out of the math.. conversion rate, traffic, even which metric moved.. and the whole calculation compresses to weeks ≈ 31 ÷ (weekly orders × lift²).

Most teams never run that number. Tests get a default 2 weeks, the order counts get read like a scoreboard, and week-1 Slack fills with screenshots of the variant "winning". So before the formula, here is what a typical 2-week read actually contains:

2 weeks in · equal traffic to both arms
Control
70 orders
Variant · "+7%!"
75 orders
1Gaps this size appear by pure chance roughly 2 times out of 3 (p ≈ 0.67).
2For +7% to be a readable signal you'd need ~3,100 orders per arm.. ~43x this sample.
=Correct read: the same as no difference. Not a win, not a loss, no information yet.
Order counts are noisy, and the noise shrinks with the square root of volume, so small tests produce differences that look decisive and mean nothing. Every "the variant is winning!" Slack message in week 1 is this exhibit.

The reason intuition fails here is that order counts are noisy and the noise shrinks slowly, with the square root of volume. A 5-order gap feels like signal because in every other part of the business 5 orders is real money, but as a statistical read it's indistinguishable from shuffling the same customers into different buckets.

Run an A/A test once

The cheapest way to make that lesson stick org-wide isn't a stats lecture, it's an A/A test: split traffic 50/50 between two identical versions of a page and report the results like a normal test. Since the true difference is 0% by construction, everything the report shows is noise and tooling, and it buys 3 things at once. It calibrates everyone on noise, because the gap between identical pages at your volume becomes the bar every future gap gets compared against. It shows metric speed: high-sample rates settle toward 0% within days while order-based metrics swing for weeks, which is why upper-funnel metrics carry early reads. And it validates the setup: the one line that should never hold a gap is the user split itself, so a 50/50 allocation sitting outside its noise band is an instrumentation problem to fix before any real test is trusted.

Here is what one looks like:

An A/A test · two identical pages, 6 weeks · gap between arms
Top-of-funnel rate (~20k events/day per arm)reads ~0% within days
+16%0-16%
Conversion rate (~30 orders/day per arm)±8% is still chance at week 6
+16%0-16%
User split, 50/50 intended (broken setup)sustained +3%: fix instrumentation first
+4%0-4%wk 1wk 2wk 3wk 4wk 5wk 6
Shaded band = the range pure chance covers (95%), shrinking as volume accumulates. The line is the measured gap between two arms with zero real difference.
The same math as a lookup: identical pages, by order volume
Orders per sideTypical CVR "difference"1 test in 20 shows over
100~10%~28%
200~7%~20%
500~4%~12%
1,500~2.5%~7%
5,000~1.3%~4%
Metrics settle in volume order: high-sample rates read clean in days, order-based metrics swing for 4+ weeks even sitewide, and any slice (one campaign, one product page) takes longer still. The third panel is the exception that reads instantly: user counts are so large that a 50/50 split holding a gap outside the band means the setup is broken.. fix that before trusting any test. [Simulated from the stated volumes; the mechanics are what an A/A on a real store shows.]

I would suggest running one sitewide for 4-6 weeks once a year or after any tooling change. It costs nothing.. both arms are the same site.. and it retires the early-read culture on its own: after everyone has watched identical pages "beat" each other by 10% in week 1, day-3 winner screenshots stop happening. It also converts "what is noise" from an argument into a screenshot everyone remembers.

From there, the only version of the run-length question with a useful answer is the one asked before launch: given our volume, how long until a lift we care about would be readable at all? That's the table. Find your row:

Weeks to a readable result · find your row
Weekly orders (in test)Detect a 20% lift10% lift5% lift
2503.4 wk13 wk50 wk
5002 wk6.4 wk25 wk
1,0002 wk3.2 wk13 wk
2,5002 wk2 wk5 wk
5,0002 wk2 wk2.5 wk
10,0002 wk2 wk2 wk
weeks ≈ 31 ÷ (weekly orders × lift²)  e.g. 500 orders, 10% lift: 31 ÷ (500 × 0.01) ≈ 6
95% confidence, 80% power, 50/50 split, whole weekly cycles with a 2-week floor. Baseline conversion rate cancels out of the math, so the table holds whether your CVR is 1% or 5%. Amber is a quarter of roadmap for one answer; red is not practically testable.. change the test, not the calendar.

What the red rows actually mean

If you're doing 250 orders a week through the tested page, a 5% conversion test needs about a year, and no ecom site holds still for a year.. cookies churn, promos shift, the site around the test keeps changing. So the honest conclusion isn't "we can't test", it's that your testing program has a minimum interesting effect size. At low volume, the only tests worth running are big swings: full page rebuilds, offer structure, template changes.. things that could plausibly move results 10-20%. Button colors and copy tweaks are tests for brands doing 5,000+ weekly orders, and running them below that produces results nobody can stand behind either way.

The other discipline the table buys you is a fixed stop date. Checking daily and stopping the moment significance appears roughly triples your false-winner rate, because noise crosses the threshold repeatedly on the way to nowhere. Compute the length, hold the length, and if the length is unacceptable, change the test rather than the calendar.

And it's worth saying plainly: statistics exists to improve a decision. If no decision changes with the result, or the readable version of the test costs a quarter of roadmap, I would suggest not running it and spending the traffic on a bigger question.

The most common way this goes wrong in practice is the test that won't die. We recently watched a paid-traffic landing page test kept alive for weeks.. first waiting for last-click orders to reach significance, then extended again to wait for attribution-modeled orders to build sample.. when both were always going to be too small to read. Meanwhile the upper-funnel metrics were already statistically significant and clearly positive: product-page view rate and add-to-cart up, purchases not moving the wrong way. That, plus the direction the team already wanted the experience to go, was the decision. The extra weeks of waiting bought nothing except delay on the next test.

Reading an inconclusive result

Even sized correctly, plenty of tests end without significance. That's information, not failure, and it has a reading procedure:

  1. Hypothesis confidence. Why did you believe it before the test? A mechanism you still believe survives an inconclusive read; a coin-flip idea doesn't.

  2. Behavioral inputs. Did add-to-cart, PDP click-through or units per transaction move in the supporting direction? Inputs are higher-volume than orders, so they read earlier and steady a noisy topline.

  3. Cost of being wrong. Shipping a false copy winner costs nothing. Shipping a false discount winner trains customers to wait for codes.. asymmetric costs get the strict read.

Strong hypothesis, supporting inputs and a cheap reversal: shipping on an inconclusive read is reasonable, and say so honestly. Weak hypothesis or costly reversal: kill it and spend the traffic on a bigger swing. The one non-option is calling it a win.

What to do with it

  1. Run an A/A test once (and after any tooling change): 4-6 weeks sitewide, reported like a real test. It validates allocation, and it calibrates the whole team on what noise looks like at your volume.

  2. Before building any test, run the formula: 31 ÷ (weekly orders × lift²). If the answer is longer than ~a quarter, redesign the test around a bigger swing rather than extending the calendar.

  3. At low volume, test fewer, bolder variants. One strong alternative against control beats 4 tweaks splitting the same traffic, since each added arm stretches every arm's timeline.

  4. Pre-commit the decision each outcome triggers.. ship, kill, iterate.. before data arrives, and hold the stop date. If you can't name a decision the test would change, that's the answer about whether to run it.

  5. Read inconclusive results with the 3 checks above, and be strictest where a false winner has a training cost. A discount test shipped on noise doesn't just misread the past, it teaches customers to wait for codes.

Happy to walk through the math for your volume and test queue.

Common questions

What is an A/A test and why run one?

An A/A test splits traffic 50/50 between two identical versions of a page, so the true difference is 0% by construction and everything the report shows is noise and tooling. It answers 3 questions no A/B test can: whether your allocation is actually random (more than ~1% uneven on a 50/50 split means a setup issue), how big the gaps between identical experiences are at your volume, and how long each metric takes to settle.. top-of-funnel rates typically converge within 1-2 weeks while conversion metrics swing for 4+ even sitewide. The result doubles as an internal calibration artifact: every future "the variant is up 8%" gets compared against what identical pages showed, which is usually the end of early winner declarations.

Can't I just run the test longer to detect a smaller lift?

In principle yes, and the formula tells you exactly how long. In practice, past ~10-12 weeks the test starts measuring a moving target: cookies churn so returning visitors get re-assigned, seasonality and promo mix shift who's arriving, and the site around the tested element keeps changing. A test that needs 6 months isn't a test, it's a queue blocking bolder ideas that could have read in 3 weeks. The workable range for most ecom tests is 2-8 weeks, which is why the effect size, not the calendar, is the lever worth adjusting.

Do I always need 95% significance?

No. Significance is a decision threshold, not a law of nature, and the right threshold depends on the cost of being wrong. A reversible, zero-cost change (copy, layout) can reasonably ship at 80-90% confidence, or even on a well-reasoned inconclusive result. An asymmetric-cost change (discounts, pricing, anything that trains customer behavior) deserves the strict read. What's not negotiable is choosing the threshold before looking at results, because a threshold chosen afterward always ends up being whatever the data happened to clear.

Should I test revenue per visitor instead of conversion rate?

Usually not as the primary metric. Revenue per visitor carries all of AOV's variance on top of conversion's, so it typically needs several times the sample.. and a single large order landing in one arm can swing the read for weeks. The decomposition works better: conversion rate as the primary metric, with AOV and units per transaction tracked as supporting reads. If the hypothesis is genuinely about basket size (bundles, thresholds), test on AOV directly with outliers capped and expect a longer run than the table shows.

See how we design experimentation programs

See how it works →
Technical notes

Sample sizes use a two-proportion test at 95% confidence (two-sided) and 80% power with a 50/50 split, run in whole weekly cycles with a 2-week floor. The weeks formula is ~31 ÷ (weekly orders × relative lift²); baseline conversion rate cancels out, so it holds across CVRs up to ~10%. Identical-page gaps are the sampling distribution of the difference between two equal conversion rates (SE of the relative gap ≈ √(2/n) at n orders per side; "typical" is the median absolute gap, "1 in 20" the 95th percentile). The 75-vs-70 example is illustrative.

SG
Sam Gil
Principal, Growth & Analytics · Meridian Growth
About the team →

Keep reading

ExperimentationSep 2026 · 6 min
How long to run an ecom A/B test
Read →
ExperimentationSep 2025 · 6 min
Why Most Startup Experiments Fail (and What to Do Instead)
Read →

Get in touch.

Tell us what you are trying to figure out. If we can help, we will say how. If we are not the right fit, we will say that too, and point you somewhere better.

01
Free work first
We start by doing a real piece of work on your data, free, so the first thing you judge is output, not a pitch.
02
First build, fixed scope
A trusted core metric layer in weeks, not quarters, at a fixed price. Small enough to be safe, real enough to matter.
03
Scale when it earns it
Month to month from there, from a few thousand a month. You own every model, dashboard and definition.