How Long Should You Run an A/B Test? Traffic, Conversion Rate and MDE Explained
A practical guide to estimating A/B test duration from sample size, eligible traffic, baseline conversion rate and the smallest effect worth detecting.
Read articleShare
A practical framework for deciding how much traffic your A/B test needs, how long it may take and whether the experiment is commercially worth running.
“How many visitors do I need for an A/B test?” sounds like a simple question. It is not. The answer depends on your current conversion rate, the smallest improvement that would actually matter, the level of uncertainty you are willing to accept and how much eligible traffic can enter the experiment.
For a business owner, CMO or growth lead, the more useful question is:
Can this experiment produce a reliable answer quickly enough to change a meaningful business decision?
That framing changes the goal. You are not trying to “get statistical significance.” You are deciding whether a test deserves traffic, calendar time and organisational attention.
There is no credible rule such as “every A/B test needs 10,000 visitors.” A test with a high baseline conversion rate and a large expected improvement may need far fewer observations than a test looking for a very small change in a low-converting funnel.
The minimum sample is therefore not a fixed traffic threshold. It is the sample required for the specific effect you want to detect under the assumptions you set before the experiment starts.
Many weak experiments begin with a page element: button colour, headline, layout or form field. Strong experiments begin with a business constraint.
Examples:
Once the business problem is clear, define what result would be large enough to matter. This is the commercial meaning behind the minimum detectable effect, or MDE.
Sample size is not the first question. Start with the business effect you care about, then decide whether your traffic can support a reliable answer in a useful timeframe.
Use the current conversion rate for the exact outcome you want to improve.
Define the smallest improvement that would change a business decision.
Estimate how many visitors or conversions each variant needs.
Translate sample size into calendar time using eligible traffic.
Run, redesign, consolidate traffic or choose another evidence method.
| Input | Business meaning | What happens to sample size | Common mistake |
|---|---|---|---|
| Baseline conversion rate | The current rate for the exact outcome being tested | Changes the amount of signal available in each visitor | Using a site-wide conversion rate for a narrower experiment |
| Minimum detectable effect | The smallest lift worth acting on | Smaller effects require substantially more observations | Choosing an unrealistically tiny effect because “more precision is better” |
| Confidence / significance threshold | How much false-positive risk the plan accepts | More stringent thresholds generally require more sample | Changing the threshold after seeing results |
| Statistical power | The probability of detecting the planned effect when it is real | Higher power requires more sample | Ignoring false-negative risk entirely |
These inputs interact. The most important practical relationship is simple: the smaller the improvement you want to detect, the more traffic you usually need.
That is why sample-size planning is also prioritisation. A low-traffic company cannot test every small optimisation independently. It must reserve experiments for changes with enough potential impact to justify the sample.
Imagine a paid-acquisition landing page with a baseline conversion rate of 3%. The team wants to know whether a new value proposition can improve that rate.
If the business would only care about a large improvement — for example, enough to materially reduce CAC — the required sample may be manageable. If the team wants to distinguish a very small improvement from normal random variation, the required sample can become much larger.
This is where commercial context matters. A small lift on a funnel producing millions in annual contribution may be worth a long experiment. The same relative lift on a low-volume page may not justify waiting months for an answer.
A sample-size result becomes useful only after you translate it into calendar time. If you need 20,000 eligible visitors but only 500 visitors per day can enter the experiment, the test has a very different operational cost than a test receiving 10,000 eligible visitors per day.
Use eligible traffic, not total website traffic. Exclude visitors who cannot enter the experience, traffic outside the relevant market, bots, employees and other populations that do not belong in the decision.
The same statistical target can be commercially useful on a high-traffic funnel and impractical on a low-traffic one. These examples are directional, not universal sample-size rules.
200 eligible visitors/day
Long
Usually too slow for small lifts
Best response
Test bigger changes or use another evidence method
1,000 eligible visitors/day
Moderate
Many practical tests become feasible
Best response
Prioritise commercially meaningful hypotheses
5,000+ eligible visitors/day
Shorter
Smaller effects become more testable
Best response
Protect quality: do not trade rigor for speed
Also remember that a calendar estimate is not permission to stop the test at an arbitrary number of days. It is a planning estimate built from traffic and sample requirements. Your experiment should still be run according to the statistical design chosen before launch.
Low traffic does not mean the business cannot learn. It means classical split-testing may not be the best method for every question.
If the calculated test would take too long, consider:
The wrong response is to lower the standard until the test “fits” the traffic. A statistically underpowered experiment does not become useful simply because the business is impatient.
One of the most common experimentation mistakes is repeatedly checking a test and stopping as soon as one variant looks ahead. Early results naturally move around as data accumulates. Treating every temporary lead as a final answer increases the chance of acting on noise.
From a management perspective, this is not merely a statistical problem. It creates false confidence: teams roll out weak changes, attribute growth to the wrong cause and build the next decision on contaminated evidence.
Define the sample plan, primary outcome and decision rule before launch. If the company needs continuous monitoring or sequential testing, use a method designed for that purpose rather than improvising a stopping rule after the results appear.
A test can be statistically clean and commercially misleading if it optimises the wrong metric.
For example, shortening a form may increase submissions but reduce qualified leads. A more aggressive discount may increase purchases but weaken margin or customer quality. A click-through improvement may not improve activation, revenue or payback.
The primary metric should therefore sit as close as practical to the business decision while still occurring frequently enough to test. When downstream outcomes matter, define guardrails such as qualified-lead rate, refund rate, margin or activation quality.
If the conversion event itself is unreliable, fix that before interpreting an experiment. The GA4 audit checklist is a useful starting point when duplicate events, missing conversions or inconsistent definitions undermine the test metric. Broader identity, CRM and revenue gaps may require marketing analytics infrastructure that connects the experiment to commercial outcomes.
A strong candidate usually has four characteristics:
If the test fails the fourth condition, do not run it. An experiment whose outcome cannot change a decision is reporting activity, not experimentation.
| Situation | Recommended action | Why |
|---|---|---|
| High-value decision + sufficient traffic | Run the test | The learning can arrive in time to change a meaningful decision |
| High-value decision + insufficient traffic | Redesign the hypothesis or evidence plan | The question matters, but the current test design is too expensive in time |
| Small expected impact + large required sample | Deprioritise | Traffic is better allocated to a larger business constraint |
| Unreliable measurement | Fix instrumentation first | More sample cannot repair a broken outcome definition |
This framework is intentionally stricter than “we have an idea, so let us test it.” The scarce resource is not ideas. It is trustworthy traffic applied to decisions with enough potential value.
When experimentation is part of a wider growth problem — weak funnel economics, unclear priorities, poor measurement or conflicting channel signals — a Growth Audit can help identify which questions deserve evidence first. If the problem is executing growth systems rather than diagnosing them, Growth Engineering provides the implementation path.
For budget decisions connected to experimentation, the marketing budget for a revenue target guide shows how measurement quality and conversion assumptions affect the amount of growth investment a business can defend.
It depends on baseline conversion rate, the minimum effect worth detecting, statistical threshold, power and traffic allocation. There is no universal visitor minimum that works for every experiment.
Enough eligible traffic to reach the planned sample within a useful business timeframe. A website can have substantial total traffic but still be a poor candidate if only a small fraction enters the tested journey.
Sometimes, especially for large expected effects or frequent outcomes. If the required sample implies an impractically long test, prioritise bigger changes, consolidate traffic or use another evidence method rather than accepting an underpowered test.
Duration follows from required sample and eligible traffic, not from a fixed “seven-day” or “two-week” rule. The test plan should also account for normal business cycles and use a pre-defined stopping approach.
A useful MDE is the smallest improvement that would actually change the business decision. Choosing a smaller MDE than the business cares about can dramatically increase the sample without adding decision value.
The best A/B test is not the one with the most statistical sophistication. It is the one that produces a trustworthy answer to a commercially meaningful question while the answer is still useful.
Start with the business decision. Define the baseline. Choose the smallest effect worth acting on. Calculate the sample. Translate it into duration using eligible traffic. Then decide whether to run, redesign or deprioritise the experiment.
Ready to plan the test? Calculate the sample size and practical traffic requirement for your A/B test.
Need to decide which growth problem deserves testing first? Discuss your growth priorities.
Share this article
Next step
Estimate the required sample, translate it into test duration and decide whether the expected learning is worth the traffic and time you are about to invest.
Related
A practical guide to estimating A/B test duration from sample size, eligible traffic, baseline conversion rate and the smallest effect worth detecting.
Read article
A practical GA4 audit framework for finding tracking, attribution, revenue and data-quality problems — and prioritising the issues that can damage business decisions.
Read articleA practical framework for turning a revenue target into a defensible marketing budget using customer volume, CAC scenarios, unit economics and decision-ready data.
Read article