A/B Test Sample Size: How Many Visitors Do You Actually Need?
A practical decision framework for estimating A/B test sample size, test duration and whether your traffic is high enough to make the experiment worth running.
Read articleShare
A practical framework for estimating how long an A/B test should run—and deciding whether your traffic can produce a useful answer before the opportunity changes.
“How long should an A/B test run?” is often answered with a calendar rule: one week, two weeks, or “until significance.” None of those is a reliable planning method on its own.
Test duration is mostly the result of two earlier decisions: how many observations the experiment needs and how much eligible traffic can reach the test each day.
The practical question for a growth team is therefore not “Should we run this for 14 days?” It is:
Can this experiment collect the required sample within a timeframe that is short enough to influence the decision we care about?
A/B test duration should be estimated after you define the baseline conversion rate, minimum detectable effect, confidence threshold and statistical power. Those inputs determine the required sample. Traffic then determines how many calendar days are needed to collect it.
Start with the effect you need to detect, calculate the required sample, then divide by the traffic that can actually enter the experiment.
Defines the starting signal level
Defines the smallest lift worth detecting
Translates the statistical plan into observations
Turns the sample requirement into calendar time
Decision rule
Do not choose a test duration first. Choose the decision-relevant effect, estimate the sample, then ask whether your traffic can deliver that sample within a timeframe the business can tolerate.
| Input | What it controls | Effect on duration |
|---|---|---|
| Baseline conversion rate | How much outcome signal exists in the current funnel | Changes the sample required for a given effect |
| Minimum detectable effect (MDE) | The smallest improvement worth detecting | Smaller MDE usually means a much larger sample and longer test |
| Confidence and power | The false-positive and false-negative risk accepted in the design | Stricter settings increase the required sample |
| Eligible experiment traffic | How many visitors can actually be allocated to the test | More daily eligible traffic shortens the calendar duration |
| Traffic allocation | The share of eligible visitors entering the variants | Partial exposure extends the test versus using all eligible traffic |
This is why traffic volume alone cannot tell you whether a test will take five days or five weeks. A high-traffic page can still need a long test if the team is trying to detect a very small lift. A lower-traffic funnel can finish sooner when the minimum effect worth acting on is large.
Once the required sample is known, the planning arithmetic is straightforward:
Estimated test days = required total sample ÷ eligible experiment visitors per day
If your sample-size calculation requires 24,000 total visitors and 2,000 eligible visitors enter the experiment each day, the statistical sample could be collected in about 12 days.
If only half of that traffic can be exposed to the experiment, the effective eligible traffic is 1,000 visitors per day and the same sample takes roughly 24 days.
The duration calculation is therefore only as good as the traffic estimate. Use traffic for the exact page, audience, device or funnel step that will enter the experiment—not total site sessions.
If you still need the sample estimate itself, the A/B test sample size guide explains how baseline conversion, MDE, confidence and power determine the number of observations required.
Assume an experiment needs 30,000 total eligible visitors under the chosen statistical design. The only thing changing below is the amount of traffic the experiment receives.
| Eligible traffic per day | Required sample | Raw duration estimate | Planning implication |
|---|---|---|---|
| 5,000 | 30,000 | 6 days | Fast enough that full-week coverage may become the bigger constraint |
| 2,000 | 30,000 | 15 days | Reasonable for many website experiments |
| 750 | 30,000 | 40 days | Opportunity cost starts to matter |
| 250 | 30,000 | 120 days | Likely too slow unless the decision is unusually valuable and stable |
These are illustrative planning examples, not universal test-duration rules. The business question is whether the expected learning is worth locking the page, feature or decision for that long.
Minimum detectable effect is one of the strongest levers in experiment planning. Trying to detect a smaller effect usually requires substantially more observations, which then increases the calendar duration.
That is not a reason to inflate the MDE to force a shorter test. The MDE should represent the smallest change that would genuinely alter the business decision.
For example, if a 2% relative lift would not change budget allocation, product direction or economics, designing a test to detect that effect creates a long experiment around a result the business may not care about. If a 10% lift would materially change CAC or revenue, a design centred on that decision threshold may be more useful.
A raw sample calculation can say the experiment needs six days. That does not automatically mean the test should stop after six calendar days.
Traffic and conversion behaviour often differ by weekday, weekend, campaign schedule, device mix or sales cycle. Ending a test before it captures the normal operating pattern can make the sample less representative of the decision environment.
For many website experiments, covering at least one complete weekly cycle is a sensible operational check. Longer B2B or subscription funnels may need a different cycle depending on how long it takes users to reach the measured outcome.
The important distinction is that a full-week rule is a coverage constraint, not a substitute for sample-size planning. A test can cover two full weeks and still be underpowered.
A statistically feasible experiment can still be a poor business experiment. The opportunity may change before the result arrives.
| Signal | Why it matters | Better response |
|---|---|---|
| Expected duration is several months | The market, campaign or product may change before the answer arrives | Test a larger decision-relevant change or use another research method |
| The experiment blocks a high-value redesign | The cost of waiting may exceed the value of precision | Prioritise the higher-value decision |
| Traffic is fragmented across many variants | Each variant receives less information | Reduce variant count or sequence experiments |
| The measured outcome is very rare | Signal accumulates slowly | Consider a validated upstream metric while protecting downstream quality |
Low traffic does not mean experimentation is impossible. It means the team must be more selective about what deserves a controlled test.
Measurement quality matters too. If conversion events are duplicated, missing or defined differently across tools, a longer experiment will not repair the underlying data problem. The GA4 audit checklist is useful when the experiment outcome depends on analytics events you do not fully trust.
This process prevents the common mistake of treating test duration as a fixed rule. The correct duration is the time needed to collect enough relevant data for a decision that is still commercially useful when the answer arrives.
Sometimes. Two weeks can be enough if the required sample is reached, the test covers the relevant traffic cycle and the experiment was planned with an appropriate MDE, confidence level and power. It is not a universal minimum.
Not automatically. The stopping rule should match the statistical method used to design the experiment. Repeatedly checking a conventional fixed-horizon test and stopping as soon as the threshold is crossed can increase error risk.
More eligible traffic reduces the calendar time needed to collect a fixed sample. But traffic does not reduce the sample requirement itself unless the statistical design or effect size changes.
Revisit the business question. Consider whether a larger meaningful change, fewer variants, a higher-volume step or another research method can produce a useful decision sooner without pretending the original test needs less evidence.
Share this article
Next step
Use your baseline conversion rate, MDE and statistical settings to estimate the required sample, then translate that sample into a realistic test duration using eligible traffic.
Related
A practical decision framework for estimating A/B test sample size, test duration and whether your traffic is high enough to make the experiment worth running.
Read article
A practical GA4 audit framework for finding tracking, attribution, revenue and data-quality problems — and prioritising the issues that can damage business decisions.
Read articleA practical framework for turning a revenue target into a defensible marketing budget using customer volume, CAC scenarios, unit economics and decision-ready data.
Read article