A/B Test Sample Size Calculator
Work out how many visitors each variant of a conversion A/B test needs, from the baseline rate, the smallest lift you care about, the significance level and the power. The answer is the smallest whole number that reaches your power, with the traffic and days it implies.
Analyzing a finished test? Use the z-test calculator for the two proportions. For a metric that is a mean rather than a rate, plan with the statistical power calculator.
The current rate of the control, for example 10 for 10%.
The smallest relative improvement worth detecting: 20 turns a 10% rate into 12%.
0.8 is the usual minimum.
Total across all variants; adds a duration estimate.
Related Calculators
Statistical Power Calculator
Compute power or required sample size for t tests, two proportions, ANOVA, and correlation with step-by-step noncentral distributions.
Confidence Interval for a Proportion Calculator
Estimate one proportion or the difference of two from counts with the Wald (1-PropZInt, 2-PropZInt), Wilson, Agresti–Coull, Clopper–Pearson, Jeffreys and Newcombe methods compared.
Z-Test Calculator
Run one-sample, two-sample, and proportion z-tests with z statistics, p-values, and critical values.
Learn More
Statistical Power Explained
What 80% power means, how effect size, sample size and alpha change it, and the sample size per group needed for small, medium and large effects.
Sample Size Explained: How Many Responses Do You Need?
Sample size formulas for a proportion and a mean, the finite population correction, an adjustment for dropout and a lookup table for margins of error, with worked examples.
What the calculator does
It plans a fixed-horizon test in which visitors are split equally between the control and each variant and every variant is compared with the control by a two-proportion z test. For your baseline rate and minimum detectable effect it finds the smallest number of visitors per variant whose power reaches the target, and shows the power at that number and at one fewer, so you can see the answer is minimal. Multiply by the number of variants for the total traffic, and divide by daily visitors for the duration.
Baseline, minimum detectable effect and why small lifts cost so much
The minimum detectable effect (MDE) is the smallest improvement you want the test to find reliably. It is not a forecast of the lift you expect; it is the size below which you would not act on the result. Enter it either as a relative lift (20% on a 10% baseline is 12%) or as absolute percentage points (2 points on a 10% baseline is also 12%).
The visitors needed grow with the inverse square of the difference between the two rates, and they grow as the baseline rate falls. Halving the lift roughly quadruples the traffic, and a rare event needs far more observations than a common one.
| Baseline rate | 5% lift | 10% lift | 20% lift | 50% lift |
|---|---|---|---|---|
| 1% | 637,009 | 163,095 | 42,693 | 7,750 |
| 2% | 315,206 | 80,682 | 21,109 | 3,826 |
| 5% | 122,124 | 31,234 | 8,158 | 1,471 |
| 10% | 57,763 | 14,751 | 3,841 | 686 |
| 20% | 25,583 | 6,510 | 1,683 | 294 |
| 50% | 6,274 | 1,565 | 388 | 58 |
Visitors per variant for a two-variant test, two-sided α = 0.05, power 0.80, relative lifts.
H₀: p₁ = p₂; reject when |p̂₁ − p̂₂| > z₁₋α/₂ · SE₀, SE₀ = √(2·p̄(1 − p̄)/n), p̄ = (p₁ + p₂)/2
SE₁ = √((p₁(1 − p₁) + p₂(1 − p₂))/n) (spread of p̂₁ − p̂₂ when the lift is real)
power(n) = Φ((|p₂ − p₁| − z·SE₀)/SE₁) + Φ((−|p₂ − p₁| − z·SE₀)/SE₁)
n = smallest whole number with power(n) ≥ target
Choosing the test duration
Days to run is the total sample divided by the visitors entering the experiment per day. Round up to whole weeks: conversion rates usually differ between weekdays and weekends, and a test that covers only part of the weekly cycle can be biased. Many teams also run for at least one or two full business cycles even when the sample size is reached sooner.
Why you should not stop early
The sample size assumes you look at the result once, when it is reached. Checking repeatedly and stopping at the first significant reading pushes the real false-positive rate well above α. If you need to monitor a test continuously, use a sequential or Bayesian design built for that, not this fixed-horizon calculator. See what a p-value means and Type I and Type II errors.
Testing more than one variant
With three or more variants every treatment is compared with the control. Turn on the Bonferroni correction to divide α by the number of treatment arms so the chance of any false winner stays at α; the sample size per variant rises, and the total is multiplied by the number of variants. Each comparison in this model uses the same equal share of traffic.
Worked example
A checkout page converts at 10%. You want to detect a relative lift of 20%, so the variant rate is 12%, with a two-sided test at α = 0.05 and power 0.80. The smallest sample that reaches it is 3,841 visitors per variant: the power there is 0.800017 and with 3,840 it is 0.799914, just below the target. Two variants need 7,682 visitors in total; at 5,000 visitors a day that is 2 days, which you would round up to a full week. Choose “Load example” to see the working.
Assumptions and limits
- The metric is a yes/no conversion and visitors are independent; clustered or repeated users need a different design.
- Traffic is split equally; unequal splits or a sample-ratio mismatch change the power.
- The z test uses a normal approximation, which is accurate for the sample sizes A/B tests need but not for tiny samples.
- For revenue per visitor or other averages, plan with the t-test option of the statistical power calculator using the metric’s standard deviation.
Software equivalents
| Software | Baseline 10%, variant 12%, power 0.8, α = 0.05 |
|---|---|
| R | power.prop.test(p1 = 0.10, p2 = 0.12, power = 0.8, sig.level = 0.05, strict = TRUE) # strict = TRUE counts both rejection tails; the default ignores the opposite tail and rounds to 3,842 |
| Python | statsmodels.stats.power.NormalIndPower().solve_power(effect_size=proportion_effectsize(0.12, 0.10), alpha=0.05, power=0.8) # arcsine (Cohen's h) approximation |
Frequently Asked Questions
What is the minimum detectable effect (MDE)?
The smallest improvement you want the test to detect with the chosen power. It is a planning choice, not a prediction: pick the lift below which you would not change the page. A relative MDE is a percentage of the baseline rate; an absolute MDE is in percentage points.
How many visitors do I need for an A/B test?
It depends on the baseline rate and the lift. A 10% baseline and a 20% relative lift (10% to 12%) needs 3,841 visitors per variant at 95% confidence and 80% power; a 2% baseline with the same relative lift needs 21,109 per variant.
Why does a low conversion rate need so much traffic?
The random variation of a rate is largest relative to the rate when the rate is small, so a fixed relative lift is a much smaller absolute difference to detect. The visitors needed grow roughly in proportion to (1 − p)/(p × lift²).
Relative or absolute lift: which should I enter?
Use whichever matches how your team states goals. They are two ways of writing the same variant rate: a 20% relative lift on a 10% baseline and a 2-point absolute lift both mean 12%. The calculator shows the variant rate so you can confirm.
Should the test be one-sided or two-sided?
Two-sided is the safe default because it also protects you from shipping a variant that is actually worse. A one-sided test needs less traffic but can only detect an improvement, and you must choose it before the test starts.
When should I use the Bonferroni correction?
When you compare several treatment variants with one control and want the probability of any false winner to stay at α. It divides α by the number of treatments, which raises the sample size per variant.
How long should I run the test?
At least until the calculated sample size is reached, rounded up to whole weeks so every weekday is represented equally. Do not stop early because the result looks significant: peeking inflates the false-positive rate.
Why does my number differ from another A/B calculator?
Calculators use slightly different formulas: pooled or unpooled variance, a continuity correction, or an arcsine transform, and some ignore the opposite rejection tail. This one uses the pooled standard error under the null and counts both tails for a two-sided test, so it can differ from a rule-of-thumb formula by a few percent.
What if I cannot reach the required sample size?
Raise the minimum detectable effect, accept lower power, run the test longer, reduce the number of variants, or test a page with more traffic. Running an underpowered test mostly produces inconclusive results.
Does this work for revenue or other averages?
No. It is for conversion rates, which are proportions. For a mean such as revenue per visitor, use the statistical power calculator with the two-sample t test and the metric's standard deviation.
Embed This Calculator
Add this free calculator to your course page or LMS.
Adjust the height value to fit your page.