A/B Test Calculator
Find out whether a variant really converts better than the control. Enter visitors and conversions for both versions to get the conversion rates, the relative uplift with its confidence interval, the p-value of the two-proportion z-test, a plain-language verdict and a check that the traffic split matches your plan.
Planning a test? The A/B test sample size calculator tells you how many visitors each version needs. The statistics behind this page: two proportion z-test.
Unique visitors who saw the current version
How many of them converted
Unique visitors who saw the new version
How many of them converted
95% means a 5% chance of a false winner
Decide before you look at the data
50 for an even split
Related Calculators
Two Proportion Z-Test Calculator
Compare two proportions with the pooled z-test: z statistic, p-value and a confidence interval for the difference (Wald, Newcombe or Agresti–Caffo).
A/B Test Sample Size Calculator
Plan fixed-horizon A/B tests: sample size per variant from baseline conversion, MDE, α, and power.
Statistical Power Calculator
Compute power or required sample size for t tests, two proportions, ANOVA, and correlation with step-by-step noncentral distributions.
How to use the A/B test calculator
- Enter the visitors and conversions of the control (A, the current version) and of the variant (B, the new version).
- Choose the confidence level. 95% is the usual standard: it accepts a 5% chance of declaring a winner when the two versions actually convert equally.
- Pick the test type. Use two-sided unless you decided before the test that only an improvement would change your decision.
- Set the share of traffic you planned to send to the variant (50% for an even split).
- Read the verdict, the p-value, the uplift and the intervals, and check the notes for a traffic split warning.
A conversion can be anything you count once per visitor: a purchase, a sign-up, a click. Enter visitors, not page views, so that the same person is not counted twice. The calculator is a two-proportion z-test with a plain-language layer; for the same test written as a textbook hypothesis test, see the two proportion z-test calculator.
How the calculator decides
p_A = x_A / n_A, p_B = x_B / n_B
p̂ = (x_A + x_B) / (n_A + n_B) (pooled rate under H₀)
SE = √( p̂ (1 − p̂) (1/n_A + 1/n_B) )
z = (p_B − p_A) / SE
Relative uplift = p_B / p_A − 1
Split check: χ² = (n_B − N·s)² / (N·s·(1 − s)), 1 degree of freedom
The p-value is the probability, if the two versions really converted equally, of a difference at least as large as the one you saw. The result is significant when the p-value is at most α, which is 1 minus the confidence level. Each tail of the normal distribution is computed directly, so even a very small p-value keeps its digits.
The interval for the absolute difference is the Newcombe hybrid score interval, which stays reliable at low conversion rates and small counts. The interval for the relative uplift is the Katz log interval for p_B / p_A with 1 subtracted; it is undefined when a group has no conversions. Both intervals are two-sided at the confidence level you chose, even for a one-sided test: a one-sided test at 95% confidence agrees with a two-sided interval at 90%.
The traffic split check is a chi-square goodness-of-fit test of the visitor counts against the planned split. A p-value below 0.001 is flagged as a possible sample ratio mismatch, the most common sign that an experiment is broken. The same test for any observed and expected counts is the chi-square goodness of fit calculator.
Worked example
The control page had 500 conversions from 10,000 visitors (5%). The new variant had 570 conversions from 10,000 visitors (5.7%). Is the variant really better?
- Pooled rate p̂ = 1,070 / 20,000 = 0.0535.
- SE = √(0.0535 × 0.9465 × (1/10,000 + 1/10,000)) = 0.003182.
- z = (0.057 − 0.05) / 0.003182 = 2.1996.
- Two-sided p-value = 2 × P(Z ≥ 2.1996) = 0.027835; one-sided p-value = 0.013917.
The absolute difference is +0.7 percentage points and the relative uplift is +14%. At 95% confidence the interval for the absolute difference is (0.0761, 1.325) pp and the interval for the uplift is (1.4279%, 28.1305%). Load example fills in these numbers. The verdict depends on the confidence level you demand:
| Confidence level | Two-sided verdict (p = 0.027835) | Interval for the difference |
|---|---|---|
| 80% | Variant wins (p is at most 0.2) | (0.2922, 1.1083) pp |
| 90% | Variant wins (p is at most 0.1) | (0.1765, 1.2243) pp |
| 95% | Variant wins (p is at most 0.05) | (0.0761, 1.325) pp |
| 99% | No significant difference (p is above 0.01) | (−0.1205, 1.5225) pp |
The interval says more than the verdict. At 95% the true gain could be as small as 0.08 percentage points or as large as 1.3, and at 99% zero is still possible. If the cost of shipping a change is high, wait for more data rather than declare a winner on a lower bound that close to zero.
Reading the result correctly
- Fix the sample size in advance and look once. Checking every day and stopping when the p-value first drops below 0.05 raises the real false-positive rate far above 5%. Work out the sample size first with the A/B test sample size calculator or the statistical power calculator.
- Correct for several variants. Each extra comparison gives another chance of a false winner. With more than one variant against the control, apply the Bonferroni correction or test all versions together with the chi-square test of independence.
- Statistical significance is not business impact. With enough traffic a trivial lift is significant. Use the interval and the effect size to judge whether the gain is worth having.
- Trust the traffic split. If the split check fails, the groups may differ for reasons other than your change, and no p-value repairs that.
- Run whole weeks. Behavior changes with the day of the week, so run at least one or two full weeks even if significance arrives sooner.
- Watch small counts. When the smallest expected count is below 10, use the Fisher exact test instead of the z-test.
For the meaning of the p-value and the two kinds of error, read what p < 0.05 actually means and Type I vs Type II errors. To put an interval around a single conversion rate, use the proportion confidence interval calculator.
Frequently Asked Questions
How do I calculate A/B test statistical significance?
Compute the conversion rate of each version, pool the two groups into one rate, and divide the difference of the rates by the pooled standard error to get a z statistic. The p-value is the area under the standard normal curve beyond z; the result is significant when it is at most 1 minus the confidence level. Enter visitors and conversions above to see each step.
What confidence level should I use for an A/B test?
95% is the common default: it accepts a 5% chance of declaring a winner when the versions are the same. Use 99% when a wrong decision is costly, and 90% or 80% only for low-stakes tests where you accept more false winners in exchange for faster decisions. Choose it before the test, not after seeing the result.
Should I use a one-sided or two-sided test?
Use a two-sided test unless you decided before collecting data that only an improvement matters. A one-sided test has more power to detect a gain but can never reveal that the variant is worse. Switching to one-sided after seeing that the two-sided result just missed significance is not valid.
How many visitors do I need for an A/B test?
It depends on your baseline conversion rate, the smallest lift you care about and the power you want. A lower baseline or a smaller lift needs many more visitors. The A/B test sample size calculator gives the number per version; run the test until you reach it instead of stopping when the p-value looks good.
Why do different A/B test calculators show different p-values?
Tools differ in the standard error (pooled or unpooled), in one-sided versus two-sided p-values, in continuity corrections, and in whether they use a frequentist test at all: some report a Bayesian chance to beat the control. This calculator uses the pooled two-proportion z-test, the same test as R's prop.test without continuity correction.
What is a sample ratio mismatch and why does it matter?
A sample ratio mismatch means the visitor counts differ from the planned split by more than chance allows, for example 10,000 against 9,500 in a 50/50 test (p = 0.0003). It usually points to a bug in assignment, redirects or tracking, and it makes the comparison unreliable. Investigate it before reading the conversion results.
Can I stop the test as soon as it is significant?
Not with a fixed-sample test like this one. The p-value wanders as data arrives, and stopping the first time it dips below 0.05 makes false winners far more common than 5%. Fix the sample size and the duration in advance and analyze once, or use a sequential method designed for repeated looks.
What is the difference between relative uplift and absolute difference?
The absolute difference subtracts the rates: 5.7% minus 5% is 0.7 percentage points. The relative uplift divides them: 5.7% is 14% higher than 5%. The same absolute difference is a much larger relative uplift on a low baseline, so report both and always give the interval.
How do I analyze a test with more than two variants?
Comparing each variant with the control multiplies the chances of a false winner. Adjust the significance level, for example with the Bonferroni correction (α divided by the number of comparisons), or first test all versions together with a chi-square test of independence on the table of conversions and non-conversions.
Embed This Calculator
Add this free calculator to your course page or LMS.
Adjust the height value to fit your page.