Statistical Power Calculator

Find the statistical power of a study, the sample size it needs, or the smallest effect it can detect. Covers t tests, two proportions, one-way ANOVA and correlation, using exact noncentral t and F distributions, with every step shown.

Planning an online experiment? The A/B test sample size calculator takes a baseline rate and a relative lift. To size an effect from data you already have, use the effect size calculator.

Mean difference divided by the standard deviation. Benchmarks: 0.2 small, 0.5 medium, 0.8 large.

Leave blank for the same size as group 1.

What statistical power means

Power is the probability that a test rejects the null hypothesis when a specific alternative is true. It equals 1 − β, where β is the Type II error rate: the chance of missing a real effect. Four quantities determine it, and fixing any three fixes the fourth: the effect size, the sample size, the significance level α and the variability of the data.

A power of 0.80, Cohen’s long-standing convention, accepts a 20% chance of missing an effect of the planned size; confirmatory studies often plan for 0.90 or 0.95. Power is a property of the design, computed before data collection from an effect size you consider worth detecting.

Three ways to use the calculator

  • Power — you have a sample size and an effect size in mind and want to know how likely the test is to find it.
  • Sample size — you know the effect and the power you want; the result is the smallest whole number per group that reaches it, together with the power at one fewer, so you can see that the answer is minimal.
  • Minimum detectable effect — you are limited to a fixed sample size (a sensitivity analysis); the result is the smallest effect the study would detect with the target power.

The tests supported are the two-sample, one-sample and paired t tests, the two-proportion z test, one-way ANOVA with equal group sizes, and the test of a Pearson correlation. One-sided tests assume the effect lies in the direction you entered.

t tests: δ = d·√(n₁n₂ / (n₁ + n₂)) (two samples), δ = d·√n (one sample, paired)

power = P(|T| > t₁₋α/₂, df) with T ~ noncentral t(df, δ)

ANOVA: λ = k·n·f², power = P(F > F₁₋α, df₁, df₂) with F ~ noncentral F(k − 1, k(n − 1), λ)

Correlation: power = P(|Z| > z₁₋α/₂) with Z ~ N(atanh(r)·√(n − 3), 1)

Two proportions: reject when |p̂₁ − p̂₂| > z₁₋α/₂·SE₀ (pooled); distribution under H₁ uses SE₁ (unpooled)

Effect size benchmarks

When nothing better is known, Cohen (1988) suggested these conventional values for a small, medium and large effect. They are starting points: the right target is the smallest effect that would matter in your field, or an estimate from a pilot study or earlier research.

MeasureSmallMediumLarge
Cohen's d (t tests)0.20.50.8
Cohen's f (one-way ANOVA)0.10.250.4
Correlation r0.10.30.5
Cohen's h (two proportions)0.20.50.8

Sample size per group for a two-sample t test

Two-sided test at α = 0.05 with equal group sizes. These figures match G*Power and R’s pwr.t.test, and the calculator above reproduces every row.

Cohen's dPower 0.80Power 0.90Power 0.95
0.2394527651
0.3176235290
0.4100133164
0.56486105
0.6456074
0.8263442
1172327
1.2121620

Sample size per group for one-way ANOVA

Equal group sizes at α = 0.05; Cohen’s f of 0.10, 0.25 and 0.40 are the small, medium and large conventions.

Cohen's fGroups (k)Power 0.80Power 0.90
0.13323423
0.14274356
0.15240310
0.2535369
0.2544558
0.2554051
0.432228
0.441924
0.451621

Sample size for detecting a correlation

Total number of pairs for a two-sided test of ρ = 0 at α = 0.05. The values use Fisher’s z approximation, which can differ by one from an exact calculation.

Correlation rPower 0.80Power 0.90
0.17831047
0.2194259
0.385113
0.44762
0.53038

Worked example: comparing two groups

You expect a medium difference, d = 0.5, between two independent groups and will use a two-sided t test at α = 0.05. With 64 people per group, δ = 0.5·√(64·64 / 128) = 2.8284 and the test has 126 degrees of freedom. The critical value is t = 1.979, and the probability that a noncentral t variable with δ = 2.8284 lands beyond it is 0.8015. Solving for the sample size returns 64 per group; with 63 per group the power is 0.7952, just under the 0.80 target. Choose “Load example” to see the full working and the power curve.

Assumptions and limits

  • The t-test formulas assume normally distributed data and, for two groups, equal variances (the pooled test). With unequal variances and unequal group sizes the Welch test has somewhat different power.
  • The two-proportion and correlation results use normal approximations (the correlation via Fisher’s z), so they can differ from exact-test software by a small amount at small sample sizes.
  • Power for the Mann–Whitney U test is close to the t test under normality: divide the t-test sample size by 0.955 to plan for it.
  • Do not compute “observed power” from the effect you measured after the study; it is a function of the p-value and adds no information. Report the confidence interval instead.
  • Add extra participants for expected dropout: divide the calculated n by the share you expect to keep.

Software equivalents

SoftwareTwo-sample t test, d = 0.5, power 0.8, α = 0.05
G*Powert tests → Means: Difference between two independent means → A priori: effect size d = 0.5
R (stats)power.t.test(delta = 0.5, sd = 1, sig.level = 0.05, power = 0.8)
R (pwr)pwr.t.test(d = 0.5, sig.level = 0.05, power = 0.8, type = "two.sample")
Pythonstatsmodels.stats.power.TTestIndPower().solve_power(effect_size=0.5, alpha=0.05, power=0.8)
Statapower twomeans 0 0.5, power(0.8)

Frequently Asked Questions

What is a good power level for a study?

Power of 0.80 is the usual minimum: it means a 20% chance of missing an effect of the planned size. Confirmatory and clinical studies often plan for 0.90 or 0.95, which needs about a third and two-thirds more participants than 0.80 for the same effect (64, 86 and 105 per group at d = 0.5).

How do I choose the effect size?

Use the smallest effect that would matter in practice, an estimate from a pilot study, or the effect reported in similar research, and check the calculator with a smaller value as a safety margin. Cohen's small, medium and large conventions (d = 0.2, 0.5, 0.8) are a fallback when nothing else is known, not a substitute for subject knowledge.

Why does the sample size grow so fast for small effects?

Sample size is inversely proportional to the square of the effect size. Halving the effect from d = 0.5 to d = 0.25 multiplies the required n by about four (from 64 to 253 per group at 80% power).

What is the difference between one-sided and two-sided power?

A one-sided test puts all of α in one tail, so it needs fewer participants for an effect in the predicted direction but cannot detect an effect in the other direction. Choose the alternative before collecting data and enter the effect as a positive number.

Do I enter n per group or the total sample size?

For the two-sample t test and two proportions, enter the size of each group; the results also list the total. For a one-sample or paired test enter the number of subjects or pairs, for ANOVA the size of each group, and for a correlation the number of pairs.

How is the paired t test different from the independent one?

The paired test analyses the within-pair differences, so the effect size is the mean difference divided by the standard deviation of those differences, and the sample size is the number of pairs. Strongly correlated pairs shrink that standard deviation and raise power.

Why might my result differ slightly from G*Power, R or statsmodels?

The t-test and ANOVA results use the exact noncentral t and F distributions and agree with those programs to many decimals. Two-proportion and correlation results use normal approximations, and different programs use different variants (pooled or unpooled variance, Fisher's z or an exact bivariate normal model), which can shift the sample size by one or a few.

What is a sensitivity analysis?

It answers the reverse question: given the sample size you can actually afford, what is the smallest effect you would reliably detect? Choose 'Minimum detectable effect' to get it, and state that value when you report a study that did not find a significant effect.

Does the calculator handle multiple comparisons?

Not directly. To plan for several tests, divide α by the number of comparisons (Bonferroni) before entering it; the required sample size then rises accordingly.

Can I use it for the Welch test or a nonparametric test?

The two-sample result assumes equal variances. Welch's test loses little power when group sizes are equal. For the Mann–Whitney U or Wilcoxon signed-rank test under roughly normal data, divide the t-test sample size by 0.955 as a planning rule.

Embed This Calculator

Add this free calculator to your course page or LMS.

Adjust the height value to fit your page.