Statistical Power Calculator
Find the statistical power of a study, the sample size it needs, or the smallest effect it can detect. Covers t tests, two proportions, one-way ANOVA and correlation, using exact noncentral t and F distributions, with every step shown.
Planning an online experiment? The A/B test sample size calculator takes a baseline rate and a relative lift. To size an effect from data you already have, use the effect size calculator.
Mean difference divided by the standard deviation. Benchmarks: 0.2 small, 0.5 medium, 0.8 large.
Leave blank for the same size as group 1.
Related Calculators
Effect Size Calculator (Cohen's d)
Quantify the difference between two groups with Cohen's d, Hedges' g, and pooled SD.
Sample Size Calculator
Plan surveys with a confidence level, margin of error, and finite population correction.
One-Sample t-Test Calculator
Test a mean against a hypothesized value from data or x̄, Sx, and n — the TI-84 T-Test output.
Learn More
Statistical Power Explained
What 80% power means, how effect size, sample size and alpha change it, and the sample size per group needed for small, medium and large effects.
Effect Size Explained: Cohen's d, r and Eta Squared
What Cohen's d, Hedges' g, r, eta squared and Cramér's V measure, how to calculate d by hand, and why a small p-value is not enough.
What statistical power means
Power is the probability that a test rejects the null hypothesis when a specific alternative is true. It equals 1 − β, where β is the Type II error rate: the chance of missing a real effect. Four quantities determine it, and fixing any three fixes the fourth: the effect size, the sample size, the significance level α and the variability of the data.
A power of 0.80, Cohen’s long-standing convention, accepts a 20% chance of missing an effect of the planned size; confirmatory studies often plan for 0.90 or 0.95. Power is a property of the design, computed before data collection from an effect size you consider worth detecting.
Three ways to use the calculator
- Power — you have a sample size and an effect size in mind and want to know how likely the test is to find it.
- Sample size — you know the effect and the power you want; the result is the smallest whole number per group that reaches it, together with the power at one fewer, so you can see that the answer is minimal.
- Minimum detectable effect — you are limited to a fixed sample size (a sensitivity analysis); the result is the smallest effect the study would detect with the target power.
The tests supported are the two-sample, one-sample and paired t tests, the two-proportion z test, one-way ANOVA with equal group sizes, and the test of a Pearson correlation. One-sided tests assume the effect lies in the direction you entered.
t tests: δ = d·√(n₁n₂ / (n₁ + n₂)) (two samples), δ = d·√n (one sample, paired)
power = P(|T| > t₁₋α/₂, df) with T ~ noncentral t(df, δ)
ANOVA: λ = k·n·f², power = P(F > F₁₋α, df₁, df₂) with F ~ noncentral F(k − 1, k(n − 1), λ)
Correlation: power = P(|Z| > z₁₋α/₂) with Z ~ N(atanh(r)·√(n − 3), 1)
Two proportions: reject when |p̂₁ − p̂₂| > z₁₋α/₂·SE₀ (pooled); distribution under H₁ uses SE₁ (unpooled)
Effect size benchmarks
When nothing better is known, Cohen (1988) suggested these conventional values for a small, medium and large effect. They are starting points: the right target is the smallest effect that would matter in your field, or an estimate from a pilot study or earlier research.
| Measure | Small | Medium | Large |
|---|---|---|---|
| Cohen's d (t tests) | 0.2 | 0.5 | 0.8 |
| Cohen's f (one-way ANOVA) | 0.1 | 0.25 | 0.4 |
| Correlation r | 0.1 | 0.3 | 0.5 |
| Cohen's h (two proportions) | 0.2 | 0.5 | 0.8 |
Sample size per group for a two-sample t test
Two-sided test at α = 0.05 with equal group sizes. These figures match G*Power and R’s pwr.t.test, and the calculator above reproduces every row.
| Cohen's d | Power 0.80 | Power 0.90 | Power 0.95 |
|---|---|---|---|
| 0.2 | 394 | 527 | 651 |
| 0.3 | 176 | 235 | 290 |
| 0.4 | 100 | 133 | 164 |
| 0.5 | 64 | 86 | 105 |
| 0.6 | 45 | 60 | 74 |
| 0.8 | 26 | 34 | 42 |
| 1 | 17 | 23 | 27 |
| 1.2 | 12 | 16 | 20 |
Sample size per group for one-way ANOVA
Equal group sizes at α = 0.05; Cohen’s f of 0.10, 0.25 and 0.40 are the small, medium and large conventions.
| Cohen's f | Groups (k) | Power 0.80 | Power 0.90 |
|---|---|---|---|
| 0.1 | 3 | 323 | 423 |
| 0.1 | 4 | 274 | 356 |
| 0.1 | 5 | 240 | 310 |
| 0.25 | 3 | 53 | 69 |
| 0.25 | 4 | 45 | 58 |
| 0.25 | 5 | 40 | 51 |
| 0.4 | 3 | 22 | 28 |
| 0.4 | 4 | 19 | 24 |
| 0.4 | 5 | 16 | 21 |
Sample size for detecting a correlation
Total number of pairs for a two-sided test of ρ = 0 at α = 0.05. The values use Fisher’s z approximation, which can differ by one from an exact calculation.
| Correlation r | Power 0.80 | Power 0.90 |
|---|---|---|
| 0.1 | 783 | 1047 |
| 0.2 | 194 | 259 |
| 0.3 | 85 | 113 |
| 0.4 | 47 | 62 |
| 0.5 | 30 | 38 |
Worked example: comparing two groups
You expect a medium difference, d = 0.5, between two independent groups and will use a two-sided t test at α = 0.05. With 64 people per group, δ = 0.5·√(64·64 / 128) = 2.8284 and the test has 126 degrees of freedom. The critical value is t = 1.979, and the probability that a noncentral t variable with δ = 2.8284 lands beyond it is 0.8015. Solving for the sample size returns 64 per group; with 63 per group the power is 0.7952, just under the 0.80 target. Choose “Load example” to see the full working and the power curve.
Assumptions and limits
- The t-test formulas assume normally distributed data and, for two groups, equal variances (the pooled test). With unequal variances and unequal group sizes the Welch test has somewhat different power.
- The two-proportion and correlation results use normal approximations (the correlation via Fisher’s z), so they can differ from exact-test software by a small amount at small sample sizes.
- Power for the Mann–Whitney U test is close to the t test under normality: divide the t-test sample size by 0.955 to plan for it.
- Do not compute “observed power” from the effect you measured after the study; it is a function of the p-value and adds no information. Report the confidence interval instead.
- Add extra participants for expected dropout: divide the calculated n by the share you expect to keep.
Software equivalents
| Software | Two-sample t test, d = 0.5, power 0.8, α = 0.05 |
|---|---|
| G*Power | t tests → Means: Difference between two independent means → A priori: effect size d = 0.5 |
| R (stats) | power.t.test(delta = 0.5, sd = 1, sig.level = 0.05, power = 0.8) |
| R (pwr) | pwr.t.test(d = 0.5, sig.level = 0.05, power = 0.8, type = "two.sample") |
| Python | statsmodels.stats.power.TTestIndPower().solve_power(effect_size=0.5, alpha=0.05, power=0.8) |
| Stata | power twomeans 0 0.5, power(0.8) |
Frequently Asked Questions
What is a good power level for a study?
Power of 0.80 is the usual minimum: it means a 20% chance of missing an effect of the planned size. Confirmatory and clinical studies often plan for 0.90 or 0.95, which needs about a third and two-thirds more participants than 0.80 for the same effect (64, 86 and 105 per group at d = 0.5).
How do I choose the effect size?
Use the smallest effect that would matter in practice, an estimate from a pilot study, or the effect reported in similar research, and check the calculator with a smaller value as a safety margin. Cohen's small, medium and large conventions (d = 0.2, 0.5, 0.8) are a fallback when nothing else is known, not a substitute for subject knowledge.
Why does the sample size grow so fast for small effects?
Sample size is inversely proportional to the square of the effect size. Halving the effect from d = 0.5 to d = 0.25 multiplies the required n by about four (from 64 to 253 per group at 80% power).
What is the difference between one-sided and two-sided power?
A one-sided test puts all of α in one tail, so it needs fewer participants for an effect in the predicted direction but cannot detect an effect in the other direction. Choose the alternative before collecting data and enter the effect as a positive number.
Do I enter n per group or the total sample size?
For the two-sample t test and two proportions, enter the size of each group; the results also list the total. For a one-sample or paired test enter the number of subjects or pairs, for ANOVA the size of each group, and for a correlation the number of pairs.
How is the paired t test different from the independent one?
The paired test analyses the within-pair differences, so the effect size is the mean difference divided by the standard deviation of those differences, and the sample size is the number of pairs. Strongly correlated pairs shrink that standard deviation and raise power.
Why might my result differ slightly from G*Power, R or statsmodels?
The t-test and ANOVA results use the exact noncentral t and F distributions and agree with those programs to many decimals. Two-proportion and correlation results use normal approximations, and different programs use different variants (pooled or unpooled variance, Fisher's z or an exact bivariate normal model), which can shift the sample size by one or a few.
What is a sensitivity analysis?
It answers the reverse question: given the sample size you can actually afford, what is the smallest effect you would reliably detect? Choose 'Minimum detectable effect' to get it, and state that value when you report a study that did not find a significant effect.
Does the calculator handle multiple comparisons?
Not directly. To plan for several tests, divide α by the number of comparisons (Bonferroni) before entering it; the required sample size then rises accordingly.
Can I use it for the Welch test or a nonparametric test?
The two-sample result assumes equal variances. Welch's test loses little power when group sizes are equal. For the Mann–Whitney U or Wilcoxon signed-rank test under roughly normal data, divide the t-test sample size by 0.955 as a planning rule.
Embed This Calculator
Add this free calculator to your course page or LMS.
Adjust the height value to fit your page.