Statistics Concepts

Statistical Power Explained

Statistical power is the probability that a test detects an effect of a given size when that effect really exists: power = 1 − β, where β is the chance of a Type II error. The convention is 80%. Power rises with the effect size and the sample size and falls as α is made stricter. For a medium effect (d = 0.5) comparing two groups at α = 0.05, 64 people per group give 80% power, while 20 per group give only 34%.

What power depends on

  • Effect size. Larger true effects are easier to detect. See effect size explained.
  • Sample size. More observations shrink the standard error, so the same effect stands out more clearly.
  • Significance level α. A stricter α means a higher bar for significance and so lower power.
  • Variability. Noisier data hide effects; a paired design or better measurements reduce noise.
  • One-sided or two-sided. A justified one-sided test has more power in its direction; see one-tailed vs two-tailed tests.

Power for a medium effect as the sample grows

Two-sample t-test, two-sided, α = 0.05, true effect d = 0.5:

People per groupPower
2033.8%
3047.8%
5069.7%
6480.1%
10094.0%

With 20 per group the study would miss a real medium effect about two times in three. That is the practical meaning of low power.

People per group for 80% and 90% power

True effect (Cohen's d)80% power90% power
0.2 (small)394527
0.5 (medium)6486
0.8 (large)2634

Halving the effect size roughly quadruples the sample needed, since the standard error shrinks only with the square root of n. Other designs at 80% power for d = 0.5: α = 0.01 needs 96 per group, a one-sided test 51 per group, and a paired design 34 pairs when d = 0.5 is measured on the within-pair differences. For a one-way ANOVA with 4 groups and a medium effect (f = 0.25), you need 179 people in total, which rounds up to 45 per group. Try your own numbers in the statistical power calculator.

Three kinds of power analysis

  • A priori: before the study, find the sample size that gives the target power for the smallest effect worth detecting. This is the useful one.
  • Sensitivity: given a fixed sample, find the smallest effect you can detect with 80% power. With 30 people per group that is d = 0.74.
  • Post hoc: power computed from the observed effect. It only restates the p-value, so avoid it; report a confidence interval instead.

Power is also the reason a non-significant result is hard to interpret: the test may simply have been too small. The two error types are covered in type 1 and type 2 errors, and the sample-size formulas in sample size explained.

Consequences of low power

  • Real effects are missed, and a negative result is uninformative.
  • The significant results that do appear tend to overstate the effect, because only unusually large sample estimates cross the significance line.
  • Small studies cost participants and time without answering the question, which is an ethical issue as well as a statistical one.

Try the Statistical Power Calculator

Compute power, or the sample size for a target power, for t-tests, proportions and ANOVA.

Try the A/B Test Sample Size Calculator

Find how many visitors per variant you need to detect a given lift.

Frequently Asked Questions

What is a good level of statistical power?

80% is the usual minimum, which accepts a 20% chance of missing a real effect of the size you planned for. Studies that will inform expensive or safety-critical decisions often aim for 90%. Going from 80% to 90% power for a medium effect raises the requirement from 64 to 86 people per group.

How are power, β and α related?

Power = 1 − β, where β is the probability of a Type II error, missing a real effect. α is the probability of a Type I error, a false positive. Lowering α makes the test stricter and so reduces power unless the sample grows: at α = 0.01 a medium effect needs 96 people per group for 80% power, against 64 at α = 0.05.

Why is post hoc (observed) power not useful?

Observed power is computed from the effect size you already measured, so it is a direct function of the p-value: a non-significant result always has low observed power and a significant one high, and it adds no information. Plan the power before collecting data, and afterwards report the confidence interval to show which effects the data can rule out.

How can I raise power without recruiting more people?

Reduce the noise: measure more precisely, control nuisance variables, or use a paired or repeated-measures design, which strips out differences between subjects. A justified one-sided test cuts the requirement for a medium effect from 64 to 51 per group. Raising α is also possible but trades power for more false positives.

Does a non-significant result with high power prove there is no effect?

It shows that an effect as large as the one you powered for is unlikely to exist, not that the effect is exactly zero. To argue for no meaningful difference, use an equivalence test or show that the confidence interval lies wholly inside a range you consider negligible.

What effect size should I use for a power analysis?

Use the smallest effect that would matter in practice, or an estimate from earlier studies, adjusted downwards because published estimates tend to be inflated. Cohen's 0.2, 0.5 and 0.8 are a last resort. With 30 people per group and 80% power you can only reliably detect a d of 0.74 or more.