Statistics Concepts
Multiple Comparisons Explained: Bonferroni, Holm and FDR
Each test at α = 0.05 has a 5% chance of a false positive when there is no effect. Run several and the chances add up: 3 independent tests give a 14.26% chance of at least one false positive, and 10 tests give 40.13%. Multiple-comparison corrections restore control, either of the family-wise error rate (Bonferroni, Holm, Šidák, Tukey) or of the false discovery rate (Benjamini-Hochberg).
How fast false positives accumulate
| Number of tests | Chance of at least one false positive at α = 0.05 | Bonferroni per-test α | Šidák per-test α |
|---|---|---|---|
| 1 | 5.00% | 0.05000 | 0.05000 |
| 3 | 14.26% | 0.01667 | 0.01695 |
| 5 | 22.62% | 0.01000 | 0.01021 |
| 10 | 40.13% | 0.00500 | 0.00512 |
| 20 | 64.15% | 0.00250 | 0.00256 |
| 100 | 99.41% | 0.00050 | 0.00051 |
The second column is 1 − 0.95^m for independent tests. The per-test α in the last two columns is what each test must beat so that the family-wise error rate stays at 0.05: Bonferroni uses α/m and Šidák uses 1 − (1 − α)^(1/m), which is very slightly less strict.
The main corrections
| Method | Controls | How it works | Best used for |
|---|---|---|---|
| Bonferroni | Family-wise error | Multiply each p-value by m (cap at 1) | A small number of tests; easy to explain |
| Holm | Family-wise error | Step-down: multiply the smallest p by m, the next by m − 1, and so on | The default choice; never less powerful than Bonferroni |
| Šidák | Family-wise error | 1 − (1 − p)^m | Independent tests |
| Tukey HSD | Family-wise error | Studentized range distribution over all pairs | All pairwise comparisons after ANOVA |
| Benjamini-Hochberg | False discovery rate | Step-up: p × m / rank, then made non-decreasing | Many tests, exploratory work |
Worked example: eight p-values
A study tests eight outcomes and gets these raw p-values. Five of them are below 0.05, but how many survive adjustment?
| Rank | Raw p | Bonferroni | Holm | Benjamini-Hochberg |
|---|---|---|---|---|
| 1 | 0.001 | 0.008 | 0.008 | 0.008 |
| 2 | 0.008 | 0.064 | 0.056 | 0.032 |
| 3 | 0.039 | 0.312 | 0.234 | 0.0672 |
| 4 | 0.041 | 0.328 | 0.234 | 0.0672 |
| 5 | 0.042 | 0.336 | 0.234 | 0.0672 |
| 6 | 0.060 | 0.480 | 0.234 | 0.0800 |
| 7 | 0.074 | 0.592 | 0.234 | 0.0846 |
| 8 | 0.205 | 1.000 | 0.234 | 0.2050 |
The adjusted p-values in the last three columns are compared with 0.05. Without correction 5 of 8 results are significant. Bonferroni and Holm keep only the first (0.008); Benjamini-Hochberg keeps two (0.008 and 0.032), because it allows a controlled share of false discoveries. Enter your own p-values in the Bonferroni correction calculator, which also offers Šidák and Hochberg.
Doing it in software
| Tool | Command |
|---|---|
| Excel or Google Sheets | =MIN(1,B2*8) for Bonferroni with 8 tests, filled down the column |
| R | p.adjust(p, method = "holm") (also "bonferroni", "BH") |
| Python (statsmodels) | multipletests(p, method='holm') (also 'bonferroni', 'fdr_bh') |
Where multiple comparisons hide
- Several outcomes: measuring five outcomes and reporting the one that reached p < 0.05.
- Subgroup analyses: splitting by age, sex and region after seeing an overall null result.
- Peeking: checking an A/B test every day and stopping when it looks significant.
- Pairwise follow-ups: comparing every pair of k groups, which gives k(k − 1)/2 tests. Four groups mean 6 tests; use Tukey HSD after the ANOVA.
The remedy is to decide the analysis in advance, correct across the family, and treat everything else as exploratory. See also type 1 and type 2 errors and which statistical test to use.
Try the Bonferroni Correction Calculator
Adjust a list of p-values with Bonferroni, Holm, Šidák, Hochberg, Benjamini-Hochberg or Benjamini-Yekutieli.
Try the Tukey HSD Calculator
Compare every pair of group means after an ANOVA while controlling the family-wise error rate.
Frequently Asked Questions
What is the family-wise error rate?
The probability of making at least one Type I error (false positive) across a whole set of tests. With m independent tests at level α it is 1 − (1 − α)^m, so 3 tests at 0.05 give 14.26%. Methods such as Bonferroni, Holm and Tukey's HSD are designed to hold this rate at α.
Should I use Bonferroni or Holm?
Holm. It controls the family-wise error rate just as strictly as Bonferroni without any extra assumptions, and it can never reject fewer hypotheses. Bonferroni is only easier to state, because it divides α by the number of tests, so it is worth knowing but rarely worth choosing.
What is the false discovery rate?
The expected share of false positives among the results you declare significant. Controlling it at 5% means about 1 in 20 of your discoveries may be wrong, which is a weaker guarantee than family-wise control but gives more power when you run hundreds of tests, as in genomics or exploratory screening. Benjamini-Hochberg is valid for independent or positively dependent tests; Benjamini-Yekutieli works under any dependence at the cost of power.
Do I need a correction after a significant ANOVA?
For the omnibus F test itself, no. But if you then compare the groups two at a time, those pairwise tests are a family and need a post hoc procedure such as Tukey's HSD, or Bonferroni or Holm applied to the pairwise p-values. Otherwise the pairwise comparisons carry the inflated error rate.
When is a correction not needed?
When there is a single pre-registered primary comparison, or a small number of planned contrasts that answer distinct questions. Any exploratory result should still be labelled as exploratory, and ideally confirmed in new data, since the correction only matters for the tests you decided on before looking.
What counts as a family of tests?
All the tests that address one scientific question and would be reported together, decided before the analysis. Testing five outcomes, or three subgroups, or checking an A/B test every day until it turns significant, all create a family. Peeking at results repeatedly is a multiple-comparisons problem even though only one test is run at a time.