Bonferroni Correction Calculator
Enter raw p-values to get Bonferroni, Holm, Šidák, Hochberg, Benjamini-Hochberg and Benjamini-Yekutieli adjusted p-values and reject or keep decisions at your α. Or leave the p-values empty and enter the number of tests to get the corrected α.
After an ANOVA, comparing every pair of group means is what Tukey HSD is built for. Use this page for p-values from separate tests.
Comma, space or new-line separated, each between 0 and 1 (1e-5 is fine). Leave empty to get only the corrected α for a number of tests.
Only needed when you have no p-values.
Related Calculators
Tukey HSD Calculator
Pairwise comparisons after ANOVA with Tukey HSD or Tukey-Kramer from raw data or an ANOVA table: differences, q statistics, adjusted p-values, and simultaneous confidence intervals.
P-Value Calculator
Find the one- or two-tailed p-value of a z, t, chi-square or F test statistic and the decision at your significance level.
ANOVA Calculator
Run one-way ANOVA calculations and inspect variance across groups.
What the Bonferroni correction does
Every hypothesis test has a small chance of a false positive. When one question is tested several times, those chances add up: with m independent tests, each run at α = 0.05, the probability of at least one false positive is 1 − (1 − α)m. That is about 22.6% for 5 tests and 64.2% for 20 tests. This probability is the family-wise error rate (FWER).
The Bonferroni correction keeps the FWER at or below α. It rests on the union bound (Boole's inequality): the chance that at least one of m true null hypotheses is rejected cannot exceed the sum of the m individual chances. Test each hypothesis at level α / m, or equivalently multiply every p-value by m (capped at 1) and compare it with α. The two give the same decisions, and neither needs the tests to be independent.
| Tests (m) | Chance of at least one false positive at α = 0.05 with no correction | Bonferroni threshold α/m |
|---|---|---|
| 1 | 5% | 0.05 |
| 3 | 14.26% | 0.016667 |
| 5 | 22.62% | 0.01 |
| 10 | 40.13% | 0.005 |
| 20 | 64.15% | 0.0025 |
| 50 | 92.31% | 0.001 |
| 100 | 99.41% | 0.0005 |
How to use the calculator
- Paste the raw p-values of your tests, separated by commas, spaces or new lines. Use the p-value your software reports (0.00031 or 3.1e-4), not a bound such as < 0.001.
- Set the family-wise α, usually 0.05. For the two Benjamini methods this number is the false discovery rate you accept (often written q).
- Press Adjust p-values. Read the corrected thresholds, the tests each method rejects and the table of adjusted p-values. A test is rejected when its adjusted p-value is at most α.
- Only need the corrected α? Leave the p-values empty and type the number of tests: you get α / m (Bonferroni), the Šidák level and the uncorrected FWER.
Formulas
With m tests and sorted p-values p(1) ≤ p(2) ≤ … ≤ p(m), every adjusted p-value is capped at 1:
Bonferroni: adjusted p = m · p, reject when p ≤ α / m
Šidák: adjusted p = 1 − (1 − p)^m, reject when p ≤ 1 − (1 − α)^(1/m)
Holm: adjusted p(i) = max over j ≤ i of (m − j + 1) · p(j)
Hochberg: adjusted p(i) = min over j ≥ i of (m − j + 1) · p(j)
Benjamini-Hochberg: adjusted p(i) = min over j ≥ i of m · p(j) / j
Benjamini-Yekutieli: adjusted p(i) = min over j ≥ i of c(m) · m · p(j) / j, c(m) = 1 + 1/2 + … + 1/m
Uncorrected family-wise error rate: 1 − (1 − α)^m
- Holm (step-down) compares the smallest p-value with α / m, the next with α / (m − 1), and so on, and stops at the first p-value that fails.
- Hochberg (step-up) starts from the largest p-value, finds the first one that is at most α / (m − i + 1) for its rank i counted from the smallest, and rejects it and every smaller p-value.
- Benjamini-Hochberg finds the largest rank i with p(i) ≤ (i / m) · α and rejects the i smallest p-values.
How to read the results
- Bonferroni threshold (α/m) is the significance level each single test must beat. The Šidák threshold is its slightly larger, exact counterpart for independent tests.
- FWER if uncorrected is the chance of at least one false positive among m independent tests if you ignored the problem and used α for each.
- Tests rejected at α lists, for each method, how many tests are still significant after the correction and which ones (numbered in the order you entered the p-values).
- Adjusted p-values can be compared directly with α, whatever the method. An adjusted p-value of 1 means the correction pushed it to the maximum, not that the effect is proven absent.
Worked example: five p-values
A study measures five outcomes and gets the p-values 0.003, 0.012, 0.021, 0.048 and 0.31. Four of them are below 0.05, but with five tests the chance of at least one false positive is 22.62%, so the correction has to be applied. Press Load example to enter them.
| Test | Raw p | Bonferroni | Holm | Šidák | Hochberg | BH | BY |
|---|---|---|---|---|---|---|---|
| 1 | 0.003 | 0.015 | 0.015 | 0.01491 | 0.015 | 0.015 | 0.03425 |
| 2 | 0.012 | 0.06 | 0.048 | 0.058577 | 0.048 | 0.03 | 0.0685 |
| 3 | 0.021 | 0.105 | 0.063 | 0.100682 | 0.063 | 0.035 | 0.079917 |
| 4 | 0.048 | 0.24 | 0.096 | 0.21804 | 0.096 | 0.06 | 0.137 |
| 5 | 0.31 | 1 | 0.31 | 0.843597 | 0.31 | 0.31 | 0.707833 |
At α = 0.05 the methods disagree in the way the theory predicts. Bonferroni (threshold 0.05 / 5 = 0.01) keeps only test 1. Holm keeps tests 1 and 2. Benjamini-Hochberg, which controls the false discovery rate instead of the family-wise error rate, keeps tests 1, 2 and 3.
Working for Holm and Benjamini-Hochberg
Holm compares the sorted p-values with 0.01, 0.0125, 0.016667, 0.025 and 0.05.
0.003 ≤ 0.01 and 0.012 ≤ 0.0125 pass, but 0.021 is above 0.016667, so Holm stops after two rejections.
Benjamini-Hochberg compares them with 0.01, 0.02, 0.03, 0.04 and 0.05.
0.003, 0.012 and 0.021 pass; 0.048 is above 0.04 and 0.31 above 0.05, so three tests are rejected.
Which correction should I use?
| Method | Controls | Valid when | Power |
|---|---|---|---|
| Bonferroni | FWER | Any dependence between the tests | Lowest of the FWER methods |
| Šidák | FWER | Independent tests (exact) and positively dependent tests (conservative) | Slightly above Bonferroni |
| Holm | FWER | Any dependence between the tests | Never below Bonferroni |
| Hochberg | FWER | Independent or non-negatively associated tests | At least Holm |
| Benjamini-Hochberg | FDR | Independent or positively dependent tests | At least Hochberg, thanks to the weaker FDR guarantee |
| Benjamini-Yekutieli | FDR | Any dependence between the tests | Never above BH; can fall below Bonferroni |
Use an FWER method (Bonferroni, Holm, Šidák, Hochberg) when a single false positive is costly, as in a confirmatory study with a handful of pre-specified comparisons. Use an FDR method when you screen many hypotheses, such as thousands of genes, and can live with a small share of false discoveries in exchange for power. R's documentation of p.adjust notes that the unmodified Bonferroni correction is dominated by Holm's method, which is valid under the same arbitrary dependence, so Holm is the better default when you only need FWER control. Bonferroni remains the simplest rule to state, apply by hand and use for simultaneous confidence intervals (build each one at level 1 − α / m).
Assumptions and pitfalls
- Fix the family before you look. m counts every test that addresses the same question, including those you did not report. Choosing which tests to count after seeing which are significant defeats the correction.
- Bonferroni is conservative. With many tests, or tests that are strongly correlated, it rejects too little: at m = 100 each test must reach 0.0005. Holm or an FDR method keeps more power.
- Mind the dependence. Šidák, Hochberg and Benjamini-Hochberg are derived for independent or positively dependent tests. When the dependence is unknown or possibly negative, use Holm or Benjamini-Yekutieli.
- It does not repair the tests. The correction assumes each p-value is valid. It cannot fix a wrong test, biased data or selective reporting.
- One primary test needs none. A single pre-specified hypothesis is tested at α. Secondary endpoints can be handled with a fixed testing order instead.
- After an ANOVA with every pair of means compared, use Tukey HSD, which exploits the structure of the comparisons and is usually more powerful than Bonferroni. For factorial designs see the two-way ANOVA calculator.
For the reasoning behind these rules read multiple comparisons explained. A p-value of a single test comes from the p-value calculator, and the chance of missing a real effect with a stricter α is the topic of statistical power explained.
Bonferroni correction in Excel, R, Python and SPSS
| Software | How to adjust the p-values |
|---|---|
| Excel | =MIN(1, A2*COUNT($A$2:$A$6)) for p-values in A2:A6, copied down the column |
| R | p.adjust(p, method = "bonferroni"); other methods: "holm", "hochberg", "hommel", "BH", "BY" (also called "fdr"), "none" |
| Python | statsmodels.stats.multitest.multipletests(pvals, alpha=0.05, method="bonferroni") returns reject, pvals_corrected, alphacSidak, alphacBonf; other methods: "holm", "sidak", "simes-hochberg", "fdr_bh", "fdr_by" |
| SPSS | Analyze > General Linear Model > Univariate > Post Hoc > Bonferroni; or Options > Compare main effects with the confidence interval adjustment set to Bonferroni |
R also offers Hommel's method and statsmodels offers Holm-Šidák and two-stage FDR procedures; this calculator covers the six methods above and matches statsmodels' multipletests for each of them.
Frequently Asked Questions
What is the Bonferroni correction?
A way to keep the family-wise error rate at or below α when you run m tests: test each hypothesis at level α / m, or multiply each p-value by m and compare it with α. It needs no assumption about how the tests depend on each other.
How do I calculate the Bonferroni corrected p-value?
Multiply the raw p-value by the number of tests and cap the result at 1. With 5 tests a raw p of 0.003 becomes 0.015 and a raw p of 0.012 becomes 0.06, so only the first is significant at 0.05.
How do I calculate the Bonferroni corrected alpha?
Divide α by the number of tests: 0.05 / 10 = 0.005. Reject a hypothesis only when its raw p-value is at most that. Leave the p-values empty in the calculator and enter the number of tests to get the value, together with the Šidák level.
Why is Bonferroni conservative?
It guards against the worst case in which the events of a false positive never overlap. Positively correlated tests overlap a lot, so the real error rate is lower than α and the method rejects too little. The more tests, the stricter the threshold: 0.0005 for 100 tests at α = 0.05.
Bonferroni or Holm?
Holm gives the same FWER guarantee under any dependence and never rejects fewer hypotheses than Bonferroni. In the worked example Holm keeps two tests where Bonferroni keeps one. Bonferroni is still handy for a quick rule and for simultaneous confidence intervals.
When should I use Benjamini-Hochberg instead?
When you test many hypotheses, for example thousands of genes or many outcomes, and accept a small share of false discoveries in return for more power. It controls the false discovery rate, the expected proportion of false positives among the rejections, not the chance of any false positive.
What is the difference between FWER and FDR?
The family-wise error rate is the probability of at least one false rejection. The false discovery rate is the expected fraction of false rejections among all rejections. FDR is the weaker condition, so FDR methods reject more hypotheses.
Should I use Bonferroni after an ANOVA?
For all pairwise comparisons of group means use Tukey HSD, which is designed for that family and is usually more powerful. Bonferroni or Holm suit a small planned set of comparisons, or p-values that come from different tests.
How many tests should I count in m?
All tests that address the same question, decided before you see the results, including the ones that are not significant and any you ran but did not report. If you enter your p-values here, m is the number of p-values.
Can an adjusted p-value be larger than 1?
No. The adjusted p-value is capped at 1. A capped value only says the correction consumed the evidence, not that the effect is zero.
Can I enter p-values such as 1e-5 or < 0.001?
Scientific notation such as 1e-5 or 3.1e-4 works. A bound such as < 0.001 does not: use the exact p-value from your software, because the adjusted values cannot be computed from a bound.
Does the order of my p-values matter, and what about ties?
No. Results are listed in the order you entered the p-values. Holm, Hochberg and the FDR methods sort them internally with a stable order, and tied p-values get the same adjusted value as in statsmodels multipletests.
Embed This Calculator
Add this free calculator to your course page or LMS.
Adjust the height value to fit your page.