Tukey HSD Calculator
Pairwise comparisons after one-way ANOVA with Tukey HSD (equal n) or Tukey-Kramer (unequal n). Enter raw data, or the group means, sizes, MSE and error df from an ANOVA table, and get the critical q, the HSD, and for every pair the difference, standard error, q, adjusted p-value and simultaneous confidence interval.
Run the omnibus test first with the one-way ANOVA calculator. Check the equal-variance assumption with the Levene test calculator. Ranks or clearly non-normal data: the Kruskal-Wallis test.
Numbers separated by commas, spaces, or new lines (at least 2 per group).
Numbers separated by commas, spaces, or new lines (at least 2 per group).
Numbers separated by commas, spaces, or new lines (at least 2 per group).
Related Calculators
ANOVA Calculator
Run one-way ANOVA calculations and inspect variance across groups.
Bonferroni Correction Calculator
Adjust multiple p-values with Bonferroni, Holm, Šidák, Hochberg, Benjamini-Hochberg, or Benjamini-Yekutieli, or find the corrected alpha for m tests.
Levene's Test Calculator
Test equal variances across groups with Levene (mean), Brown-Forsythe (median) or a 10% trimmed mean: W, p-value, F critical value, group SDs and the ANOVA on absolute deviations.
What Tukey's HSD test does
Tukey's honestly significant difference (HSD) test compares every pair of group means after a one-way ANOVA and keeps the chance of at least one false positive across all the pairs at α. Testing each pair with its own t-test would let that chance grow with the number of pairs. Tukey's test instead refers each difference to the studentized range distribution, the distribution of the gap between the largest and the smallest of k means measured in standard errors, so the whole family of comparisons holds its level exactly when the groups have equal sizes. With unequal sizes the same formula is called the Tukey-Kramer method. It errs on the safe side: Hayter (1984) proved that its error rate never exceeds α.
Use it when the group variances are similar and you want to know which means differ, usually after a significant F in the one-way ANOVA calculator. The reasoning behind adjusted p-values is in multiple comparisons explained, and ANOVA explained covers the test that comes first.
How to use the Tukey HSD calculator
- Raw data: enter each group's observations separated by commas, spaces or new lines, at least 2 per group. Add another group gives you up to 10 groups. An empty group is left out and the others keep their own numbers in the results.
- From an ANOVA table: choose Means + MSE and df (from ANOVA) and enter the group means, the sample sizes, the within-group mean square (MSE, called Mean Square Within or Error in the output) and its degrees of freedom N − k. This also works for the marginal means of a two-way ANOVA.
- Set the family-wise significance level α, usually 0.05. It fixes the critical q and the confidence level 1 − α of the intervals.
- Press Calculate Tukey HSD. Read the critical q, the HSD (only when all groups have the same size) and the table with the difference, standard error, q, adjusted p-value and interval of every pair.
- Load example enters the PlantGrowth data used below in the current input mode, and Copy link to this calculation shares your exact analysis.
Tukey HSD formulas
q = |x̄ᵢ − x̄ⱼ| / √( (MSE / 2) · (1/nᵢ + 1/nⱼ) )
p = P( Q(k, N − k) ≥ q ), Q = studentized range with k groups and N − k df
HSD = q(α; k, N − k) · √( MSE / n ), for equal group sizes n
CI: (x̄ᵢ − x̄ⱼ) ± q(α; k, N − k) · √( (MSE / 2) · (1/nᵢ + 1/nⱼ) )
SE of a difference = √( MSE · (1/nᵢ + 1/nⱼ) ), so q = √2 · |x̄ᵢ − x̄ⱼ| / SE
MSE = SSW / (N − k), the within-group mean square of the one-way ANOVA
Significant at family-wise α ⇔ p < α ⇔ |x̄ᵢ − x̄ⱼ| > HSD (equal n) ⇔ the interval excludes 0
Two groups: q = √2 · |t|, so Tukey's test is the pooled two-sample t-test
The critical value q(α; k, N − k) is the upper α point of the studentized range distribution. The calculator computes it, and every p-value, by numerical integration of that distribution (relative error below 1e-12, checked against 30-digit reference values and SciPy) instead of reading a printed q table, so any number of groups up to 50 and any error df from 2 to 10 million work.
How to read the results
- Critical q and HSD. With equal group sizes any two means further apart than the HSD are significantly different at the family-wise level; with unequal sizes each pair has its own threshold and the HSD card shows n/a.
- Diff. The mean of the first group minus the mean of the second (the SciPy and SPSS convention). R's TukeyHSD subtracts the earlier group from the later one, so its differences have the opposite sign.
- SE and q. SE is the standard error of the difference and q = √2 · |Diff| / SE is the studentized range statistic for the pair.
- Adj. p. The probability of a gap at least this large between two of the k means when all population means are equal. It is already adjusted for the number of pairs, so compare it with α directly.
- Confidence interval. A simultaneous interval: all the pair intervals together cover the true differences with probability 1 − α, so each is wider than an ordinary t interval. It excludes 0 exactly when the pair is significant.
- Not significant is not equal. A pair whose interval contains 0 is compatible with no difference; with small groups it is compatible with a large one too.
Worked example by hand: three groups of three
The groups 4, 5, 6 (group 1), 6, 7, 8 (group 2) and 8, 9, 10 (group 3) are what the calculator holds when the page opens. The means are 5, 7 and 9, each group has n = 3, and the within-group sum of squares is 2 + 2 + 2 = 6 on 9 − 3 = 6 degrees of freedom, so MSE = 1.
Working
Denominator of q: √(MSE / 2 · (1/3 + 1/3)) = √(1/3) = 0.5774.
Critical value q(0.05; 3, 6) = 4.3392, so HSD = 4.3392 × 0.5774 = 2.5052.
The differences are −2 for 1 vs 2, −4 for 1 vs 3 and −2 for 2 vs 3. Only |−4| exceeds 2.5052, so only groups 1 and 3 differ.
q = |difference| / 0.5774 = 3.4641 for 1 vs 2 and for 2 vs 3, and 6.9282 for 1 vs 3. With k = 3 and 6 degrees of freedom, q = 3.4641 gives p = 0.1089 and q = 6.9282 gives p = 0.0065.
The intervals are the difference ± 2.5052: (−4.5052, 0.5052) for 1 vs 2, (−6.5052, −1.4948) for 1 vs 3 and (−4.5052, 0.5052) for 2 vs 3. Only the interval for 1 vs 3 excludes 0.
Worked example with real data: PlantGrowth
The PlantGrowth data set that ships with R (Dobson, 1983) holds the dry weight of 30 plants, ten controls and ten for each of two treatments. The one-way ANOVA calculator gives F(2, 27) = 4.846 and p = 0.0159 with MSE = 0.3886 on 27 degrees of freedom, so the means differ somewhere. Press Load example to enter the weights, or switch to the ANOVA-table input, where the example holds the means 5.032, 4.661 and 5.526, n = 10, MSE = 0.3886 and df = 27.
The critical value is q(0.05; 3, 27) = 3.5064, so HSD = 3.5064 × √(0.3886 / 10) = 0.6912. The three comparisons:
| Pair | Diff | q | Adj. p | 95% CI | Significant |
|---|---|---|---|---|---|
| 1 vs 2 | 0.371 | 1.882 | 0.390871 | (-0.3202, 1.0622) | No |
| 1 vs 3 | -0.494 | 2.506 | 0.197996 | (-1.1852, 0.1972) | No |
| 2 vs 3 | -0.865 | 4.388 | 0.012006 | (-1.5562, -0.1738) | Yes |
Only treatment 1 against treatment 2 (groups 2 and 3) is significant: its gap of 0.865 exceeds the HSD of 0.6912, its adjusted p-value is 0.012 and its interval excludes 0. The control differs from neither treatment at α = 0.05 with ten plants per group. R's TukeyHSD(aov(weight ~ group, data = PlantGrowth)) reports the same comparisons as trt1-ctrl = −0.371, trt2-ctrl = 0.494 and trt2-trt1 = 0.865 with adjusted p = 0.3908711, 0.1979960 and 0.0120064: R subtracts the earlier group from the later one, so its signs are the opposite of the Diff column and its lower and upper limits are the negated upper and lower limits here. Typed from the rounded ANOVA table, the example differs from the raw-data result only in the sixth decimal of the p-values.
Unequal group sizes: the Tukey-Kramer version
With unequal sample sizes the standard error of each pair uses its own two sizes, and there is no single HSD. Take four groups from an ANOVA table with means 10, 12, 15 and 11, sizes 6, 8, 5 and 7, MSE = 4 and 22 error degrees of freedom. The critical value is q(0.05; 4, 22) = 3.927. Three of the six comparisons:
| Pair | Diff | SE | q | Adj. p | 95% CI | Significant |
|---|---|---|---|---|---|---|
| 1 vs 2 | -2 | 1.0801 | 2.6186 | 0.277139 | (-4.9993, 0.9993) | No |
| 1 vs 3 | -5 | 1.2111 | 5.8387 | 0.002312 | (-8.3629, -1.6371) | Yes |
| 3 vs 4 | 4 | 1.1711 | 4.8305 | 0.012277 | (0.7481, 7.2519) | Yes |
Groups 1 and 2 are two units apart and not significantly different, while the larger gaps between groups 1 and 3 (5 units) and between groups 3 and 4 (4 units) are. Hayter (1984) proved that the Tukey-Kramer intervals are conservative for every choice of sizes: the simultaneous coverage is at least 1 − α, and its minimum, exactly 1 − α, occurs when the sizes are equal.
Tukey HSD compared with other post-hoc tests
| Test | Best for | Compared with Tukey HSD |
|---|---|---|
| Bonferroni or Holm | A few planned comparisons, or p-values from any tests | Valid for any set of tests but usually less powerful when every pair among many groups is compared |
| Scheffé | All possible contrasts, not only pairs | More conservative for pairwise comparisons |
| Dunnett | Several treatments against one control | More powerful for that question, because it makes fewer comparisons |
| Games-Howell | Unequal variances and unequal sizes | Does not pool the variance. Not offered here; SPSS, the rstatix package in R and recent SciPy releases have it |
| Fisher's LSD | Three groups, after a significant F | No adjustment for the number of pairs; with more than three groups the family-wise error can exceed α |
| Dunn's test | Ranks or clearly non-normal data | The nonparametric follow-up of the Kruskal-Wallis test |
For a chosen set of p-values from any source use the Bonferroni correction calculator. The Kruskal-Wallis test is the rank-based alternative to the ANOVA that comes before Tukey's test: Kruskal-Wallis test.
Assumptions of Tukey HSD
- Independent observations within and between the groups. This comes from the design.
- Roughly normal values within each group. Moderate departures do little harm; check with the Shapiro-Wilk test or a histogram.
- Equal variances. The MSE pools all the groups. With equal sizes, a group with a much larger spread makes the comparisons that involve it too liberal and the comparisons among the other groups too conservative; with unequal sizes the direction depends on which groups are larger. Test the spreads with the Levene test calculator and switch to Games-Howell if it fails.
- One factor. For a factor in a two-way ANOVA enter its marginal means, the number of observations behind each mean, and the error mean square and error df of the two-way table.
Tukey HSD in Excel, R, Python and SPSS
| Software | How to run it |
|---|---|
| Excel | No built-in Tukey test. The Analysis ToolPak ANOVA gives MSE (MS Within) and df; take q(α; k, df) from a studentized range table or an add-in such as Real Statistics, then HSD = q * SQRT(MSE / n) |
| R | TukeyHSD(aov(y ~ g, data = d)); range distribution: qtukey(0.95, nmeans = 3, df = 27) and ptukey(q, nmeans = 3, df = 27, lower.tail = FALSE) |
| Python | scipy.stats.tukey_hsd(a, b, c); statsmodels.stats.multicomp.pairwise_tukeyhsd(endog=y, groups=g, alpha=0.05); scipy.stats.studentized_range.isf(0.05, 3, 27) |
| SPSS | Analyze > Compare Means and Proportions (Compare Means before version 29) > One-Way ANOVA > Post Hoc > Tukey; for a factorial model use General Linear Model > Univariate > Post Hoc > Tukey |
Frequently Asked Questions
What is Tukey's HSD test?
Tukey's honestly significant difference (HSD) test compares every pair of group means after a one-way ANOVA. It refers each difference to the studentized range distribution, so the chance of at least one false positive across all the pairs stays at α. Every pair gets a difference, an adjusted p-value and a simultaneous confidence interval.
When should I use Tukey HSD instead of t-tests?
Use it when you want all pairwise comparisons among three or more groups. Separate t-tests let the false-positive rate grow: at α = 0.05 three groups (3 pairs) give about a 14.26% chance of at least one false positive and five groups (10 pairs) about 40.13%. Tukey's test holds the family-wise rate at α.
How do I interpret the adjusted p-value and the confidence interval?
The adjusted p-value is the probability, when all population means are equal, of a gap between two of the k means at least as large as the one observed, so compare it with α directly without any further correction. The interval is simultaneous: all the pair intervals together cover the true differences with probability 1 − α. An interval that excludes 0 goes with a p-value below α.
What is the HSD value?
The honestly significant difference is the smallest gap between two means that Tukey's test calls significant: HSD = q × √(MSE / n), where q is the critical value of the studentized range at α for k groups and N − k degrees of freedom. It exists only when all groups have the same size n. With unequal sizes each pair has its own threshold and the calculator shows n/a.
What is the difference between Tukey HSD and Tukey-Kramer?
Tukey-Kramer is Tukey's test for unequal group sizes: the standard error of each pair uses its own sizes, √(MSE / 2 × (1/nᵢ + 1/nⱼ)). Hayter (1984) proved that it is conservative, so its family-wise error rate never exceeds α. With equal sizes the two are the same test, and this calculator switches between them automatically.
Do I need a significant ANOVA before running Tukey HSD?
Not for the error rate: Tukey's procedure controls the family-wise error rate by itself. Running the F test first is a common convention, and the two can disagree, for example a significant F with no significant pair. Decide the order before you look at the data and report both results.
Should I use Tukey HSD or Bonferroni?
For all pairwise comparisons among the groups Tukey is usually more powerful, because it uses the exact joint distribution of the means rather than a worst-case bound. Bonferroni or Holm is the better choice when only a few comparisons were planned or when the p-values come from different kinds of tests; the Bonferroni correction calculator adjusts any list of p-values.
What if the group variances are unequal?
Tukey's test pools the variance of all groups, so it is unreliable when the spreads differ a lot. Check them with the Levene test calculator or the Brown-Forsythe line of the ANOVA calculator, and use Games-Howell for the pairs and Welch ANOVA for the overall test. Games-Howell is not part of this calculator; SPSS, the rstatix package in R and recent SciPy releases offer it.
Why do my results differ from R, SPSS or Python?
The tests agree; the sign and the direction of the intervals differ. This calculator and SciPy subtract the second group from the first, SPSS reports the mean difference as I − J (first minus second), and R's TukeyHSD subtracts the earlier group from the later one, so its differences have the opposite sign and its lower and upper limits are swapped. R's qtukey is documented to be accurate to about four decimals, so R's interval limits can differ from the exact values in the fourth decimal.
Can I run Tukey HSD from means and standard deviations?
Yes. Choose the ANOVA-table input, enter the group means and sizes, and use the pooled variance as MSE: MSE = Σ (nᵢ − 1) sᵢ² / (N − k), with N − k degrees of freedom. For the PlantGrowth groups (SD 0.5831, 0.7937 and 0.4426, n = 10 each) this is (9 × 0.5831² + 9 × 0.7937² + 9 × 0.4426²) / 27 = 0.3886 with 27 df.
How many groups can I compare?
Raw-data mode takes up to 10 groups (45 pairs) and the ANOVA-table input up to 50 means. The number of pairs is k(k − 1)/2, and the more pairs there are, the larger the critical q and the wider the intervals.
Can I use Tukey HSD after a two-way ANOVA?
Yes, for a main effect: enter the marginal means of the factor, the number of observations behind each mean, and the error mean square and error degrees of freedom of the two-way ANOVA table. Comparing individual cells when there is an interaction needs more care; see the two-way ANOVA calculator.
Is Tukey HSD a one-tailed or a two-tailed test?
It is two-sided. The statistic q uses the absolute difference between the two means, so the adjusted p-value tests a difference in either direction and the confidence intervals are two-sided.
Embed This Calculator
Add this free calculator to your course page or LMS.
Adjust the height value to fit your page.