Levene's Test Calculator
Test whether two or more groups have equal variances with Levene's test (mean), the Brown-Forsythe test (median, the default) or the 10% trimmed-mean version. Enter up to 10 groups and get W, both degrees of freedom, the p-value, the F critical value, each group's spread and the ANOVA on absolute deviations that produces W.
Levene's test checks the equal-variance assumption of the one-way ANOVA calculator, the Tukey HSD calculator and the pooled t-test. For exactly two groups of normal data see the F-test for two variances. Whether to test the variances first at all is discussed below.
Numbers separated by commas, spaces, or new lines (at least 2 per group).
Numbers separated by commas, spaces, or new lines (at least 2 per group).
Numbers separated by commas, spaces, or new lines (at least 2 per group).
Related Calculators
What Levene's test checks
Levene's test asks whether two or more groups have the same variance, the assumption of homogeneity of variance (homoscedasticity) behind the pooled two-sample t-test, the classical one-way ANOVA and Tukey's HSD test. The null hypothesis is H₀: σ₁² = σ₂² = … = σₖ²; the alternative is that at least one pair of variances differs.
The idea is simple. Replace every observation by its absolute distance from the centre of its group, then run a one-way ANOVA on those distances. A group with a larger spread has larger distances, so the F ratio of that ANOVA, called W here, is large when the variances differ. Levene (1960) measured the distance from the group mean; Brown and Forsythe (1974) proposed the group median and a trimmed mean, which makes the test much less sensitive to skewed or heavy-tailed data. Bartlett's test is more powerful when the data really are normal but far more sensitive to departures from normality.
How to use the Levene test calculator
- Enter each group's observations separated by commas, spaces or new lines, at least 2 per group (at least 3 are needed for a usable statistic). Add another group gives you up to 10 groups. An empty group is left out and the others keep their own numbers in the results.
- Choose the center: the median (Brown-Forsythe, the default), the mean (Levene's original) or the 10% trimmed mean.
- Set the significance level α, usually 0.05. It fixes the F critical value and the wording of the verdict.
- Press Calculate Levene Test. Read W, the two degrees of freedom, the p-value, the group table and the ANOVA table on the absolute deviations that W comes from.
- Load example enters the PlantGrowth data, Load InsectSprays example six groups with clearly unequal spread and Load NIST gear example the ten batches of the NIST handbook. Copy link to this calculation shares your exact analysis.
Levene test formulas
Zᵢⱼ = |Yᵢⱼ − Ȳᵢ| (mean), |Yᵢⱼ − Ỹᵢ| (median) or |Yᵢⱼ − Ȳᵢ′| (10% trimmed mean)
Z̄ᵢ = mean of the Zᵢⱼ in group i, Z̄ = mean of all N values Zᵢⱼ
W = [ (N − k) / (k − 1) ] · Σᵢ Nᵢ (Z̄ᵢ − Z̄)² / Σᵢ Σⱼ (Zᵢⱼ − Z̄ᵢ)²
df₁ = k − 1, df₂ = N − k
p = P( F(df₁, df₂) ≥ W ); reject H₀ at α when W > F(α; k − 1, N − k)
W is the F statistic of the one-way ANOVA on the Zᵢⱼ: (SS between / df₁) / (SS within / df₂)
The 10% trimmed mean drops the ⌊0.1 · Nᵢ⌋ smallest and the same number of largest values of a group before averaging, so a group with fewer than 10 observations loses nothing and its trimmed mean is the ordinary mean. Only the centre is trimmed: all N distances still enter the ANOVA. Large values of W mean unequal spread, so the p-value is the upper tail of the F distribution, but the test is two-sided in the variances: it detects any difference in spread, not a direction.
Mean, median or trimmed mean: which center to use
| Center | Test name | Works best when | In software |
|---|---|---|---|
| Median | Brown-Forsythe test (1974) | The distribution is skewed or unknown; the safe general-purpose choice | The default here, in R's car::leveneTest and in SciPy's levene |
| Mean | Levene's original test (1960) | The data are symmetric with moderate tails, such as roughly normal samples: best power there | The first row (Based on Mean) of the SPSS homogeneity table |
| 10% trimmed mean | Brown-Forsythe (1974) | The tails are heavy, with outliers in either direction | SPSS and SciPy trim 5% by default, not 10% |
Brown and Forsythe compared the three in Monte Carlo studies: the trimmed mean performed best for a heavy-tailed Cauchy distribution, the median for a skewed χ² distribution with 4 degrees of freedom, and the mean had the best power for symmetric, moderate-tailed distributions. The NIST/SEMATECH handbook recommends the median as the choice that keeps good robustness against many kinds of non-normal data while retaining good power.
The choice can change the verdict. For the handbook's GEAR.DAT data (ten batches of ten gear diameters) the median version gives W = 1.7059 (p = 0.0991), and the handbook concludes that there is insufficient evidence of unequal variances. The mean version gives W = 2.1595 (p = 0.0322) and the 10% trimmed version W = 2.1537 (p = 0.0327), both below 0.05. Decide on the center before you look at the results, and report it.
Worked example by hand: three groups of five
The groups 23, 25, 24, 26, 22 (group 1), 20, 30, 15, 35, 25 (group 2) and 24, 24, 25, 23, 24 (group 3) are what the calculator holds when the page opens. Group 2 is clearly the most spread out: the sample standard deviations are 1.5811, 7.9057 and 0.7071.
Working
The centers are 24, 25 and 24. In each group the mean and the median coincide, and with 5 values nothing is trimmed, so all three center options give the same result here.
Absolute deviations from the center: 1, 1, 0, 2, 2 in group 1 (mean 1.2), 5, 5, 10, 10, 0 in group 2 (mean 6) and 0, 0, 1, 1, 0 in group 3 (mean 0.4). The mean of all 15 distances is 38 / 15 = 2.5333.
SS between = 5 × [(1.2 − 2.5333)² + (6 − 2.5333)² + (0.4 − 2.5333)²] = 91.7333 and SS within = 2.8 + 70 + 1.2 = 74, with df₁ = 3 − 1 = 2 and df₂ = 15 − 3 = 12.
W = (91.7333 / 2) / (74 / 12) = 45.8667 / 6.1667 = 7.4378. The upper 5% point of F(2, 12) is 3.8853 and the p-value is 0.007924.
W is above 3.8853 and p is below 0.05, so the variances differ significantly, as the standard deviations suggested.
Worked example with real data: PlantGrowth and InsectSprays
The PlantGrowth data set that ships with R holds the dry weight of 30 plants, ten in each of three groups (a control and two treatments). The sample standard deviations are 0.5831, 0.7937 and 0.4426, a ratio of 1.7933. Press Load example: with the median as center W = 1.1192 on 2 and 27 degrees of freedom, below the F critical value of 3.3541, with p = 0.341227. R's car::leveneTest(weight ~ group, data = PlantGrowth) prints the same F value, 1.1192, and Pr(>F) = 0.3412. No version of the test finds evidence of unequal variances, so the pooled analysis in the ANOVA calculator is not contradicted.
The InsectSprays data set holds the insect counts in 72 plots treated with six sprays (A to F), twelve plots per spray. The standard deviations run from 1.7321 (spray E) to 6.2134 (spray F), a ratio of 3.5873, far above the rule of thumb of 2. Press Load InsectSprays example: all three centers reject at 0.05 and only the size of W depends on the center. R's car::leveneTest(count ~ spray, data = InsectSprays) prints F = 3.8214 with Pr(>F) = 0.004223, and F = 6.4554 with Pr(>F) = 6.1e-05 for center = mean. Counts often have a variance that grows with the mean, and taking square roots, the transformation R's documentation applies to these data, evens the spread out: the median-centered W falls to 0.8836 (p = 0.497135).
| Data set | Center | W | df1, df2 | p |
|---|---|---|---|---|
| PlantGrowth | Median | 1.1192 | 2, 27 | 0.341227 |
| PlantGrowth | Mean | 1.237 | 2, 27 | 0.306195 |
| PlantGrowth | 10% trimmed mean | 1.2777 | 2, 27 | 0.294985 |
| InsectSprays | Median | 3.8214 | 5, 66 | 0.004223 |
| InsectSprays | Mean | 6.4554 | 5, 66 | 0.000061 |
| InsectSprays | 10% trimmed mean | 5.8928 | 5, 66 | 0.000146 |
| InsectSprays, square roots | Median | 0.8836 | 5, 66 | 0.497135 |
How to read the result
- W and the p-value. W is the F ratio of the ANOVA on the absolute deviations. A p-value below α, or a W above the F critical value, means the variances differ significantly. It does not say which group is more variable: compare the standard deviations in the group table.
- Not significant is not equal. With small groups the test has little power, so unequal variances can go undetected; with very large groups it flags differences too small to matter.
- Max SD / Min SD. A rough yardstick that does not depend on the sample size: textbooks such as Moore and McCabe accept the classical procedures when the largest sample standard deviation is less than twice the smallest.
- Group table. It lists each group's center, standard deviation and mean absolute deviation from the center. W grows with the differences among those mean deviations relative to the variation of the distances within the groups.
- Undefined statistic. When every group holds only two values, or the distances inside every group are all equal, there is no variation within the groups to compare against and W is undefined. The calculator says so instead of printing a number.
Should you test the variances before a t-test or ANOVA?
Many courses teach a two-step routine: run Levene's test, then choose the pooled or the Welch version of the t-test or ANOVA by its p-value. Statisticians increasingly argue against it. Zimmerman (2004) reviewed the evidence, advised against preliminary tests of equal variances and, for unequal group sizes, recommended the separate-variances test unconditionally. Delacre, Lakens and Leys (2017) recommend Welch's t-test by default for two groups, and Delacre and colleagues (2019) Welch's F-test in one-way ANOVA: it costs almost nothing when the variances happen to be equal and protects the error rate when they are not. Both Welch versions are options in the ANOVA calculator and the t-test calculator.
Levene's test remains the right tool when the variances are the question, for example the consistency of two machines, methods or raters, when you need to check an assumption that has no Welch version, such as the pooled error of Tukey's HSD test and of the two-way ANOVA, and as a diagnostic next to a plot of the residuals.
Levene's test compared with other tests of equal variances
| Test | Groups | Sensitivity to non-normality | Notes |
|---|---|---|---|
| Levene (mean) | 2 or more | Moderate | The 1960 original; best power for symmetric data |
| Brown-Forsythe (median) | 2 or more | Low | Levene's ANOVA on distances from the median |
| Bartlett | 2 or more | High | Most powerful when the data really are normal; R bartlett.test, SciPy bartlett |
| F-test for two variances | Exactly 2 | High | Ratio of the two sample variances; see the F-test calculator |
| Fligner-Killeen | 2 or more | Low | Rank-based and robust; R fligner.test, SciPy fligner |
For exactly two normal samples, the F-test for two variances is the classical choice. Which assumptions the other procedures make is summarised in parametric vs nonparametric tests, and how to check normality in normality tests explained.
Levene's test in Excel, R, Python and SPSS
| Software | How to run it |
|---|---|
| Excel | No built-in test, but it is an ANOVA on absolute deviations: in a helper column type =ABS(B2 - MEDIAN(B$2:B$11)) for each group (AVERAGE instead of MEDIAN for the mean version), then Data > Data Analysis > Anova: Single Factor on the helper columns. Its F and P-value are W and p |
| R | library(car); leveneTest(y ~ g, data = d) uses the median (Brown-Forsythe); leveneTest(y ~ g, data = d, center = mean) is Levene's original; leveneTest(y ~ g, data = d, center = mean, trim = 0.1) is the 10% trimmed version. Base R: bartlett.test(y ~ g, data = d), fligner.test(y ~ g, data = d) |
| Python | scipy.stats.levene(a, b, c) uses the median; center='mean' is Levene's original. For the 10% trimmed mean use statsmodels.stats.oneway.test_scale_oneway([a, b, c], method='equal', center='trimmed', trim_frac_mean=0.1). SciPy's own center='trimmed' cuts 5% by default and was incorrect before the fix for scipy issue 23551 |
| SPSS | Analyze > Compare Means and Proportions (Compare Means before version 29) > One-Way ANOVA > Options > Homogeneity of variance test. Recent versions print four rows: Based on Mean, Based on Median, Based on Median and with adjusted df, and Based on trimmed mean (5% trim). The Independent-Samples T Test prints Levene's F and Sig. above its two t-test rows |
Frequently Asked Questions
What is Levene's test?
Levene's test checks whether two or more groups have equal variances (homogeneity of variance). It runs a one-way ANOVA on the absolute deviations of the observations from the center of their group; a significant result means that at least two group variances differ.
What are the null and alternative hypotheses of Levene's test?
H₀: all group variances are equal (σ₁² = σ₂² = … = σₖ²). H₁: at least two of the variances differ. The test does not say which groups differ or which one has the larger variance.
How do I interpret Levene's test?
Compare the p-value with α, usually 0.05. If p is below α, reject H₀: the variances differ significantly, and procedures that pool them (Student's t-test, classical ANOVA, Tukey HSD) are questionable. If p is not below α there is no evidence of unequal variances, which is not proof that they are equal.
What is the difference between Levene's test and the Brown-Forsythe test?
Both are the same ANOVA on absolute deviations; they differ in the center. Levene (1960) used the group mean, Brown and Forsythe (1974) the group median and a trimmed mean. The median version is more robust to skewed data, and R's car::leveneTest and SciPy's levene use it by default. Many textbooks and programs call the mean version Levene's test and the median version the Brown-Forsythe test.
Which center should I choose: mean, median or trimmed mean?
The median is the safe default; the NIST handbook recommends it for its robustness against many kinds of non-normal data with good power. The mean has the best power for symmetric, moderate-tailed data and the 10% trimmed mean does best for heavy-tailed data. The choice can change the verdict, as it does for the handbook's gear data (median p = 0.0991, mean p = 0.0322), so decide before you look at the results and report the center.
What should I do if Levene's test is significant?
Use methods that do not assume equal variances: Welch's t-test for two groups, or Welch's ANOVA followed by Games-Howell comparisons for three or more. If the spread grows with the mean, as it does for counts, a log or square-root transformation can even it out; in the InsectSprays example the median-centered W drops from 3.8214 (p = 0.004223) to 0.8836 (p = 0.497135) after taking square roots.
Should I run Levene's test before a t-test or ANOVA?
Not as an automatic gate. Zimmerman (2004) advised against preliminary tests of equal variances, and Delacre, Lakens and Leys (2017, 2019) recommend Welch's t-test and Welch's F-test by default instead of choosing between the pooled and the Welch version by a Levene p-value. Use Levene's test when the variances themselves matter, or to check an assumption that has no Welch version.
Levene's test or Bartlett's test?
Bartlett's test is more powerful when the data really are normal but very sensitive to departures from normality, so it can reject for reasons other than unequal variances. Levene's test with the median is the safer general-purpose choice; the NIST handbook presents it as the alternative to Bartlett's test.
How many observations per group do I need?
At least 2 per group for the calculation and at least 3 for a usable statistic. If every group holds two values, the two distances in each group are equal, there is no variation within the groups and W is undefined; the calculator tells you so. The test also has little power with small groups, so a non-significant result from 5 or 10 observations per group says little.
Does Levene's test assume normality?
It assumes independent observations and reasonably regular distributions for the F distribution of W to be accurate, but it is much less sensitive to non-normality than Bartlett's test, especially with the median or the trimmed mean as center. With strongly skewed data and small groups the p-value is approximate.
Why do my results differ from SPSS, R or Python?
Check the center first. SPSS prints four rows (Based on Mean, Based on Median, Based on Median and with adjusted df, and Based on trimmed mean with a 5% trim); R's leveneTest and SciPy's levene default to the median; SciPy's trimmed option cuts 5% by default and was incorrect before the fix for scipy issue 23551. With the same center W and p match this calculator, which agrees with SciPy (mean and median), R's car::leveneTest and statsmodels (all three centers, including the 10% trimmed mean).
Can I use Levene's test with two groups?
Yes. With two groups W is the square of the two-sample t statistic computed on the absolute deviations, and the test compares the two variances. For two normal samples the F-test for two variances is the classical alternative, but it is far more sensitive to non-normality.
Is Levene's test one-tailed or two-tailed?
The p-value is the upper tail of the F distribution, because a large W signals a difference in spread, but the test is two-sided in the variances: it detects any difference in spread, not a direction. To see which group is more variable, look at the standard deviations in the group table.
How do I do Levene's test in Excel?
Excel has no built-in Levene test, but the test is an ANOVA on absolute deviations. For each value compute =ABS(value - group median) (or the group mean), then run Data > Data Analysis > Anova: Single Factor on those columns. The F and P-value in the output are Levene's W and p. The calculator above does all of this in one step.
How many groups can I compare?
Up to 10 groups in this calculator, each with any number of observations. The test itself works for any number of groups from two upward.
Embed This Calculator
Add this free calculator to your course page or LMS.
Adjust the height value to fit your page.