Statistics Concepts
Degrees of Freedom Explained: What df Means in Statistics
Degrees of freedom (df) count the values in a calculation that are free to vary once its constraints are fixed. If you know the mean of five numbers and four of them, the fifth is forced, so the deviations from the mean have 5 − 1 = 4 degrees of freedom. The t, chi-square and F distributions all depend on df, so the df decides the critical value and the p-value of a test.
The idea in one example
Take the five values 4, 8, 6, 5 and 12. Their mean is 7, so the deviations from the mean are −3, 1, −1, −2 and 5. Deviations from the sample mean always add up to zero, so once four of them are known the fifth is fixed: if you are told −3, 1, −1 and −2, the last must be +5. Only four of the five deviations carry independent information. The data have n − 1 = 4 degrees of freedom for variability.
That is why the sample variance divides the sum of squares, 40, by 4 and not by 5, giving s² = 10 (see why sample variance divides by n − 1). In general, degrees of freedom = the number of observations − the number of quantities you had to estimate from them.
Degrees of freedom for common tests
| Test | Degrees of freedom | Example |
|---|---|---|
| One-sample t-test | n − 1 | 20 observations: 19 |
| Paired t-test | n − 1, with n the number of pairs | 12 pairs: 11 |
| Two-sample t-test, pooled variance | n₁ + n₂ − 2 | 10 and 15 observations: 23 |
| Two-sample t-test, Welch | Welch-Satterthwaite formula, usually a fraction | n₁ = 10, s₁ = 2, n₂ = 15, s₂ = 5: 19.76 |
| One-way ANOVA | Between k − 1, within N − k | 3 groups of 10: 2 and 27 |
| Chi-square test of independence | (rows − 1)(columns − 1) | 2 × 3 table: 2 |
| Chi-square goodness of fit | k − 1, minus each parameter estimated | 6 die faces: 5 |
| Correlation test and simple regression | n − 2 | 5 points: 3 |
| Multiple regression with p predictors | n − p − 1 | 30 observations, 3 predictors: 26 |
Each row follows the same rule: count the observations, then subtract one for every parameter the test estimates from the data. The degrees of freedom calculator applies these rules, and the calculators for the t-test, ANOVA, chi-square and regression show the df they use.
The Welch degrees of freedom
When two groups have different variances, Welch's t-test does not pool them and estimates the degrees of freedom from the data:
df = (s₁²/n₁ + s₂²/n₂)² / [ (s₁²/n₁)²/(n₁ − 1) + (s₂²/n₂)²/(n₂ − 1) ]
Example: n₁ = 10, s₁ = 2, n₂ = 15, s₂ = 5 gives df = 19.76
The result always lies between the smaller of n₁ − 1 and n₂ − 1 (here 9) and n₁ + n₂ − 2 (here 23). It approaches 23 when the two variances are equal and the group sizes match, and falls toward the smaller group when the smaller group is also the noisier one. A smaller df widens the tails of the t distribution, which is how the test protects you against unequal spreads.
How df changes the critical value
The t distribution has heavier tails than the normal distribution when df is small, because the sample standard deviation is a shaky estimate. A test therefore needs a more extreme t to reach 5% significance. The two-tailed 5% critical values are:
| Degrees of freedom | Critical t (two-tailed, α = 0.05) |
|---|---|
| 1 | 12.706 |
| 2 | 4.303 |
| 5 | 2.571 |
| 10 | 2.228 |
| 20 | 2.086 |
| 30 | 2.042 |
| 60 | 2.000 |
| 120 | 1.980 |
| Infinite (the normal z) | 1.960 |
A t statistic of 2.1 is significant with 30 df (critical value 2.042) but not with 10 df (critical value 2.228), so the same t needs more data behind it to count. Look up any value in the t-table, or read how to read a t-table.
Finding critical values by software
| Tool | Command for the 5% two-tailed t critical value with 20 df |
|---|---|
| Excel / Google Sheets | =T.INV.2T(0.05,20) gives 2.086 |
| R | qt(0.975, df = 20) gives 2.085963 |
| Python (SciPy) | from scipy.stats import t; t.ppf(0.975, 20) gives 2.085963 |
| TI-84 | invT(0.975, 20) gives 2.085963 |
Common misunderstandings
- df is not the sample size. A t-test on 30 observations has 29 df. The tables and software want the df, not n.
- df is not always n − 1. That is the rule for one mean. Two pooled groups lose two (n₁ + n₂ − 2), a regression line loses two, and a chi-square test depends on the shape of the table.
- More df means more information, not a bigger effect. It makes the test more sensitive by thinning the tails of the reference distribution, and the estimate of the standard deviation less noisy.
- Do not round a Welch df up. Use the fractional value in software; if you must use a table, round down, which is conservative.
Try the Degrees of Freedom Calculator
Find the degrees of freedom for a t-test, ANOVA, chi-square test or regression from your sample sizes.
Frequently Asked Questions
Why is the sample variance divided by n − 1?
Because the deviations from the sample mean always add up to zero, so only n − 1 of them are free. Dividing the sum of squares by its degrees of freedom, n − 1, makes the sample variance an unbiased estimate of the population variance. Dividing by n underestimates it, by a factor of (n − 1)/n on average.
Can degrees of freedom be a decimal?
Yes. Welch's t-test uses the Welch-Satterthwaite formula, which usually gives a fraction such as 19.76. The t distribution is defined for any positive df, so software uses the fractional value directly, and you should not round it when you look up a p-value. Tables need a whole number, so round down to be safe.
What happens to the t distribution as the degrees of freedom increase?
Its tails get thinner and it approaches the standard normal distribution. The two-tailed 5% critical value falls from 12.706 at 1 df to 2.086 at 20 df, 1.980 at 120 df and 1.960 in the limit. That is why a z-test and a t-test agree closely for large samples.
Are degrees of freedom the same as the sample size?
No. Degrees of freedom are the sample size minus the number of quantities estimated from the data (n − 1 for one mean, n − 2 for a regression line with two estimated coefficients). They are always smaller than n whenever anything is estimated, and a t-test with 30 observations has 29 df, not 30.
How many degrees of freedom does a chi-square test have?
For a test of independence on a table with r rows and c columns it is (r − 1)(c − 1), because once the row and column totals are fixed only that many cells can vary freely. A 2 × 2 table has 1 df and a 3 × 4 table 6 df. For a goodness-of-fit test on k categories it is k − 1, minus one for each parameter estimated from the data.
What are the degrees of freedom in a one-way ANOVA?
There are two numbers. The numerator (between-groups) df is k − 1 for k groups, and the denominator (within-groups) df is N − k for N observations in total. Three groups of 10 give 2 and 27 df, and the F statistic is compared with the F distribution for that pair. The total df, N − 1 = 29, is the sum of the two.