Statistics Concepts
Parametric vs Nonparametric Tests: What Is the Difference?
Parametric tests such as the t-test, ANOVA and Pearson correlation assume the data come from a distribution described by parameters, usually the normal, and compare means. Nonparametric tests such as the Mann-Whitney U, Wilcoxon signed-rank, Kruskal-Wallis and Spearman tests replace the values with their ranks and assume much less. Use the parametric test when its assumptions are reasonable, and the rank-based one for ordinal data, outliers or small skewed samples.
Each parametric test has a rank-based counterpart
| Question | Parametric test | Nonparametric alternative |
|---|---|---|
| One sample against a value | One-sample t-test | Wilcoxon signed-rank test |
| Two paired measurements | Paired t-test | Wilcoxon signed-rank test on the differences |
| Two independent groups | Two-sample t-test | Mann-Whitney U test |
| Three or more independent groups | One-way ANOVA | Kruskal-Wallis test |
| Association between two variables | Pearson correlation | Spearman correlation |
What each family assumes
- Parametric tests assume interval or ratio data, independent observations, roughly normal data or residuals (or a large sample), and for several tests similar spread across groups. In return they use the actual values, which gives them the most power when the assumptions hold.
- Nonparametric tests assume independent observations and data that can at least be ordered. They use ranks, so they work for ordinal ratings, are not thrown off by a single extreme value and give exact p-values in small samples without ties.
The details of the normality assumption, and how to check it, are in normality tests explained.
A worked comparison
Two groups of 7 scores: group A is 12, 15, 14, 10, 13, 16, 11 and group B is 9, 11, 10, 8, 12, 10, 9. Pool the 14 values and rank them from 1 (smallest) to 14, giving tied values the average rank. Group A's ranks add up to 72 and group B's to 33, out of a total of 105.
| Test | Statistic | p-value |
|---|---|---|
| Two-sample t-test (pooled) | t = 3.267, 12 df | 0.0067 |
| Mann-Whitney U test | U = 44 out of a maximum of 49 | 0.0144 |
Both reject equality at the 5% level, which is typical when the data are well behaved; the rank test returns a slightly larger p-value. The probability that a random score from A beats one from B, counting ties as half, is 44/49 = 0.898. To reproduce this, use the Mann-Whitney calculator and the two-sample t-test calculator.
When to choose which
| Situation | Prefer |
|---|---|
| Measurements, roughly normal, 30 or more per group | Parametric |
| Ordinal ratings such as 1 to 5 stars | Nonparametric |
| Small sample and clearly skewed data | Nonparametric |
| Genuine outliers that cannot be removed | Nonparametric, or a parametric test with a robust check |
| Small counts in a table | Fisher's exact test rather than chi-square |
| You need a confidence interval for a mean difference | Parametric |
Nonparametric tests give up little when the data are normal: their efficiency against the t-test is about 95.5% (3/π). Read more on the choice in which statistical test to use, and on the parametric side in hypothesis testing.
Try the Mann-Whitney U Test Calculator
Compare two independent groups by ranks, with U, the p-value and an effect size.
Try the Kruskal-Wallis Test Calculator
Compare three or more independent groups by ranks, with post hoc pairs.
Frequently Asked Questions
Are nonparametric tests assumption-free?
No, they are distribution-free, not assumption-free. They do not require normal data, but the observations still have to be independent, and comparisons of two groups are cleanest when the two distributions have a similar shape. Without similar shapes a Mann-Whitney test compares the tendency to produce larger values rather than the medians.
Do nonparametric tests test the same hypothesis as the t-test?
Not exactly. A t-test compares means. The Mann-Whitney U test asks whether values from one group tend to be larger than values from the other, which becomes a comparison of medians only when the shapes match, and the Kruskal-Wallis test asks the same of several groups. Report the hypothesis that the test actually addresses.
Do I lose power by using a nonparametric test?
Very little when the data really are normal: the asymptotic relative efficiency of the Wilcoxon-Mann-Whitney test against the t-test is 3/π, about 0.955, so it needs roughly 5% more observations for the same power. For skewed or heavy-tailed data the rank test can be the more powerful of the two.
Can I use the t-test with non-normal data?
Often yes. With 30 or more observations per group, the sampling distribution of the mean is close to normal by the central limit theorem, so the t-test is fairly robust. It is less safe with small samples, strong skew or influential outliers, which is exactly when the rank-based test earns its place.
What is the nonparametric alternative to the paired t-test?
The Wilcoxon signed-rank test, applied to the within-pair differences, or the simpler sign test that uses only the direction of each difference. The signed-rank test also uses the size of the differences through their ranks, so it is usually the more powerful of the two.
Should I run both tests and report the better one?
No. Choose the primary test before looking at the results, based on the design and the data type, and report it. Running both and picking the smaller p-value is a form of p-hacking. It is fine to report the second test as a sensitivity analysis, provided you state that the primary test was decided in advance.