Statistics Concepts
Normality Tests Explained: Shapiro-Wilk, Q-Q Plots and Skewness
To check normality, combine three views: a Q-Q plot to see the shape, skewness and kurtosis to measure it, and a formal test such as Shapiro-Wilk to judge it. The null hypothesis of Shapiro-Wilk is that the data come from a normal distribution, so a small p-value is evidence against normality, and a large one only means that no departure was detected.
Three ways to check normality
- Q-Q plot. Plot the sorted data against the values a normal distribution would give at the same positions. Points that follow a straight line indicate normality; a curve indicates skewness, and points that peel away at both ends indicate heavy or light tails. The Shapiro-Wilk calculator draws one with Blom plotting positions.
- Skewness and kurtosis. A normal distribution has skewness 0 and excess kurtosis 0. Dividing the sample skewness by its standard error gives an approximate z-score.
- Shapiro-Wilk test. H0: the data are normal. Reject H0 when the p-value is below α. See null and alternative hypothesis for how to read such a decision.
SE(skewness) = √( 6n(n − 1) / ((n − 2)(n + 1)(n + 3)) )
n = 10 gives SE = 0.687
Worked examples
Three small datasets, all checked with the sample skewness and excess kurtosis (the same versions Excel returns with SKEW and KURT) and the Shapiro-Wilk test:
| Dataset | n | W | p-value | Skewness | Excess kurtosis | Skewness ÷ SE |
|---|---|---|---|---|---|---|
| A: 72, 75, 78, 80, 81, 83, 85, 88, 90, 94 | 10 | 0.9893 | 0.9959 | 0.1345 | −0.6150 | 0.20 |
| B: 1, 1, 2, 2, 3, 4, 6, 9, 15, 30 | 10 | 0.7280 | 0.0019 | 2.0772 | 4.4169 | 3.02 |
| C: 4, 8, 6, 5, 12 | 5 | 0.9124 | 0.4822 | 1.1859 | 1.0500 | 1.30 |
Dataset A shows no sign of departure. Dataset B is strongly right-skewed, and Shapiro-Wilk rejects normality at 0.05 and even at 0.01. Dataset C has skewness above 1, yet the test cannot reject with only five values. The result is consistent with normality, but it says almost nothing because the sample is so small.
The large-sample and small-sample traps
- Small n: the test rarely rejects, so a non-significant p-value is weak reassurance. Rely on the plot and on knowledge of how the measurement behaves.
- Large n: the test rejects almost everything. A simulated sample of 4,000 values from a t distribution with 10 degrees of freedom, which is nearly indistinguishable from a normal curve, gives W = 0.9963 and p = 2.3e-8 (generated with scipy.stats.t.rvs(df=10, size=4000, random_state=1)). The departure is real but harmless for most methods; see the central limit theorem.
What to test
The normality assumption of a t-test concerns each group (or the paired differences), and for regression and ANOVA it concerns the residuals. Pooling groups with different means makes the combined data look non-normal even when every group is normal. Read about the exact assumptions in which statistical test to use.
If the data are not normal
- Transform: a logarithm often fixes right skew in positive data such as incomes and waiting times; then re-check the plot.
- Switch to a rank-based test: Mann-Whitney U, Wilcoxon signed-rank or Kruskal-Wallis; the trade-offs are in parametric vs nonparametric tests.
- Check for outliers or mixtures: a few extreme values or two hidden subgroups are the most common reasons for a failed test, and fixing the cause is better than choosing a different test.
In software
| Tool | Command |
|---|---|
| R | shapiro.test(x) (3 to 5,000 values); qqnorm(x); qqline(x) |
| Python (SciPy) | scipy.stats.shapiro(x) ; scipy.stats.probplot(x, plot=plt) |
| Excel | SKEW(range) and KURT(range) for the shape; no built-in Shapiro-Wilk test |
Try the Shapiro-Wilk Test Calculator
W statistic, p-value, skewness, kurtosis and a normal Q-Q plot for your data.
Try the Skewness and Kurtosis Calculator
Sample and population skewness and excess kurtosis with plain-language labels.
Frequently Asked Questions
Which normality test is the best?
Shapiro-Wilk is the most widely recommended for small and moderate samples because it has good power against many kinds of departure from normality. Anderson-Darling and Jarque-Bera are common alternatives. Whichever test you use, look at a Q-Q plot as well, because a p-value alone does not show how the data depart from normal.
Does a p-value above 0.05 prove that my data are normal?
No. It means the test found no significant evidence against normality. With a very small sample the test has almost no power: the five values 4, 8, 6, 5 and 12 give p = 0.4822 although their skewness is 1.19. Absence of evidence is not evidence of absence.
Do I need normal data for a t-test?
The t-test assumes the sampling distribution of the mean (or of the mean difference) is normal. With moderate or large samples, mild skewness has little effect because of the central limit theorem. With small samples and clear skewness or outliers, use a rank-based test such as Mann-Whitney or Wilcoxon signed-rank instead.
Should I test the raw data or the residuals?
For regression and ANOVA the assumption concerns the residuals (errors), so test those. For a t-test, check each group, or the paired differences for a paired design. Testing all groups pooled together can show non-normality that is only caused by the group difference you are trying to detect.
Why does Shapiro-Wilk reject almost everything with big samples?
Its power grows with n, so it detects departures that are real but too small to matter. A simulated sample of 4,000 values from a t distribution with 10 degrees of freedom, which looks almost normal, gives W = 0.9963 and p = 2.3e-8 (SciPy, random_state = 1). With large samples judge the size of the departure on a Q-Q plot, not only the p-value.
What sample sizes does Shapiro-Wilk allow?
From 3 up to 5,000 observations in R's shapiro.test. SciPy documents that for more than 5,000 values the W statistic is still accurate but the p-value may not be. With small samples the test has little power to detect anything but severe departures.