Statistics How-To
Which Statistical Test Should I Use?
Choose the test from three facts about your data: what kind of outcome you measured, how many groups or variables you compare, and whether the observations are independent or paired. Compare two group means with a t-test, three or more with ANOVA, two categorical variables with a chi-square test and two numeric variables with correlation or regression. If the data are ranks, or clearly not normal in a small sample, switch to the rank-based test in the same row.
Three questions that pick the test
- What type is the outcome? A numeric outcome is a measurement such as a height, a time or a score. An ordinal outcome is a ranking or a rating on a short scale, such as 1 to 5 stars. A categorical outcome is a label: yes or no, a colour, a treatment group.
- How many groups or variables are compared? One sample against a known value, two groups, three or more groups, or two variables measured on the same subjects.
- Are the observations independent or paired? Different subjects in each group are independent. The same subjects measured twice, or subjects matched in pairs, are paired, and the test must use the differences within each pair.
Once you have the three answers, the table below names the usual test and the rank-based test to use when the usual test's assumptions fail. Rank-based tests are also called nonparametric; the parametric vs nonparametric guide explains what you give up by using them.
Decision table
| You want to compare | Usual test | If the assumptions fail |
|---|---|---|
| One mean with a known value | One-sample t-test | Wilcoxon signed-rank test |
| One mean, population standard deviation known | Z-test | Wilcoxon signed-rank test |
| Two independent groups (different subjects) | Two-sample t-test (Welch) | Mann-Whitney U test |
| Two paired measurements (same subjects) | Paired t-test | Wilcoxon signed-rank test on the differences |
| Three or more independent groups | One-way ANOVA | Kruskal-Wallis test |
| Which pairs differ after a significant ANOVA | Tukey HSD | Dunn's test after Kruskal-Wallis |
| Two factors at once | Two-way ANOVA | A regression model with both factors |
| Two numeric variables, how strongly related | Pearson correlation | Spearman correlation |
| Predict a numeric outcome from another variable | Linear regression | Transform the variables, or fit a curve |
| Two categorical variables | Chi-square test of independence | Fisher's exact test |
| One categorical variable against expected shares | Chi-square goodness of fit | An exact multinomial test |
| One proportion with a target value | One-proportion z-test | An exact binomial test |
| Two proportions | Two-proportion z-test | Fisher's exact test |
| Agreement between two raters | Cohen's kappa | Weighted kappa for ordered categories |
Check the assumptions before you trust the p-value
- Independence. One observation must not influence another. Repeated measures on the same person, students within one classroom or the same customer counted twice break it, and no later step repairs it. Use a paired or repeated-measures design instead.
- Roughly normal data, or a large sample. The t-test and ANOVA assume the data (more exactly, the residuals) are close to normal, which matters most below about 30 per group. Look at a histogram or box plot, and use the Shapiro-Wilk test as a supplement, not a replacement, for looking.
- Similar spread in the groups. Welch's version of the two-sample t-test does not assume equal variances, so it is the safe default. For ANOVA, check that the largest standard deviation is not more than about twice the smallest.
- Large enough expected counts. The chi-square test needs an expected count of about 5 or more in every cell of the table. With fewer, use Fisher's exact test.
- No influential outliers. A single extreme value can create or hide a difference in means. Check it against the source, and if it is genuine, run a rank-based test as well and report both.
Five worked decisions
- Does a diet change weight? Each of 40 people is weighed before and after. The outcome is numeric, there are two measurements, and they come from the same people, so the data are paired. Use the paired t-test on the 40 differences.
- Do two landing pages convert differently? 48 of 400 visitors convert on A and 66 of 400 on B. The outcome is categorical (converted or not) and the groups are independent, so use the two-proportion z-test, or the chi-square test on the 2 × 2 table.
- Do three teaching methods give different exam scores? A numeric outcome and three independent groups call for a one-way ANOVA. If it is significant, run Tukey HSD to see which methods differ.
- Do two groups rate a product differently on a 1 to 5 scale? Ratings are ordinal, and the gaps between the points are not guaranteed to be equal. Use the Mann-Whitney U test.
- Do students who study longer score higher? Two numeric variables measured on the same students: report the correlation and fit a regression line to describe the slope.
Mistakes that change the answer
- Running many t-tests instead of one ANOVA. At the 5% level, 3 comparisons carry a 14.26% chance of a false positive and 10 comparisons 40.13%. See multiple comparisons explained.
- Treating paired data as two independent groups. The independent test ignores the pairing and throws away the power that comes from comparing each subject with themselves.
- Choosing the tail after seeing the data. Decide between a one-tailed and a two-tailed test first; see one-tailed vs two-tailed tests.
- Reading a non-significant result as no effect. A small sample may simply lack the power to see it. Compute the power and the effect size as well as the p-value.
- Using percentages or ratings as if they were unbounded measurements. Proportions near 0 or 1 and ordinal ratings often break the normality assumption, which is why the table gives them their own rows.
For the ideas behind every test in the table, read hypothesis testing, the p-value explained and type 1 and type 2 errors.
Try the T-Test Calculator
Compare two group means with a t-test, with the t statistic, degrees of freedom and p-value shown.
Try the ANOVA Calculator
Compare the means of three or more groups with a one-way ANOVA table and F test.
Frequently Asked Questions
What is the difference between a t-test and ANOVA?
A t-test compares the means of two groups, or one group with a fixed value. ANOVA compares the means of three or more groups in a single test. With exactly two groups, ANOVA gives the same p-value as the pooled-variance t-test, and the F statistic is the square of the t statistic. ANOVA only says that at least one mean differs, so a follow-up such as Tukey's HSD is needed to say which.
When should I use a z-test instead of a t-test?
Use a z-test for a mean only when the population standard deviation is known, which is rare, or for proportions. When the standard deviation is estimated from the sample, use the t-test: it allows for that extra uncertainty with a slightly wider distribution. For samples of a few hundred the two give almost identical answers, because the t distribution approaches the normal.
Why not run several t-tests instead of one ANOVA?
Each test at the 5% level has a 5% chance of a false positive, and the chances add up. Three pairwise comparisons carry a 14.26% chance of at least one false positive when nothing is going on, and ten comparisons 40.13%. One ANOVA keeps the overall error at 5%, and a post hoc test such as Tukey's HSD then compares the pairs while controlling the family-wise error rate.
Chi-square test or Fisher's exact test?
Both test whether two categorical variables are associated. The chi-square test relies on a large-sample approximation and is unreliable when any expected count in the table is below about 5. Fisher's exact test computes the probability exactly, so it is the safe choice for small counts, and it works for large tables too, at a higher computing cost.
What if my data are not normally distributed?
With about 30 or more observations per group the t-test and ANOVA are fairly robust, because the sample means are close to normal even when the data are not. With small samples, strong skew or outliers, use the rank-based alternative in the table (Mann-Whitney U, Wilcoxon signed-rank, Kruskal-Wallis, Spearman), or transform the data, for example with a logarithm.
Does the test I choose depend on the sample size?
Only through its assumptions. Small samples need the data to be roughly normal for the t-test to be valid, while large samples relax that. What sample size does change is the power: a test can be the right one and still miss a real effect if the sample is too small, so plan the sample size before collecting data.