Statistics Concepts
Effect Size Explained: Cohen's d, r, Eta Squared and Cramér's V
An effect size measures how big a difference or relationship is, in a way that does not depend on the sample size. A p-value only says whether an effect is distinguishable from chance. For two group means the standard measure is Cohen's d, the difference in means divided by the pooled standard deviation. With means 78 and 71 and standard deviations 10 and 12 in groups of 30, d = 0.63, a medium-to-large effect.
The common effect sizes
| Measure | Used with | Small, medium, large (Cohen) | What it says |
|---|---|---|---|
| Cohen's d | Two group means | 0.2, 0.5, 0.8 | Difference in means in standard-deviation units |
| Hedges' g | Two group means, small samples | Same as d | d with a small-sample bias correction |
| Pearson r | Two numeric variables | 0.1, 0.3, 0.5 | Strength and direction of a linear relationship |
| Eta squared η² | ANOVA | 0.01, 0.06, 0.14 | Share of the total variance explained by the groups |
| Cramér's V | Chi-square tests | 0.1, 0.3, 0.5 for tables with two rows or columns | Strength of association between two categorical variables |
| Odds ratio | 2 × 2 tables | No fixed benchmark | Ratio of the odds in the two groups; 1 means no effect |
The benchmarks for V shrink as the table grows, because V is Cohen's w divided by the square root of the smaller side of the table minus 1. The Cramér's V calculator applies the benchmarks that fit your table and also gives the bias-corrected V for small samples.
Worked example: Cohen's d by hand
Two teaching methods are compared with 30 students each. Method 1 has a mean score of 78 with a standard deviation of 10, and method 2 has a mean of 71 with a standard deviation of 12.
Pooled SD: s_p = √(((30 − 1)·10² + (30 − 1)·12²) / (30 + 30 − 2)) = 11.0454
Cohen's d = (78 − 71) / 11.0454 = 0.6338
Hedges' g = d × 0.9870 = 0.6255
A d of 0.63 means the group means are about 0.63 standard deviations apart. The same data give t = 2.4545 with 58 degrees of freedom and p = 0.0171, so the effect is both statistically significant and of medium-to-large size. Verify the steps with the effect size calculator. When the two groups have very different spreads, some authors instead use Glass's Δ, dividing by the control group's standard deviation alone; taking method 2 as the control here gives 7 / 12 = 0.58.
What d means in everyday terms
| Cohen's d | The average member of the higher group beats | Chance a random member of the higher group beats a random member of the other | Equivalent r |
|---|---|---|---|
| 0.2 | 57.9% of the lower group | 55.6% | 0.10 |
| 0.5 | 69.1% of the lower group | 63.8% | 0.24 |
| 0.8 | 78.8% of the lower group | 71.4% | 0.37 |
These figures assume two normal distributions with equal spread. They show that even a medium effect leaves heavy overlap, so most individuals cannot be classified by group membership alone.
Why a small p-value is not an effect size
Significance depends on both the size of the effect and the sample size. The two studies below show how far the two can drift apart.
| Study | Cohen's d | Per group | t and p-value | Reading |
|---|---|---|---|---|
| A | 0.2 (small) | 400 | t = 2.83, p = 0.0048 | Significant, yet a small effect |
| B | 0.8 (large) | 5 | t = 1.26, p = 0.2415 | Not significant, yet potentially a large effect |
Study B is not evidence of no effect; it is too small to tell, which is a question of statistical power. Study A shows that with enough data trivial differences become significant, so always read the p-value together with the size. See the p-value explained.
Reporting checklist
- State which measure you used and how it was calculated, for example d with the pooled standard deviation.
- Give a confidence interval for the raw difference, in the original units, next to the standardised value.
- Interpret the size in context; use the benchmarks only as a last resort.
- Use the sample size to justify the study in advance; see sample size explained and the A/B test sample size calculator.
Try the Effect Size Calculator (Cohen's d)
Compute Cohen's d, Hedges' g and the pooled standard deviation from two groups' summary statistics.
Try the Cramér's V Calculator
Get the effect size of a chi-square test for a table of counts, with the benchmark that fits the table.
Frequently Asked Questions
What is a good effect size?
There is no universally good value. Cohen suggested d = 0.2, 0.5 and 0.8 as small, medium and large, r = 0.1, 0.3 and 0.5, and η² = 0.01, 0.06 and 0.14, and he meant them only as a fallback when nothing better is known. Judge an effect against what matters in your field: a d of 0.2 can be worth having for a cheap intervention applied to millions of people.
What is the difference between Cohen's d and Hedges' g?
Both divide the difference in means by a pooled standard deviation. Hedges' g multiplies d by a small correction factor that removes the upward bias of d in small samples. For 30 observations per group the factor is 0.987, turning d = 0.634 into g = 0.626, and for large samples the two are practically identical.
How do I calculate Cohen's d from a t-test?
For two independent groups of sizes n₁ and n₂, d = t × √(1/n₁ + 1/n₂). With t = 2.4545 and 30 per group, d = 2.4545 × √(2/30) = 0.634. For a paired design divide t by √n instead, which gives the standardised mean difference of the changes.
Should I report an effect size with every test?
Yes, alongside the p-value and a confidence interval. The p-value says whether the result could plausibly be chance, the confidence interval shows the range of plausible sizes in the original units, and the standardised effect size lets readers compare studies that use different scales.
What effect size goes with a chi-square test?
Cramér's V, which rescales the chi-square statistic to a 0 to 1 range: √(χ² / (n × min(rows − 1, columns − 1))). For a 2 × 3 table with χ² = 16.67 and n = 150 it is 0.333. For a 2 × 2 table the odds ratio and the phi coefficient are also common choices.
How is effect size used to plan a study?
The sample size you need depends on the smallest effect you want to detect. For a two-sample t-test with 80% power at α = 0.05, d = 0.8 needs 26 people per group, d = 0.5 needs 64 and d = 0.2 needs 394. Guessing the effect size well is the hard part, and it is better taken from earlier studies or from the smallest difference that would matter in practice.