Two Sample T-Test Calculator
Compare two independent sample means with Welch’s t-test (default) and the pooled Student test shown together, from pasted data or n, x̄, and Sx for each group.
Hub for quick tests: general t-test calculator. Check equal variances with the F-test or Levene test.
Comma- or space-separated values (≥ 2).
Comma- or space-separated values (≥ 2).
Related Calculators
T-Test Calculator
Compare sample means with common t-test workflows and interpretable outputs.
F-Test for Two Variances
Compare two variances with F, p-value, and df from data or Sx and n — the TI-84 2-SampFTest.
Effect Size Calculator (Cohen's d)
Quantify the difference between two groups with Cohen's d, Hedges' g, and pooled SD.
Learn More
T-Test in Excel: T.TEST and the Data Analysis ToolPak
Run a t-test in Excel with T.TEST or the Analysis ToolPak: the type and tails arguments, paired and unequal-variance tests, a checked example and how to read the output.
T-Test in R and Python: t.test() and scipy.stats
Run one-sample, two-sample, Welch and paired t-tests in R and Python side by side, see why the default variance option differs and how one outlier can flip the result — checked against SciPy.
Hypotheses
H₀: μ₁ − μ₂ = δ₀ (often δ₀ = 0). The alternative follows your selection: two-sided, μ₁ − μ₂ > δ₀, or μ₁ − μ₂ < δ₀. The confidence interval is for μ₁ − μ₂ using the same method and tail convention as scipy.stats.ttest_ind.
Welch vs pooled
By default this page highlights Welch’s t-test, which does not assume equal variances and uses the Welch–Satterthwaite degrees of freedom. The pooled Student test is still shown because it matches classical textbooks and Excel’s T.TEST type 2 when variances are similar. Modern practice favors Welch unless you have a strong design reason to pool; Delacre et al. (2017) recommend Welch as the default for two-sample comparisons in behavioral research.
Welch t = (x̄₁ − x̄₂ − δ₀) / √(Sx₁²/n₁ + Sx₂²/n₂)
df_W = (Sx₁²/n₁ + Sx₂²/n₂)² / [(Sx₁²/n₁)²/(n₁−1) + (Sx₂²/n₂)²/(n₂−1)]
Pooled: sp² = ((n₁−1)Sx₁² + (n₂−1)Sx₂²)/(n₁+n₂−2), t = (x̄₁−x̄₂−δ₀) / (sp√(1/n₁+1/n₂)), df = n₁+n₂−2
Assumptions and checks
- Independent observations between and within groups (random sampling, no pairing).
- Approximate normality within each group matters most for small n; inspect histograms or run a Shapiro–Wilk test per group.
- Similar variances matter only for the pooled test; compare SDs, use Levene’s test or the F-test, or rely on Welch when unsure.
- If normality fails badly, consider the Mann–Whitney U test.
Worked example (Load example)
Group 1: 5, 6, 7, 8, 9 (x̄₁ = 7, Sx₁ = √2.5). Group 2: 3, 4, 5, 6, 7 (x̄₂ = 5, Sx₂ = √2.5). Mean difference = 2. Welch and pooled both give t = 2.0 with df = 8 and two-sided p ≈ 0.0805; the 95% CI for the difference is about (−0.306, 4.306). Cohen’s d ≈ 0.894.
Software equivalents
| Software | Command |
|---|---|
| Excel | T.TEST(array1, array2, tails, type): type 2 = homoscedastic, type 3 = Welch |
| R | t.test(x, y, var.equal = FALSE) # Welch; TRUE for pooled |
| Python | scipy.stats.ttest_ind(a, b, equal_var=False) |
| TI-84 | STAT → TESTS → 4:2-SampTTest (set Pooled: No for Welch) |
Common mistakes
- Using an independent test on paired/repeated measures — use the paired t-test instead.
- Choosing a one-sided alternative after seeing which mean is larger (inflates Type I error).
- Equating “not significant” with “no effect”; report the CI and effect size.
Frequently Asked Questions
Why show both Welch and pooled results?
So you can see whether the equal-variance assumption changes the conclusion; when group SDs or sizes differ, Welch df and p often differ from the pooled test.
Which method is the headline result?
Welch by default (unequal variances allowed); switch Headline method to pooled to match a classical Student test or Excel T.TEST type 2.
What does the variance ratio tell me?
It is the larger sample SD divided by the smaller; large ratios warn that pooling variances may be questionable, not a formal test by itself.
How is Cohen's d computed here?
Standardized mean difference using the pooled sample SD with n₁ + n₂ − 2 degrees of freedom.
Does a significant p-value mean the groups differ practically?
No — check Cohen's d or Hedges' g and the confidence interval width; large n can make tiny differences significant.
Can I test μ₁ − μ₂ = 5 instead of 0?
Yes — enter that value as δ₀; t uses (x̄₁ − x̄₂ − δ₀) / SE.
What if one group has n = 2?
The formulas still run, but normality and variance estimates are very unstable; treat p-values cautiously and prefer larger samples.
Excel T.TEST type 2 vs type 3?
Type 2 assumes equal variances (pooled); type 3 is Welch. Match this page's method to the Excel type you need.
Embed This Calculator
Add this free calculator to your course page or LMS.
Adjust the height value to fit your page.