Two Sample T-Test Calculator

Compare two independent sample means with Welch’s t-test (default) and the pooled Student test shown together, from pasted data or n, x̄, and Sx for each group.

Hub for quick tests: general t-test calculator. Check equal variances with the F-test or Levene test.

Comma- or space-separated values (≥ 2).

Comma- or space-separated values (≥ 2).

Hypotheses

H₀: μ₁ − μ₂ = δ₀ (often δ₀ = 0). The alternative follows your selection: two-sided, μ₁ − μ₂ > δ₀, or μ₁ − μ₂ < δ₀. The confidence interval is for μ₁ − μ₂ using the same method and tail convention as scipy.stats.ttest_ind.

Welch vs pooled

By default this page highlights Welch’s t-test, which does not assume equal variances and uses the Welch–Satterthwaite degrees of freedom. The pooled Student test is still shown because it matches classical textbooks and Excel’s T.TEST type 2 when variances are similar. Modern practice favors Welch unless you have a strong design reason to pool; Delacre et al. (2017) recommend Welch as the default for two-sample comparisons in behavioral research.

Welch t = (x̄₁ − x̄₂ − δ₀) / √(Sx₁²/n₁ + Sx₂²/n₂)

df_W = (Sx₁²/n₁ + Sx₂²/n₂)² / [(Sx₁²/n₁)²/(n₁−1) + (Sx₂²/n₂)²/(n₂−1)]

Pooled: sp² = ((n₁−1)Sx₁² + (n₂−1)Sx₂²)/(n₁+n₂−2), t = (x̄₁−x̄₂−δ₀) / (sp√(1/n₁+1/n₂)), df = n₁+n₂−2

Assumptions and checks

  • Independent observations between and within groups (random sampling, no pairing).
  • Approximate normality within each group matters most for small n; inspect histograms or run a Shapiro–Wilk test per group.
  • Similar variances matter only for the pooled test; compare SDs, use Levene’s test or the F-test, or rely on Welch when unsure.
  • If normality fails badly, consider the Mann–Whitney U test.

Worked example (Load example)

Group 1: 5, 6, 7, 8, 9 (x̄₁ = 7, Sx₁ = √2.5). Group 2: 3, 4, 5, 6, 7 (x̄₂ = 5, Sx₂ = √2.5). Mean difference = 2. Welch and pooled both give t = 2.0 with df = 8 and two-sided p ≈ 0.0805; the 95% CI for the difference is about (−0.306, 4.306). Cohen’s d ≈ 0.894.

Software equivalents

SoftwareCommand
ExcelT.TEST(array1, array2, tails, type): type 2 = homoscedastic, type 3 = Welch
Rt.test(x, y, var.equal = FALSE) # Welch; TRUE for pooled
Pythonscipy.stats.ttest_ind(a, b, equal_var=False)
TI-84STAT → TESTS → 4:2-SampTTest (set Pooled: No for Welch)

Common mistakes

  • Using an independent test on paired/repeated measures — use the paired t-test instead.
  • Choosing a one-sided alternative after seeing which mean is larger (inflates Type I error).
  • Equating “not significant” with “no effect”; report the CI and effect size.

Frequently Asked Questions

Why show both Welch and pooled results?

So you can see whether the equal-variance assumption changes the conclusion; when group SDs or sizes differ, Welch df and p often differ from the pooled test.

Which method is the headline result?

Welch by default (unequal variances allowed); switch Headline method to pooled to match a classical Student test or Excel T.TEST type 2.

What does the variance ratio tell me?

It is the larger sample SD divided by the smaller; large ratios warn that pooling variances may be questionable, not a formal test by itself.

How is Cohen's d computed here?

Standardized mean difference using the pooled sample SD with n₁ + n₂ − 2 degrees of freedom.

Does a significant p-value mean the groups differ practically?

No — check Cohen's d or Hedges' g and the confidence interval width; large n can make tiny differences significant.

Can I test μ₁ − μ₂ = 5 instead of 0?

Yes — enter that value as δ₀; t uses (x̄₁ − x̄₂ − δ₀) / SE.

What if one group has n = 2?

The formulas still run, but normality and variance estimates are very unstable; treat p-values cautiously and prefer larger samples.

Excel T.TEST type 2 vs type 3?

Type 2 assumes equal variances (pooled); type 3 is Welch. Match this page's method to the Excel type you need.

Embed This Calculator

Add this free calculator to your course page or LMS.

Adjust the height value to fit your page.