Statistics How-To

How to Run a T-Test in R and Python

In R, t.test(x, y) runs Welch's two-sample t-test. In Python, scipy.stats.ttest_ind(x, y, equal_var=False) does the same, but without equal_var=False SciPy assumes equal variances and runs the pooled test. For x = 5, 6, 7, 8, 9 and y = 3, 4, 5, 6, 7 both give t = 2, df = 8 and p = 0.0805. This guide shows the one-sample, two-sample and paired versions side by side.

Two independent groups

The data are x = 5, 6, 7, 8, 9 and y = 3, 4, 5, 6, 7, the same groups used in the t-test in Excel guide.

# R

x <- c(5, 6, 7, 8, 9)

y <- c(3, 4, 5, 6, 7)

t.test(x, y) # Welch: t = 2, df = 8, p-value = 0.08052

t.test(x, y, var.equal = TRUE) # pooled: same result for these groups

# Python

from scipy import stats

x = [5, 6, 7, 8, 9]

y = [3, 4, 5, 6, 7]

res = stats.ttest_ind(x, y, equal_var=False) # Welch

print(res.statistic, res.df, res.pvalue) # 2.0 8.0 0.0805

low, high = res.confidence_interval() # -0.306, 4.306

Welch and pooled agree here because the groups have the same size and the same variance. The 95% confidence interval for the difference in means is −0.306 to 4.306, which includes zero, in line with p = 0.0805. The statistic is positive because the mean of the first sample is larger.

Paired and one-sample tests

Six people are weighed before (82, 90, 75, 88, 79, 85) and after (78, 85, 74, 80, 77, 79) a programme:

# R

t.test(before, after, paired = TRUE) # t = 4.111, df = 5, p-value = 0.009255

# Python

res = stats.ttest_rel(before, after) # 4.111, p = 0.009255, df = 5

res.confidence_interval() # 1.624 to 7.043 (mean difference 4.333)

The one-sample test compares a mean with a fixed value. For x above tested against 5:

# R

t.test(x, mu = 5) # t = 2.8284, df = 4, p-value = 0.04742

# Python

stats.ttest_1samp(x, popmean=5) # 2.8284, p = 0.04742, df = 4

The 95% interval for the mean is 5.037 to 8.963, which excludes 5, in line with p below 0.05.

The arguments side by side

PurposeR t.test()SciPy
Paired datapaired = TRUEttest_rel(a, b)
Pooled variancevar.equal = TRUEttest_ind(a, b) (the default) or equal_var=True
Welch's testvar.equal = FALSE (the default)ttest_ind(a, b, equal_var=False)
Directionalternative = "two.sided", "less" or "greater"alternative='two-sided', 'less' or 'greater'
Hypothesised meanmu = 5popmean=5 in ttest_1samp
Confidence levelconf.level = 0.95res.confidence_interval(confidence_level=0.95)

The mean-difference direction is the same in both: a positive statistic means the first sample has the larger mean. For the example groups, greater gives a one-sided p-value of 0.0403 and less gives 0.9597.

Why the default matters, and a warning about outliers

The R documentation's own examples show both points. It annotates t.test(1:10, y = c(7:20)) with P = .00001855, and SciPy's Welch test on the same data gives t = −5.435, df = 21.98 and p = 1.855e-05. It then adds one value of 200 to the second group, t.test(1:10, y = c(7:20, 200)), annotates P = .1245 and remarks that the result is not significant anymore; SciPy gives t = −1.633, df = 14.16 and p = 0.1245. One outlier turns a very significant result into a non-significant one.

When you see such sensitivity, check the value against its source and run the rank-based alternative as well; see the Mann-Whitney U test and parametric vs nonparametric tests.

Mistakes to avoid

  • Comparing R and Python without matching the variance option. Set var.equal and equal_var explicitly so both run the same test.
  • Running the unpaired test on paired data. Use paired = TRUE or ttest_rel; see which statistical test to use.
  • Choosing the one-sided alternative after seeing the data. Decide the direction beforehand; see one-tailed vs two-tailed tests.
  • Reporting only the p-value. Report the estimate, its confidence interval and an effect size. R's t.test and SciPy's ttest_ind return the first two but not Cohen's d, so compute it from the group summaries with the effect size calculator.

Try the T-Test Calculator

Check any R or Python result with a one-sample, two-sample or paired t-test in the browser.

Try the Two-Sample T-Test Calculator

Compare two independent groups with Welch's correction and see the confidence interval.

Frequently Asked Questions

Why do R and Python give different p-values for the same two groups?

Almost always because of the default. R's t.test() runs Welch's test unless you set var.equal = TRUE, while SciPy's ttest_ind() assumes equal variances unless you pass equal_var=False. With unequal group sizes and unequal variances the two can differ a lot. Set the argument explicitly so both run the same test.

How do I run a one-tailed t-test?

In R add alternative = "greater" or alternative = "less", and in Python pass alternative='greater' or alternative='less' to the SciPy function. "Greater" tests whether the first sample's mean is larger than the second's. For the example groups the one-sided p-value is 0.0403 for greater and 0.9597 for less.

How do I get the confidence interval?

R prints it automatically as the 95 percent confidence interval, and you can change it with conf.level. In SciPy call .confidence_interval(confidence_level=0.95) on the result, which returns the low and high limits of the interval for the difference in means (available from SciPy 1.11).

How do I run a t-test on a data frame with a grouping column?

In R use the formula interface, t.test(score ~ group, data = df), where group has exactly two levels. In Python with pandas, split the column first, for example a = df.loc[df.group == 'A', 'score'] and b = df.loc[df.group == 'B', 'score'], and pass a and b to ttest_ind.

How do I check the assumptions before the t-test?

Plot the data, and use shapiro.test(x) in R or scipy.stats.shapiro(x) for a formal normality check. For equal variances use var.test(x, y) in R or scipy.stats.levene(x, y). With large samples the t-test tolerates moderate non-normality, and with small skewed samples a rank test is the safer choice.

What is the nonparametric alternative in R and Python?

For two independent groups use wilcox.test(x, y) in R or scipy.stats.mannwhitneyu(x, y) in Python. For paired data use wilcox.test(x, y, paired = TRUE) or scipy.stats.wilcoxon(x, y). They compare ranks rather than means, so a single extreme value cannot dominate the result.