Statistics How-To
How to Run a T-Test in R and Python
In R, t.test(x, y) runs Welch's two-sample t-test. In Python, scipy.stats.ttest_ind(x, y, equal_var=False) does the same, but without equal_var=False SciPy assumes equal variances and runs the pooled test. For x = 5, 6, 7, 8, 9 and y = 3, 4, 5, 6, 7 both give t = 2, df = 8 and p = 0.0805. This guide shows the one-sample, two-sample and paired versions side by side.
Two independent groups
The data are x = 5, 6, 7, 8, 9 and y = 3, 4, 5, 6, 7, the same groups used in the t-test in Excel guide.
# R
x <- c(5, 6, 7, 8, 9)
y <- c(3, 4, 5, 6, 7)
t.test(x, y) # Welch: t = 2, df = 8, p-value = 0.08052
t.test(x, y, var.equal = TRUE) # pooled: same result for these groups
# Python
from scipy import stats
x = [5, 6, 7, 8, 9]
y = [3, 4, 5, 6, 7]
res = stats.ttest_ind(x, y, equal_var=False) # Welch
print(res.statistic, res.df, res.pvalue) # 2.0 8.0 0.0805
low, high = res.confidence_interval() # -0.306, 4.306
Welch and pooled agree here because the groups have the same size and the same variance. The 95% confidence interval for the difference in means is −0.306 to 4.306, which includes zero, in line with p = 0.0805. The statistic is positive because the mean of the first sample is larger.
Paired and one-sample tests
Six people are weighed before (82, 90, 75, 88, 79, 85) and after (78, 85, 74, 80, 77, 79) a programme:
# R
t.test(before, after, paired = TRUE) # t = 4.111, df = 5, p-value = 0.009255
# Python
res = stats.ttest_rel(before, after) # 4.111, p = 0.009255, df = 5
res.confidence_interval() # 1.624 to 7.043 (mean difference 4.333)
The one-sample test compares a mean with a fixed value. For x above tested against 5:
# R
t.test(x, mu = 5) # t = 2.8284, df = 4, p-value = 0.04742
# Python
stats.ttest_1samp(x, popmean=5) # 2.8284, p = 0.04742, df = 4
The 95% interval for the mean is 5.037 to 8.963, which excludes 5, in line with p below 0.05.
The arguments side by side
| Purpose | R t.test() | SciPy |
|---|---|---|
| Paired data | paired = TRUE | ttest_rel(a, b) |
| Pooled variance | var.equal = TRUE | ttest_ind(a, b) (the default) or equal_var=True |
| Welch's test | var.equal = FALSE (the default) | ttest_ind(a, b, equal_var=False) |
| Direction | alternative = "two.sided", "less" or "greater" | alternative='two-sided', 'less' or 'greater' |
| Hypothesised mean | mu = 5 | popmean=5 in ttest_1samp |
| Confidence level | conf.level = 0.95 | res.confidence_interval(confidence_level=0.95) |
The mean-difference direction is the same in both: a positive statistic means the first sample has the larger mean. For the example groups, greater gives a one-sided p-value of 0.0403 and less gives 0.9597.
Why the default matters, and a warning about outliers
The R documentation's own examples show both points. It annotates t.test(1:10, y = c(7:20)) with P = .00001855, and SciPy's Welch test on the same data gives t = −5.435, df = 21.98 and p = 1.855e-05. It then adds one value of 200 to the second group, t.test(1:10, y = c(7:20, 200)), annotates P = .1245 and remarks that the result is not significant anymore; SciPy gives t = −1.633, df = 14.16 and p = 0.1245. One outlier turns a very significant result into a non-significant one.
When you see such sensitivity, check the value against its source and run the rank-based alternative as well; see the Mann-Whitney U test and parametric vs nonparametric tests.
Mistakes to avoid
- Comparing R and Python without matching the variance option. Set var.equal and equal_var explicitly so both run the same test.
- Running the unpaired test on paired data. Use paired = TRUE or ttest_rel; see which statistical test to use.
- Choosing the one-sided alternative after seeing the data. Decide the direction beforehand; see one-tailed vs two-tailed tests.
- Reporting only the p-value. Report the estimate, its confidence interval and an effect size. R's t.test and SciPy's ttest_ind return the first two but not Cohen's d, so compute it from the group summaries with the effect size calculator.
Try the T-Test Calculator
Check any R or Python result with a one-sample, two-sample or paired t-test in the browser.
Try the Two-Sample T-Test Calculator
Compare two independent groups with Welch's correction and see the confidence interval.
Frequently Asked Questions
Why do R and Python give different p-values for the same two groups?
Almost always because of the default. R's t.test() runs Welch's test unless you set var.equal = TRUE, while SciPy's ttest_ind() assumes equal variances unless you pass equal_var=False. With unequal group sizes and unequal variances the two can differ a lot. Set the argument explicitly so both run the same test.
How do I run a one-tailed t-test?
In R add alternative = "greater" or alternative = "less", and in Python pass alternative='greater' or alternative='less' to the SciPy function. "Greater" tests whether the first sample's mean is larger than the second's. For the example groups the one-sided p-value is 0.0403 for greater and 0.9597 for less.
How do I get the confidence interval?
R prints it automatically as the 95 percent confidence interval, and you can change it with conf.level. In SciPy call .confidence_interval(confidence_level=0.95) on the result, which returns the low and high limits of the interval for the difference in means (available from SciPy 1.11).
How do I run a t-test on a data frame with a grouping column?
In R use the formula interface, t.test(score ~ group, data = df), where group has exactly two levels. In Python with pandas, split the column first, for example a = df.loc[df.group == 'A', 'score'] and b = df.loc[df.group == 'B', 'score'], and pass a and b to ttest_ind.
How do I check the assumptions before the t-test?
Plot the data, and use shapiro.test(x) in R or scipy.stats.shapiro(x) for a formal normality check. For equal variances use var.test(x, y) in R or scipy.stats.levene(x, y). With large samples the t-test tolerates moderate non-normality, and with small skewed samples a rank test is the safer choice.
What is the nonparametric alternative in R and Python?
For two independent groups use wilcox.test(x, y) in R or scipy.stats.mannwhitneyu(x, y) in Python. For paired data use wilcox.test(x, y, paired = TRUE) or scipy.stats.wilcoxon(x, y). They compare ranks rather than means, so a single extreme value cannot dominate the result.