Mann-Whitney U Test Calculator

Enter two independent samples to run the Mann-Whitney U (Wilcoxon rank-sum) test with pooled ranks, U1 and U2, z, exact or asymptotic p-value, effect sizes, and a ranking table.

Comma, space, or newline separated.

Hypotheses

H0: The two distributions are equal (no stochastic difference). Ha depends on the alternative: a two-sided test detects any shift; one-sided tests detect stochastic dominance of group A over B or the reverse. Equal shapes are not required, but interpreting U as a median difference needs similar shapes.

U1 = R1 − n1(n1 + 1)/2

R1 = sum of ranks of group A in the pooled sample

Z uses tie correction T = 1 − Σ(t³ − t)/(N³ − N) on ranks

Worked example (matches Load example)

Group A: 8, 10, 12, 15; group B: 5, 7, 9, 11. U1 = 13, U2 = 3, exact two-sided p = 0.2 at α = 0.05 (do not reject). Probability of superiority ≈ 0.8125; rank-biserial r ≈ 0.625.

When to use another test

Three or more groups: Kruskal-Wallis calculator. Paired measurements: Wilcoxon signed-rank test. Parametric mean comparison (planned): two-sample t-test.

Software equivalents

  • Excel: no built-in Mann-Whitney; use R/Python add-ins or rank manually.
  • R: wilcox.test(x, y, exact = TRUE, correct = TRUE)
  • Python: scipy.stats.mannwhitneyu(x, y, use_continuity=True, method="auto")

Related guides and calculators

For normal data the two-sample t-test compares the means instead, and for more than two groups the Kruskal-Wallis test is the rank-based counterpart of ANOVA. The Shapiro-Wilk test and the Levene test check normality and equal spread. Read parametric vs nonparametric tests and which statistical test to use.

Frequently Asked Questions

What does the Mann-Whitney U test compare?

It tests whether one group tends to produce larger values than the other (stochastic ordering), not whether medians differ unless both distributions have the same shape.

When is the exact p-value used here?

Exact null distribution (dynamic programming) is used when there are no ties and each group has at most 50 values — wider than scipy's method="auto" cutoff of 8×8 but chosen so the DP stays fast in the browser; otherwise the normal approximation with tie and continuity correction is reported.

How are tied values ranked?

Tied values receive the average of the ranks they would occupy if broken by a small amount, matching scipy.stats.rankdata(method="average").

What is probability of superiority?

U1 / (n1 n2) estimates P(a random draw from group A exceeds a draw from group B); it equals the common-language effect size for this test.

When should I use a two-sample t-test instead?

Use a t-test when both samples are independent, approximately normal, and you want inference about means; use Mann-Whitney when normality is doubtful or outliers dominate.

How do I run this in R or Python?

R: wilcox.test(x, y, exact = TRUE, correct = TRUE). Python: scipy.stats.mannwhitneyu(x, y, alternative="two-sided", method="auto", use_continuity=True).

What is rank-biserial correlation?

On this page it is 2U1/(n1 n2) − 1 = P(A > B) − P(A < B), a signed effect size between −1 and 1: positive when group A tends to be larger, negative when group B does.

Can I enter more than 10,000 values?

No. The calculator accepts at most 10,000 observations in total across both groups to keep ranking and exact DP tables feasible in the browser.

Embed This Calculator

Add this free calculator to your course page or LMS.

Adjust the height value to fit your page.