McNemar's Test Calculator

Enter the 2×2 table of paired binary outcomes to test whether the proportion changed between the two classifications. McNemar's test uses only the discordant pairs b and c; you get the chi-square, continuity-corrected, exact and mid-p p-values and the odds ratio b / c with its intervals.

For the same subjects measured twice, or for matched pairs. For two independent groups use the chi-square test of independence or Fisher's exact test.

Paste a 2×2 table of counts, one row per line, the counts separated by tabs, commas or spaces:

What McNemar's test does

McNemar's test compares two proportions measured on the same subjects, or on matched pairs: a symptom before and after a treatment, a diagnosis by two tests on the same patients, an answer to the same question at two times. Each subject falls in one cell of a 2×2 table. The subjects who gave the same answer twice (cells a and d) say nothing about a change; the test looks only at the discordant pairs, b (first positive, second negative) and c (first negative, second positive). If nothing changed on average, b and c should be about equal.

Ignoring the pairing is the common mistake. The chi-square test of independence assumes that every observation is independent; with the same subjects counted twice it tests whether the first answer predicts the second, which it nearly always does, and not whether the proportion changed.

DesignTest
Two independent groups, binary outcomeChi-square test of independence, or Fisher's exact test for small counts
The same subjects twice, binary outcomeMcNemar's test (this page)
The same subjects twice, numeric outcomePaired t-test, or the Wilcoxon signed-rank test
Two raters classifying the same subjectsCohen's kappa for agreement, McNemar's test for a systematic difference

How to use the McNemar test calculator

  1. Enter the four counts of the 2×2 table of pairs. Rows are the first classification and columns the second; put the same category (for example positive) in row 1 and column 1. Rename the rows and columns for your data, or paste a table.
  2. Choose the confidence level of the intervals for the odds ratio b / c.
  3. Press Calculate McNemar Test. The results follow every later change to the table.
  4. Load example enters the worked example below and Load small-sample example a table in which the p-values fall on both sides of 0.05. Copy link to this calculation shares your table and settings.

McNemar test formulas

χ² = (b − c)² / (b + c), 1 degree of freedom

With Edwards' continuity correction: χ² = (|b − c| − 1)² / (b + c)

Exact: p = min(1, 2 × P(X ≤ min(b, c))), X ~ Binomial(b + c, 0.5)

Mid-p: p = 2 × P(X < min(b, c)) + P(X = min(b, c)), and 1 − P(X = b) / 2 when b = c

Discordant odds ratio = b / c, Wald interval: exp( ln(b / c) ± z × √(1/b + 1/c) )

Under the null hypothesis that the two proportions are equal, each discordant pair is equally likely to be a b or a c, so b follows a binomial distribution with b + c trials and probability 0.5. The exact test uses that distribution directly (it is the sign test of the pairs that changed); the chi-square test is its large-sample approximation. The continuity correction of Edwards (1948) is applied only when b ≠ c, as R's mcnemar.test does: with b = c the bare formula would give a statistic above 0 for data that show no change at all. The exact interval is the Clopper-Pearson interval of b / (b + c) turned into odds.

Worked example: 10, 5, 3, 12

30 patients are classified positive or negative before and after a treatment. Press Load example: 10 are positive both times, 12 negative both times, 5 changed from positive to negative (b) and 3 from negative to positive (c).

Working

The 22 patients who did not change do not enter the test: b + c = 8.

χ² = (5 − 3)² / 8 = 0.5, p = 0.4795.

With the continuity correction: χ² = (|5 − 3| − 1)² / 8 = 0.125, p = 0.7237.

Exact: p = 2 × P(X ≤ 3) for X ~ Binomial(8, 0.5) = 0.7266. Mid-p = 0.5078.

The odds ratio b / c = 5 / 3 = 1.6667, with a Wald 95% interval of 0.3983 to 6.9739 and an exact 95% interval of 0.3243 to 10.7325.

Every p-value is far above 0.05 and both intervals contain 1: there is no evidence that the proportion of positive patients changed. The share positive fell from 15 / 30 = 50% to 13 / 30 = 43.33%, a difference of (5 − 3) / 30 = 6.67 percentage points that eight changed pairs cannot distinguish from chance.

Which p-value should you report?

With many discordant pairs the four p-values agree. With few they do not, and the choice can decide whether a result is significant. Load small-sample example enters the study of Bentur et al. (b = 1, c = 7) that Fagerland, Lydersen and Laake use in their Table 6:

P-valueValueAt the 0.05 level
Chi-square, no correction0.0339significant
Chi-square, continuity correction0.0771not significant
Exact binomial0.0703not significant
Mid-p0.0391significant

The exact test guarantees that its error rate never exceeds the nominal level, but for small counts it is well below it: the test is conservative and often fails to find a real change. The continuity correction has the same problem. Fagerland and colleagues compared the tests by simulation and recommend the mid-p test, which counts only half the probability of the observed table, and the asymptotic test without correction; they advise against the exact conditional test and the corrected test. For the study of Cavo et al. (b = 6, c = 16) the values are 0.0330, 0.0550, 0.0525 and 0.0347 in the same order. Whichever you choose, decide before you look at the data and say which one you used.

The odds ratio b / c

Among the pairs that changed, b / c is the odds that a pair moved in one direction rather than the other, and it is the conditional odds ratio of the matched-pairs design. b / c = 1 means no net change; with b = 5 and c = 3 it is 1.6667, so a change from positive to negative was 1.67 times as likely as the reverse. The Wald interval needs b and c above 0; the exact interval works when one of them is 0 (its limit is then 0 or infinity). To describe the size of the change, also report the two proportions, which the result text under the cards gives.

McNemar test in Excel, R, Python and SPSS

SoftwareHow to get it
ExcelThere is no built-in McNemar function. With b in B2 and c in C2: chi-square =(B2-C2)^2/(B2+C2), its p-value =CHISQ.DIST.RT((B2-C2)^2/(B2+C2),1), exact p-value =MIN(1,2*BINOM.DIST(MIN(B2,C2),B2+C2,0.5,TRUE))
Rmcnemar.test(matrix(c(10, 3, 5, 12), nrow = 2)) gives χ² = 0.125 and p = 0.7237 for the worked example, with the continuity correction that R applies by default; correct = FALSE gives 0.5 and 0.4795. binom.test(3, 8, 0.5)$p.value is the exact p-value, and exact2x2::mcnemar.exact adds the exact interval
Pythonfrom statsmodels.stats.contingency_tables import mcnemar; mcnemar([[10, 5], [3, 12]], exact=False, correction=False) gives 0.5 and 0.4795; exact=True gives the binomial p-value 0.7266. With correction=True statsmodels also corrects b = c, which R does not
SPSSAnalyze > Descriptive Statistics > Crosstabs > Statistics > McNemar reports an exact binomial significance. The PROPORTIONS command with /PAIREDSAMPLES offers the exact, mid-p, McNemar and corrected McNemar tests

Frequently Asked Questions

What is McNemar's test used for?

It tests whether the proportion of a binary outcome differs between two measurements on the same subjects or matched pairs, for example before and after a treatment or between two diagnostic tests applied to the same patients. It is the paired counterpart of the chi-square test for two proportions.

When should I use McNemar's test instead of the chi-square test?

Whenever the two classifications come from the same subjects or matched pairs. The chi-square test of independence assumes independent observations; applied to paired data it tests whether the first answer predicts the second, not whether the proportion changed, and it is usually wrong.

How do I set up the 2×2 table for McNemar's test?

Rows are the first classification and columns the second, with the same category first in both. Cell a counts subjects positive twice, b those positive then negative, c those negative then positive and d those negative twice. Only b and c enter the test.

What are discordant pairs?

The pairs whose two classifications differ: b and c. Concordant pairs (a and d) give the same answer twice and carry no information about a change. McNemar's test is a test on b + c pairs, so its power depends on how many pairs changed, not on how many subjects there are.

Which p-value should I report: chi-square, corrected, exact or mid-p?

Fagerland, Lydersen and Laake (2013) recommend the mid-p test and the chi-square test without correction, and advise against the exact conditional and the continuity-corrected tests because they are conservative. With many discordant pairs the four values are close. Choose one before you look at the data and report which one.

What does the continuity correction do?

Edwards' correction subtracts 1 from |b − c| before squaring, which makes the chi-square approximation of the discrete binomial distribution more cautious. It gives larger p-values and is conservative. The calculator, like R, applies it only when b ≠ c; statsmodels applies it even when b = c, which gives a statistic of 1 / (b + c) for data with no change.

How many discordant pairs do I need?

At least one: with b = c = 0 the statistic is 0/0 and the test cannot be done. The chi-square p-values are large-sample approximations, and a common rule asks for at least 25 discordant pairs. With fewer, use the mid-p or the exact p-value, which do not rely on an approximation.

What does the odds ratio b / c tell me?

Among the pairs that changed, it is how much more often they moved from row 1 to row 2 than the other way. It equals 1 when b = c. An interval that excludes 1 agrees with a significant test. The exact interval is available even when b or c is 0, the Wald interval is not.

Is McNemar's test one-tailed or two-tailed?

The p-values are two-sided. If you predicted the direction of the change before seeing the data and the data go that way, half of the exact, mid-p or uncorrected chi-square p-value is the one-sided p-value; if they go the other way, it is above 0.5.

How is McNemar's test related to the sign test?

The exact McNemar p-value is the two-sided sign test p-value of the b + c pairs that changed, with the unchanged pairs dropped as ties: each changed pair is a plus or a minus, and under the null hypothesis both are equally likely. The sign test calculator on this site gives the same p-value for data with b differences of one sign and c of the other.

Can I use McNemar's test to compare two diagnostic tests?

Yes, when both tests are applied to the same patients. Run it on the patients with the condition to compare the sensitivities and on those without it to compare the specificities. The sensitivity and specificity calculator gives each test's measures with intervals.

What if I have more than two categories or more than two measurements?

For a k × k table of paired categorical data the extension is the McNemar-Bowker test of symmetry, or the Stuart-Maxwell test of marginal homogeneity. For a binary outcome measured three or more times on the same subjects use Cochran's Q test. This calculator handles the 2×2 case.

How do I run McNemar's test in Excel, R, Python and SPSS?

Excel has no function, but the statistic and the p-values take one formula each from b and c. In R use mcnemar.test, in Python statsmodels.stats.contingency_tables.mcnemar, and in SPSS Crosstabs with the McNemar statistic. The table above gives the calls and the results for the worked example.

Embed This Calculator

Add this free calculator to your course page or LMS.

Adjust the height value to fit your page.