Statistics Concepts

Sensitivity, Specificity, PPV and NPV Explained

Sensitivity is the share of people with a condition whom a test detects, specificity the share of people without it whom the test clears, and the positive and negative predictive values (PPV and NPV) are the shares of positive and negative results that are correct. Sensitivity and specificity belong to the test; the predictive values also depend on how common the condition is. A test with 95% sensitivity and 90% specificity has a PPV of only 8.8% when 1% of the people tested have the condition.

The 2×2 table of a diagnostic test

A test is judged against a reference standard that tells who really has the condition. Every person then falls in one of four cells.

Condition presentCondition absent
Test positiveTrue positive (TP)False positive (FP)
Test negativeFalse negative (FN)True negative (TN)

Sensitivity = TP / (TP + FN) (of the people with the condition, the share who test positive)

Specificity = TN / (TN + FP) (of the people without it, the share who test negative)

PPV = TP / (TP + FP) (of the positive results, the share that are correct)

NPV = TN / (TN + FN) (of the negative results, the share that are correct)

Accuracy = (TP + TN) / (TP + FN + FP + TN)

Sensitivity and specificity read down the columns, so they do not depend on how many people with and without the condition were tested. The predictive values read along the rows, so they do.

Worked example: 80, 20, 10, 890

A test is given to 1,000 people, 100 of whom have the condition. It is positive for 80 of them (TP = 80, FN = 20) and for 10 of the other 900 (FP = 10, TN = 890).

MeasureCalculationValue95% CI (Wilson)
Sensitivity80 / 1000.80.7112 to 0.8666
Specificity890 / 9000.98890.9797 to 0.994
PPV80 / 900.88890.8074 to 0.9385
NPV890 / 9100.9780.9663 to 0.9857
Accuracy970 / 1,0000.970.9575 to 0.9789

The interval of the sensitivity is wide, 0.71 to 0.87, because it rests on only 100 people with the condition, while the specificity uses 900 people and is far more precise. Always report how many people each measure is based on. The sensitivity and specificity calculator gives these intervals for any counts, and the proportion confidence interval calculator explains the Wilson method.

Why the predictive values depend on prevalence

Take a screening test with sensitivity 0.95 and specificity 0.90. Bayes' theorem gives the PPV for any prevalence p: sensitivity × p divided by sensitivity × p plus (1 − specificity) × (1 − p).

PrevalencePPVNPV
0.1%0.94%99.99%
1%8.76%99.94%
5%33.33%99.71%
10%51.35%99.39%
20%70.37%98.63%
50%90.48%94.74%

With the same test, a positive result is almost always a false alarm in a screening population where 1 in 1,000 people has the condition, and is right nine times out of ten where half the people have it. Reading a positive result as if the sensitivity were the probability of having the condition is the base-rate fallacy: it confuses P(positive | condition) with P(condition | positive). The Bayes' theorem article works through the reversal in general.

Natural frequencies: 10,000 people at a prevalence of 1%

Counts of people are easier to think about than conditional probabilities (Gigerenzer and Hoffrage, 1995). Send 10,000 people through the same test:

Test positiveTest negativeTotal
Condition present (1%)955100
Condition absent (99%)9908,9109,900
Total1,0858,91510,000

Of the 1,085 positive results only 95 are true, a PPV of 95 / 1,085 = 8.76%, while 8,910 of the 8,915 negative results are true, an NPV of 99.94%. The quickest sanity check on any diagnostic claim is to draw this table.

Likelihood ratios: from pre-test to post-test probability

LR+ = sensitivity / (1 − specificity), LR− = (1 − sensitivity) / specificity

Post-test odds = pre-test odds × LR, odds = probability / (1 − probability)

For sensitivity 0.95 and specificity 0.90, LR+ = 9.5 and LR− = 0.0556. At a prevalence of 1% the pre-test odds are 1 / 99, a positive result gives odds of 9.5 / 99 (a probability of 8.76%) and a negative result gives odds of 0.0556 / 99 (a probability of 0.056%). The likelihood ratios do not change with the prevalence, so one pair of numbers serves every population. As a rule of thumb, an LR+ above 10 or an LR− below 0.1 gives a large, often conclusive shift in the probability of the condition (Jaeschke et al., JAMA 1994).

Trade-offs and the same idea in hypothesis testing

  • Cut-offs. For a test on a continuous measurement, moving the cut-off raises sensitivity and lowers specificity, or the reverse. Each cut-off is one point of the ROC curve. Youden's J, sensitivity + specificity − 1, is one way to choose it: 0.85 for the screening test above.
  • Rule out, rule in. A negative result of a highly sensitive test rules the condition out (SnNout); a positive result of a highly specific test rules it in (SpPin).
  • Hypothesis tests. If "condition present" means that the alternative hypothesis is true, sensitivity is the power of the test and specificity is 1 − α; see statistical power and type I and type II errors.
  • Comparing tests. When two tests are applied to the same patients, compare their sensitivities with McNemar's test on the patients with the condition.

Common mistakes

  • Quoting the PPV of one study for another population. It holds only at the prevalence of the sample, so enter the prevalence you care about.
  • Computing the predictive values from a case-control study. The investigators chose the numbers with and without the condition, so the prevalence in the sample is arbitrary.
  • Trusting accuracy. When the condition is rare a test that is always negative has an accuracy of 99% at a prevalence of 1%, and a sensitivity of 0.
  • Ignoring the uncertainty. A sensitivity of 0.90 from 10 patients is compatible with a wide range of true values; see confidence intervals.
  • Mixing up the ratio scales. The diagnostic odds ratio, LR+ / LR−, is an odds ratio of test result and condition; see relative risk vs odds ratio for how odds ratios differ from risks.

Try the Sensitivity and Specificity Calculator

Sensitivity, specificity, PPV, NPV, likelihood ratios and confidence intervals from counts or from rates and prevalence.

Try the Bayes' Theorem Calculator

Update a prior probability with a test result: the same calculation as the positive predictive value.

Frequently Asked Questions

What is the difference between sensitivity and specificity?

Sensitivity is the share of people who have the condition and test positive, TP / (TP + FN). Specificity is the share of people who do not have it and test negative, TN / (TN + FP). A sensitive test rarely misses the condition; a specific test rarely raises a false alarm. With TP = 80, FN = 20, FP = 10 and TN = 890, sensitivity is 0.8 and specificity 0.9889.

How do I calculate PPV and NPV from sensitivity, specificity and prevalence?

PPV = sensitivity × p / [sensitivity × p + (1 − specificity) × (1 − p)] and NPV = specificity × (1 − p) / [specificity × (1 − p) + (1 − sensitivity) × p], with p the prevalence. For sensitivity 0.95, specificity 0.90 and p = 0.01, PPV = 0.0095 / 0.1085 = 0.0876 and NPV = 0.891 / 0.8915 = 0.9994.

Why is the PPV low although the test has high sensitivity and specificity?

Because most of the people tested do not have the condition, so even a small false positive rate produces many false positives. At a prevalence of 1% and a specificity of 90%, 990 of 9,900 healthy people out of 10,000 test positive, against 95 true positives from the 100 people with the condition. Only 95 of the 1,085 positive results are correct.

What is a good PPV or NPV?

There is no fixed threshold, because both depend on the prevalence and on what a wrong result costs. A screening test for a serious, treatable condition can accept a low PPV if a confirmatory test follows, but it needs a high NPV so that few cases are sent home. Judge the predictive values at the prevalence of the population where the test will be used.

Are sensitivity and recall, and PPV and precision, the same thing?

Yes. In machine learning and information retrieval sensitivity is called recall or the true positive rate, and the positive predictive value is called precision. The false positive rate is 1 − specificity. Like the PPV, precision depends on how common the positive class is.

What are likelihood ratios?

LR+ = sensitivity / (1 − specificity) says how much a positive result multiplies the odds of the condition, and LR− = (1 − sensitivity) / specificity how much a negative result multiplies them. For sensitivity 0.95 and specificity 0.90, LR+ = 9.5 and LR− = 0.0556. They do not depend on the prevalence, which makes them convenient for moving from a pre-test to a post-test probability.

Do sensitivity and specificity depend on the prevalence?

Not by definition: they are conditional on the true status. In practice they can differ between populations, because the mix of mild and severe cases and of easy and hard non-cases changes with the setting (spectrum bias). The predictive values change with the prevalence even when sensitivity and specificity do not.