Sensitivity and Specificity Calculator
Enter the counts of a diagnostic test (TP, FN, FP, TN), or its sensitivity, specificity and the prevalence, to get sensitivity, specificity, PPV, NPV, likelihood ratios, accuracy and the diagnostic odds ratio, with confidence intervals for the counts.
Sensitivity and specificity describe the test; the predictive values describe a result in a population and change with its prevalence, as the tables below the results show. See sensitivity, specificity, PPV and NPV explained and the Bayes' theorem calculator. Educational only, not medical advice.
Related Calculators
Bayes' Theorem Calculator
Update a prior probability with new evidence to get the posterior P(A|B) and total P(B).
Confidence Interval for a Proportion Calculator
Estimate one proportion or the difference of two from counts with the Wald (1-PropZInt, 2-PropZInt), Wilson, Agresti–Coull, Clopper–Pearson, Jeffreys and Newcombe methods compared.
McNemar's Test Calculator
Test paired binary outcomes: chi-square, corrected, exact and mid-p p-values, and the odds ratio of the discordant pairs.
What sensitivity and specificity measure
A diagnostic test is judged against a reference standard, which says who really has the condition. The sensitivity (true positive rate, recall) is the share of people with the condition whom the test detects: TP / (TP + FN). The specificity (true negative rate) is the share of people without the condition whom the test clears: TN / (TN + FP). Both describe the test itself and do not depend on how common the condition is.
The positive predictive value (PPV) answers the question a patient asks: given a positive result, how likely is the condition? It is TP / (TP + FP), and the negative predictive value (NPV), TN / (TN + FN), is its counterpart for a negative result. Unlike sensitivity and specificity, both depend on the prevalence of the condition in the people tested.
How to use the sensitivity and specificity calculator
- With TP, FN, FP, TN counts, enter the four cells of the test result against the reference standard: true positives, false negatives, false positives and true negatives. Choose the confidence level and the interval for the proportions.
- With Sensitivity, specificity, prevalence, enter three numbers between 0 and 1, for example a published sensitivity and specificity and the prevalence in your population. You get the predictive values without a sample, so there are no confidence intervals.
- Press Calculate Diagnostic Metrics. The results follow every later change.
- Under the results, type another prevalence to see the PPV, the NPV and the expected results for 1,000 people. Load example enters the worked example below, and Load screening example a test for a rare condition.
Sensitivity and specificity formulas
Sensitivity = TP / (TP + FN)
Specificity = TN / (TN + FP)
PPV = TP / (TP + FP), NPV = TN / (TN + FN)
False positive rate = 1 − specificity, false negative rate = 1 − sensitivity
Accuracy = (TP + TN) / (TP + FN + FP + TN)
LR+ = sensitivity / (1 − specificity), LR− = (1 − sensitivity) / specificity
Diagnostic odds ratio = LR+ / LR− = (TP × TN) / (FP × FN)
Youden's J = sensitivity + specificity − 1
PPV at prevalence p = sensitivity × p / [ sensitivity × p + (1 − specificity)(1 − p) ]
The intervals for sensitivity, specificity, PPV, NPV and accuracy treat each as a binomial proportion: the Wilson score interval (Wilson 1927), which has better coverage than the textbook Wald interval, or the exact Clopper-Pearson interval. The interval for a likelihood ratio uses the log method of Simel, Samsa and Matchar (1991), and the one for the diagnostic odds ratio Woolf's logit method. The log intervals need all four counts to be above 0.
Worked example: 80, 20, 10, 890
A test is given to 1,000 people, 100 of whom have the condition. Press Load example to enter TP = 80, FN = 20, FP = 10 and TN = 890.
Working
Sensitivity = 80 / 100 = 0.8, with a Wilson 95% interval of 0.7112 to 0.8666.
Specificity = 890 / 900 = 0.9889, with a Wilson 95% interval of 0.9797 to 0.994.
PPV = 80 / 90 = 0.8889 (0.8074 to 0.9385) and NPV = 890 / 910 = 0.978 (0.9663 to 0.9857).
Accuracy = 970 / 1,000 = 0.97 (0.9575 to 0.9789) and Youden's J = 0.8 + 0.9889 − 1 = 0.7889.
LR+ = 0.8 / (10 / 900) = 72 (38.5741 to 134.3906), LR− = 0.2 / (890 / 900) = 0.2022 (0.1367 to 0.2993) and the diagnostic odds ratio = (80 × 890) / (10 × 20) = 356 (161.1148 to 786.6192).
A positive result multiplies the odds of the condition by 72 and a negative result multiplies them by 0.2022. The Clopper-Pearson interval of the sensitivity is a little wider, 0.7082 to 0.8733.
Why the positive predictive value depends on prevalence
Take the same test, sensitivity 0.8 and specificity 0.9889, and apply it where the condition is less common than in the sample. The false positives come from the people without the condition, and there are far more of them when the condition is rare, so they swamp the true positives.
| Prevalence | PPV | NPV |
|---|---|---|
| 1% | 42.13% | 99.8% |
| 5% | 79.14% | 98.95% |
| 10% | 88.9% | 97.8% |
| 20% | 94.74% | 95.19% |
| 50% | 98.63% | 83.18% |
The effect is stronger for a screening test. Press Load screening example: a test with sensitivity 0.95 and specificity 0.90 sounds excellent, but at a prevalence of 1% the PPV is only 8.76%. Out of 1,000 people 10 have the condition and 9.5 of them test positive; of the 990 without it 99 also test positive, so 9.5 of the 108.5 positive results are correct. A negative result is reassuring, NPV 99.94%, so screening tests are usually followed by a more specific confirmatory test. The Bayes' theorem calculator does the same calculation with a prior probability.
Likelihood ratios and the diagnostic odds ratio
A likelihood ratio says how much a result changes the odds of the condition: post-test odds = pre-test odds × LR. LR+ tells you how much a positive result raises them and LR− how much a negative result lowers them. Unlike the predictive values, the likelihood ratios do not depend on the prevalence. For the screening example LR+ = 9.5 and LR− = 0.0556: at a prevalence of 1% the pre-test odds are 1 / 99, a positive result gives odds of 9.5 / 99, a probability of 8.76%, and a negative result lowers the probability to 0.056%.
As a rule of thumb a positive likelihood ratio above 10 or a negative one below 0.1 gives a large, often conclusive change in the probability of the condition (Jaeschke et al. 1994; Deeks and Altman 2004), while values between 0.5 and 2 rarely change anything. The diagnostic odds ratio, LR+ / LR−, is a single number for the discrimination of the test: 1 means a useless test. It is the odds ratio of the 2×2 table of test result and condition.
How to read the result
- Trade-off. Moving the cut-off of a test raises sensitivity at the cost of specificity or the reverse, and each cut-off is one point of the ROC curve. Youden's J, sensitivity + specificity − 1, is one way to pick a cut-off.
- Rule out and rule in. The mnemonic SnNout says that a negative result of a highly sensitive test rules the condition out; SpPin says that a positive result of a highly specific test rules it in.
- Predictive values need the right prevalence. The PPV and NPV of a table of counts hold for the prevalence of the sample. A study that chose its numbers of diseased and healthy subjects, such as a case-control study, has an arbitrary prevalence, so its PPV and NPV mean nothing: enter the prevalence of the population you care about instead.
- Accuracy can mislead. When the condition is rare, a test that is always negative is right for almost everyone: at a prevalence of 1% its accuracy is 99%, with a sensitivity of 0.
- Intervals. With few subjects with the condition the interval for the sensitivity is wide, and with few without it the interval for the specificity is wide, whatever the estimates are.
Sensitivity and specificity in Excel, R, Python and SPSS
| Software | How to get it |
|---|---|
| Excel | With TP, FN, FP, TN in A2:D2: sensitivity =A2/(A2+B2), specificity =D2/(D2+C2), PPV =A2/(A2+C2), NPV =D2/(D2+B2), accuracy =(A2+D2)/SUM(A2:D2), LR+ =(A2/(A2+B2))/(1-D2/(D2+C2)) |
| R | Plain arithmetic on the four counts: tp / (tp + fn), tn / (tn + fp), tp / (tp + fp) and tn / (tn + fn). The packages epiR (epi.tests) and caret (confusionMatrix) report the same measures, epiR with confidence intervals |
| Python | from sklearn.metrics import confusion_matrix, recall_score; tn, fp, fn, tp = confusion_matrix(y_true, y_pred).ravel(); sensitivity = recall_score(y_true, y_pred); specificity = recall_score(y_true, y_pred, pos_label=0). Interval: statsmodels.stats.proportion.proportion_confint(80, 100, method="wilson") |
| SPSS | Analyze > Descriptive Statistics > Crosstabs gives the 2×2 table of test result against reference; Analyze > Classify > ROC Curve lists sensitivity and 1 − specificity for every cut-off of a continuous test |
Frequently Asked Questions
What are sensitivity and specificity?
Sensitivity is the share of people with the condition whom the test correctly calls positive, TP / (TP + FN). Specificity is the share of people without the condition whom it correctly calls negative, TN / (TN + FP). Together they describe how well a test separates the two groups.
How do I calculate sensitivity and specificity?
Arrange the results in a 2×2 table against a reference standard and divide. With TP = 80, FN = 20, FP = 10 and TN = 890, sensitivity is 80 / 100 = 0.8 and specificity is 890 / 900 = 0.9889.
What are the positive and negative predictive values?
The PPV is the probability that a person with a positive result has the condition, TP / (TP + FP); the NPV is the probability that a person with a negative result does not, TN / (TN + FN). In the worked example they are 0.8889 and 0.978.
Why does the positive predictive value change with prevalence?
Because the false positives come from the people without the condition. When the condition is rare there are many more of them than people with it, so even a small false positive rate produces more false than true positives. With sensitivity 0.95 and specificity 0.90 the PPV is 8.76% at a prevalence of 1% and 90.48% at a prevalence of 50%.
What is a good sensitivity and specificity?
There is no universal threshold; it depends on the cost of a missed case against that of a false alarm. A screening test should be highly sensitive so that few cases are missed, a confirmatory test highly specific so that few healthy people are treated. Values above 0.9 are usually considered good, but a test also has to be useful at the prevalence where it will be used.
What are SnNout and SpPin?
Mnemonics for using a test: a highly Sensitive test that is Negative rules the condition out (SnNout), and a highly Specific test that is Positive rules it in (SpPin). They work because a highly sensitive test has few false negatives and a highly specific one few false positives.
What are likelihood ratios and how do I interpret them?
LR+ = sensitivity / (1 − specificity) and LR− = (1 − sensitivity) / specificity show how much a positive or a negative result changes the odds of the condition. LR+ above 10 or LR− below 0.1 give large, often conclusive changes; values close to 1 change little. They do not depend on the prevalence.
Can I calculate the PPV from sensitivity, specificity and prevalence?
Yes: PPV = sensitivity × p / [sensitivity × p + (1 − specificity)(1 − p)] with p the prevalence. Choose the input mode with sensitivity, specificity and prevalence, or type a prevalence under the results of a table of counts. It is Bayes' theorem applied to a test result.
What is the difference between sensitivity and the positive predictive value?
They condition on different things. Sensitivity is the probability of a positive result given the condition, and the PPV is the probability of the condition given a positive result. The second depends on the prevalence, the first does not, and confusing the two is the base-rate fallacy.
Is sensitivity the same as recall, and what is the false positive rate?
Yes: sensitivity is called recall or the true positive rate in machine learning. The false positive rate is 1 − specificity, the share of people without the condition who test positive, and the false negative rate is 1 − sensitivity. The ROC curve plots sensitivity against the false positive rate.
How are the confidence intervals calculated?
Sensitivity, specificity, PPV, NPV and accuracy are proportions, so they use the Wilson score interval or the exact Clopper-Pearson interval, as you choose. The likelihood ratios use the log method of Simel and colleagues (1991) and the diagnostic odds ratio the Woolf interval. All are computed at the confidence level you select.
Why is there no interval for the likelihood ratios when a count is zero?
The log intervals need TP, FN, FP and TN all above 0, because with a zero count the ratio is 0 or infinite and its variance cannot be estimated. The calculator says so and still shows the estimates that exist; the intervals of the proportions are available whenever they can be computed.
Can I use the predictive values of a case-control diagnostic study?
Not directly. In such a study the numbers of subjects with and without the condition are chosen by the investigators, so the prevalence in the sample is arbitrary and the PPV and NPV computed from it are meaningless. Use its sensitivity and specificity with the prevalence of the population where the test will be used.
How do I calculate sensitivity and specificity in Excel?
With TP, FN, FP and TN in cells A2 to D2, type =A2/(A2+B2) for the sensitivity and =D2/(D2+C2) for the specificity. The PPV is =A2/(A2+C2) and the NPV =D2/(D2+B2). The calculator adds the likelihood ratios, the diagnostic odds ratio and the confidence intervals.
Embed This Calculator
Add this free calculator to your course page or LMS.
Adjust the height value to fit your page.