Cohen's Kappa Calculator
Enter a k×k agreement table or paired category ratings (2–8 levels) for two raters: Cohen's κ, SE under H₀ for the z test, general SE for the 95% CI, and linear/quadratic weighted κ.
Paste a table of counts, one row per line, the counts separated by tabs, commas or spaces. The grid takes the size of the paste, from 2×2 up to 8×8:
Related Calculators
McNemar's Test Calculator
Test paired binary outcomes: chi-square, corrected, exact and mid-p p-values, and the odds ratio of the discordant pairs.
Chi-Square Test of Independence Calculator
Run Pearson's chi-square test on an r×c contingency table with expected counts, residuals, Cramér's V, and p-value.
Correlation Coefficient Calculator
Calculate Pearson's r with its p-value, Fisher z confidence interval, R², covariance, worked steps and a scatter plot.
What kappa corrects for
κ = (po − pe) / (1 − pe) adjusts for agreement expected by chance. Landis & Koch benchmarks (0.41–0.60 moderate, etc.) are arbitrary—report κ with CI. See McNemar for paired binary disagreement.
κ = (p_o − p_e) / (1 − p_e)
Kappa is for raters who assign items to categories. To judge the internal consistency of a questionnaire scale answered by many respondents, use the Cronbach's alpha calculator.
Worked example
2×2 table in Load example: po = 0.85, κ = 0.70, strong agreement beyond chance for H₀: κ = 0.
Related guides and calculators
Kappa measures agreement between two raters; for the accuracy of a test against a reference see sensitivity and specificity, for a change in paired ratings the McNemar test, and for the reliability of a scale with several items Cronbach's alpha (see Cronbach's alpha explained). The chi-square test of independence tests association, not agreement.
Frequently Asked Questions
Which SE is used for the z test?
SE under H₀ (Fleiss, Cohen & Everitt 1969); the 95% CI uses the general SE.
Kappa paradox?
High p_o but low κ can happen when marginal totals are skewed and p_e is large.
Ordinal categories?
Use linear or quadratic weighted κ to penalize disagreements by distance.
sklearn match?
Unweighted κ matches sklearn.metrics.cohen_kappa_score; weights match weights='linear'|'quadratic'.
Paired ratings list?
Switch to paired ratings and enter parallel lists with category codes 1…k; the calculator builds the k×k table with buildAgreementTable before computing κ.
What is maximum κ?
With fixed marginals, κ cannot exceed (P_max − p_e)/(1 − p_e), where P_max is perfect agreement subject to those margins (Sim & Wright 2005).
Is κ = 1 perfect agreement?
Only if every count lies on the diagonal; κ = 1 implies p_o = 1.
How do I interpret negative κ?
Negative κ means observed agreement is below chance; check for systematic disagreement or coding errors.
Embed This Calculator
Add this free calculator to your course page or LMS.
Adjust the height value to fit your page.