Cohen’s Kappa
Chance-corrected agreement between two raters, from a 2×2 table.
Observed agreement 85%, chance agreement 56% across 100 cases.
The raw percentage the two raters agreed on — before correcting for chance.
Kappa corrects raw agreement for the agreement expected by chance: κ = (p_o − p_e) ÷ (1 − p_e). With 25 both-yes, 5 and 10 disagreements, and 60 both-no, observed agreement is 85% and chance agreement 56%, giving κ = .29 ÷ .44 = .66 — substantial.
Why raw agreement is not enough
Two coders who agree on 85% of cases sound reliable. But if a category is rare — say only 5% of cases are positive — two coders who both guessed “negative” every time would agree on 95% of cases while sharing no judgement at all. Raw agreement rewards skewed categories, which is exactly where reliability matters most.
Kappa removes that free agreement. It calculates how often the two raters would coincide by chance alone, given how often each of them uses each category, and then asks how much of the remaining possible agreement was actually achieved. A kappa of .66 means the raters captured 66% of the agreement available beyond chance.
Reading the table
The four cells are the two agreements — both yes, both no — and the two disagreements. Only the diagonal counts towards observed agreement, but all four shape the chance term, because the marginal totals describe how freely each rater used each category.
p_o is observed agreement, p_e the agreement expected by chance from the marginal totals. Kappa runs from −1 to 1, where 0 is exactly chance-level agreement.
Worked example: 25 / 5 / 10 / 60
Observed first, chance second, then the ratio:
- 1 Build the 2×2 table. Both yes = 25, rater 1 yes only = 5, rater 2 yes only = 10, both no = 60. Total n = 100.
- 2 Compute observed agreement. (25 + 60) ÷ 100 = .85 — the raters agreed on 85 of 100 cases.
- 3 Find the marginal totals. Rater 1 said yes 30 times and no 70; rater 2 said yes 35 times and no 65.
- 4 Compute chance agreement. (30 × 35 + 70 × 65) ÷ 100² = (1050 + 4550) ÷ 10,000 = .56.
- 5 Apply the formula. (.85 − .56) ÷ (1 − .56) = .29 ÷ .44 = .66.
- 6 Interpret it. On the Landis & Koch bands, .66 is substantial agreement.
Landis & Koch (1977) interpretation bands
The most widely cited benchmarks. They are conventions rather than statistical thresholds, and stricter cut-offs are common in clinical work.
| Kappa | Strength of agreement |
|---|---|
| Below .00 | Poor — worse than chance |
| .00 – .20 | Slight |
| .21 – .40 | Fair |
| .41 – .60 | Moderate |
| .61 – .80 | Substantial |
| .81 – 1.00 | Almost perfect |
The kappa paradox
Kappa has a well-documented quirk: with very skewed marginals, high observed agreement can produce a surprisingly low kappa. If both raters call 95% of cases negative and agree on 90% of the sample, chance agreement is already so high that little room remains to beat it, and kappa collapses. The measure is not broken — it is telling you that agreement on a near-constant variable carries little information.
Because of this, report observed agreement alongside kappa and describe the prevalence of each category. A reader can then see whether a low kappa reflects genuine disagreement or a base-rate problem.
Two limits worth knowing. Kappa treats all disagreements as equally serious, so for ordered categories — mild, moderate, severe — weighted kappa is the better choice, since it penalises a two-step disagreement more than a one-step one. And kappa is defined for exactly two raters; with three or more, Fleiss’ kappa generalises it.