Skip to content
K Knidox Search…
Statistics · Agreement

Cohen’s Kappa

Chance-corrected agreement between two raters, from a 2×2 table.

Rater 1 yes, rater 2 yes.
Rater 1 no, rater 2 no.
Cohen’s kappa
0.659κSubstantial

Observed agreement 85%, chance agreement 56% across 100 cases.

Observed agreement
85%

The raw percentage the two raters agreed on — before correcting for chance.

Kappa corrects raw agreement for the agreement expected by chance: κ = (p_o − p_e) ÷ (1 − p_e). With 25 both-yes, 5 and 10 disagreements, and 60 both-no, observed agreement is 85% and chance agreement 56%, giving κ = .29 ÷ .44 = .66 — substantial.

Why raw agreement is not enough

Two coders who agree on 85% of cases sound reliable. But if a category is rare — say only 5% of cases are positive — two coders who both guessed “negative” every time would agree on 95% of cases while sharing no judgement at all. Raw agreement rewards skewed categories, which is exactly where reliability matters most.

Kappa removes that free agreement. It calculates how often the two raters would coincide by chance alone, given how often each of them uses each category, and then asks how much of the remaining possible agreement was actually achieved. A kappa of .66 means the raters captured 66% of the agreement available beyond chance.

Reading the table

The four cells are the two agreements — both yes, both no — and the two disagreements. Only the diagonal counts towards observed agreement, but all four shape the chance term, because the marginal totals describe how freely each rater used each category.

κ = (p_o − p_e) ÷ (1 − p_e) p_o = (a + d) ÷ n p_e = (row₁ × col₁ + row₂ × col₂) ÷ n²

p_o is observed agreement, p_e the agreement expected by chance from the marginal totals. Kappa runs from −1 to 1, where 0 is exactly chance-level agreement.

Worked example: 25 / 5 / 10 / 60

Observed first, chance second, then the ratio:

  1. 1
    Build the 2×2 table. Both yes = 25, rater 1 yes only = 5, rater 2 yes only = 10, both no = 60. Total n = 100.
  2. 2
    Compute observed agreement. (25 + 60) ÷ 100 = .85 — the raters agreed on 85 of 100 cases.
  3. 3
    Find the marginal totals. Rater 1 said yes 30 times and no 70; rater 2 said yes 35 times and no 65.
  4. 4
    Compute chance agreement. (30 × 35 + 70 × 65) ÷ 100² = (1050 + 4550) ÷ 10,000 = .56.
  5. 5
    Apply the formula. (.85 − .56) ÷ (1 − .56) = .29 ÷ .44 = .66.
  6. 6
    Interpret it. On the Landis & Koch bands, .66 is substantial agreement.

Landis & Koch (1977) interpretation bands

The most widely cited benchmarks. They are conventions rather than statistical thresholds, and stricter cut-offs are common in clinical work.

KappaStrength of agreement
Below .00Poor — worse than chance
.00 – .20Slight
.21 – .40Fair
.41 – .60Moderate
.61 – .80Substantial
.81 – 1.00Almost perfect

The kappa paradox

Kappa has a well-documented quirk: with very skewed marginals, high observed agreement can produce a surprisingly low kappa. If both raters call 95% of cases negative and agree on 90% of the sample, chance agreement is already so high that little room remains to beat it, and kappa collapses. The measure is not broken — it is telling you that agreement on a near-constant variable carries little information.

Because of this, report observed agreement alongside kappa and describe the prevalence of each category. A reader can then see whether a low kappa reflects genuine disagreement or a base-rate problem.

Two limits worth knowing. Kappa treats all disagreements as equally serious, so for ordered categories — mild, moderate, severe — weighted kappa is the better choice, since it penalises a two-step disagreement more than a one-step one. And kappa is defined for exactly two raters; with three or more, Fleiss’ kappa generalises it.

What is a good kappa value?
On the Landis & Koch bands, .61 to .80 is substantial and above .81 almost perfect, with .41 to .60 moderate. Clinical and diagnostic work often demands stricter values, so the field convention matters more than the generic bands.
Why is my kappa low when agreement looks high?
This is the kappa paradox. When one category dominates, chance agreement is already very high, leaving little room to exceed it. Report observed agreement and category prevalence alongside kappa so the reason is visible.
Can kappa be negative?
Yes. A negative kappa means the raters agreed less often than chance alone would predict, which usually points to a coding scheme being applied in opposite directions rather than to random error.
What is weighted kappa?
A version for ordered categories that penalises large disagreements more than small ones. If ratings run mild, moderate, severe, a mild-versus-severe disagreement should count more heavily than mild-versus-moderate, which plain kappa cannot express.
Can I use kappa for three or more raters?
No — Cohen’s kappa is defined for exactly two. Fleiss’ kappa generalises the idea to any number of raters, and the intraclass correlation is used when ratings are continuous rather than categorical.
Does kappa work with more than two categories?
Yes. The same formula extends to any square table: observed agreement is the sum of the diagonal, and chance agreement sums the products of matching row and column totals. This calculator handles the common 2×2 case.
How do I report kappa in a paper?
Give kappa to two decimals without a leading zero, the number of cases, and the observed agreement — for example “κ = .66, 85% agreement, n = 100”. Naming the interpretation scheme you used is good practice.