Skip to content
K Knidox Search…
Statistics · Assessment

Test Item Analysis

Difficulty, discrimination, point-biserial, and KR-20 from a grid of right and wrong answers.

One row per student, one column per question. 1 for correct, 0 for incorrect. Rows must all be the same length.
Test reliability (KR-20)
0.497Low

10 students, 5 questions. Mean score 2.20 of 5. KR-20 is Cronbach’s alpha for right/wrong items.

Per-item statistics
ItemCorrectDifficulty (p)Discrimination (D)r₍pb₎
18/100.80 · Easy0.67 · Excellent0.681
26/100.60 · Moderate0.67 · Excellent0.621
34/100.40 · Moderate1.00 · Excellent0.686
43/100.30 · Moderate0.33 · Good0.419
51/100.10 · Hard0.33 · Good0.480

Discrimination uses the conventional upper and lower 27% of students by total score. With a small class those groups are tiny, so treat D as indicative rather than precise — the point-biserial, which uses every student, is more stable at low sample sizes.

Item analysis asks which questions actually worked. Difficulty is the proportion answering correctly; discrimination compares the strongest and weakest students. A question everyone answers correctly has a difficulty of 1.00 and a discrimination of 0 — it separates nobody.

Two numbers per question

After marking a test, the scores tell you how students did. Item analysis tells you how the test did — whether each question earned its place on the paper.

The difficulty index (p) is simply the proportion of students who answered correctly. The name is counterintuitive: a high p means an easy item, because more people got it right. A p of 0.85 means 85% succeeded.

The discrimination index (D) asks whether the question separated strong students from weak ones. Rank the class by total score, take the top 27% and the bottom 27%, and subtract the proportion correct in the lower group from the proportion in the upper. A question the best students answer and the weakest miss has high discrimination; one where both groups do the same has none.

The case that should worry you

A negative discrimination is the finding that matters most. It means the weakest students outperformed the strongest on that question, which almost never happens by chance in a well-written item. The usual causes are a miskeyed answer, an ambiguous distractor that the better students overthink, or a question testing something the strong students learned differently. A negative D is a reason to re-read the item before it counts toward anyone's grade.

p = correct ÷ students D = p(upper 27%) − p(lower 27%) r₍pb₎ = [(M₁ − M₀) ÷ SD] × √(p × q)

M₁ and M₀ are the mean total scores of students who got the item right and wrong. SD is the population standard deviation of total scores, and q = 1 − p.

How to read the output

Difficulty and discrimination interact — you need both to judge an item:

  1. 1
    Build the response matrix. One row per student, one column per question, 1 for correct and 0 for incorrect. Every row must be the same length.
  2. 2
    Read the difficulty first. Around 0.5 is ideal for maximising discrimination, though a spread from about 0.3 to 0.9 is healthy across a paper.
  3. 3
    Then the discrimination. Above 0.3 is good. Below 0.2 means the item is not separating students, whatever its difficulty.
  4. 4
    Investigate anything negative. Negative discrimination usually means a miskeyed answer or a misleading option. Check the key before anything else.
  5. 5
    Check the point-biserial. It uses every student rather than just the extremes, so it is more stable in a small class. It should broadly agree with D.
  6. 6
    Look at KR-20 for the whole test. Above 0.8 is good for a classroom test; below 0.7 suggests the items are not measuring one consistent thing.

Interpreting the indices

Conventional bands from classical test theory. They are guidance for reviewing items, not rules for discarding them automatically.

ValueDifficulty (p)Discrimination (D)
0.90 and aboveVery easy — may be a warm-up itemExcellent
0.70 – 0.89EasyExcellent
0.40 – 0.69Moderate — best for discriminatingExcellent
0.30 – 0.39Moderate to hardGood
0.20 – 0.29HardAcceptable
Below 0.20Very hard — check it is answerablePoor — revise the item
NegativeNot possibleReview immediately — likely miskeyed

Difficulty and discrimination are linked

An item everyone answers correctly cannot discriminate, because there is no variation to correlate with anything — its D is necessarily 0. The same is true of an item nobody answers correctly. Discrimination is mathematically greatest when difficulty is near 0.5, which is why assessment texts recommend aiming there on average.

That does not mean every question should sit at 0.5. A test needs some easy items so weaker students can demonstrate what they know, and some hard ones to separate the top of the class. What it means is that a paper made entirely of very easy or very hard items will fail to rank students reliably regardless of how many questions it has.

Two cautions on sample size. The 27% split is a convention chosen to balance group size against how extreme the groups are, but in a class of twelve it means groups of three, and D computed from three students is noisy. The point-biserial correlation uses the whole class and is the more trustworthy figure in small groups. And KR-20 is Cronbach's alpha for right/wrong scoring — the same measure, specialised to dichotomous items.

Why is a high difficulty index an easy question?
Because p is defined as the proportion answering correctly, not the proportion failing. A p of 0.90 means 90% got it right. The name is a long-standing convention rather than a description.
What does negative discrimination mean?
That weaker students outperformed stronger ones on that item, which rarely happens by chance. The usual causes are a miskeyed answer or an ambiguous option that better students read too carefully.
Why 27% for the upper and lower groups?
It is the conventional split, chosen to balance having enough students in each group against making the groups genuinely extreme. In small classes the resulting groups are tiny, so treat D cautiously.
Should every item have a difficulty around 0.5?
On average, since discrimination peaks there. But a good paper spreads difficulty so weaker students can show what they know and the strongest are still separated at the top.
What is the difference between D and the point-biserial?
D compares only the extreme groups; the point-biserial correlates the item against total score using every student. The point-biserial is more stable in small classes and should broadly agree with D.
What is KR-20?
A reliability coefficient for the whole test — Cronbach’s alpha specialised to items scored right or wrong. Above 0.8 is good for a classroom test; below 0.7 suggests the items are not measuring one thing consistently.
Should I drop items with poor statistics?
Review them rather than deleting automatically. A low-discrimination item may still cover content you need to assess, and the fix is often rewording rather than removal.