Test Item Analysis
Difficulty, discrimination, point-biserial, and KR-20 from a grid of right and wrong answers.
10 students, 5 questions. Mean score 2.20 of 5. KR-20 is Cronbach’s alpha for right/wrong items.
| Item | Correct | Difficulty (p) | Discrimination (D) | r₍pb₎ |
|---|---|---|---|---|
| 1 | 8/10 | 0.80 · Easy | 0.67 · Excellent | 0.681 |
| 2 | 6/10 | 0.60 · Moderate | 0.67 · Excellent | 0.621 |
| 3 | 4/10 | 0.40 · Moderate | 1.00 · Excellent | 0.686 |
| 4 | 3/10 | 0.30 · Moderate | 0.33 · Good | 0.419 |
| 5 | 1/10 | 0.10 · Hard | 0.33 · Good | 0.480 |
Discrimination uses the conventional upper and lower 27% of students by total score. With a small class those groups are tiny, so treat D as indicative rather than precise — the point-biserial, which uses every student, is more stable at low sample sizes.
Item analysis asks which questions actually worked. Difficulty is the proportion answering correctly; discrimination compares the strongest and weakest students. A question everyone answers correctly has a difficulty of 1.00 and a discrimination of 0 — it separates nobody.
Two numbers per question
After marking a test, the scores tell you how students did. Item analysis tells you how the test did — whether each question earned its place on the paper.
The difficulty index (p) is simply the proportion of students who answered correctly. The name is counterintuitive: a high p means an easy item, because more people got it right. A p of 0.85 means 85% succeeded.
The discrimination index (D) asks whether the question separated strong students from weak ones. Rank the class by total score, take the top 27% and the bottom 27%, and subtract the proportion correct in the lower group from the proportion in the upper. A question the best students answer and the weakest miss has high discrimination; one where both groups do the same has none.
The case that should worry you
A negative discrimination is the finding that matters most. It means the weakest students outperformed the strongest on that question, which almost never happens by chance in a well-written item. The usual causes are a miskeyed answer, an ambiguous distractor that the better students overthink, or a question testing something the strong students learned differently. A negative D is a reason to re-read the item before it counts toward anyone's grade.
M₁ and M₀ are the mean total scores of students who got the item right and wrong. SD is the population standard deviation of total scores, and q = 1 − p.
How to read the output
Difficulty and discrimination interact — you need both to judge an item:
- 1 Build the response matrix. One row per student, one column per question, 1 for correct and 0 for incorrect. Every row must be the same length.
- 2 Read the difficulty first. Around 0.5 is ideal for maximising discrimination, though a spread from about 0.3 to 0.9 is healthy across a paper.
- 3 Then the discrimination. Above 0.3 is good. Below 0.2 means the item is not separating students, whatever its difficulty.
- 4 Investigate anything negative. Negative discrimination usually means a miskeyed answer or a misleading option. Check the key before anything else.
- 5 Check the point-biserial. It uses every student rather than just the extremes, so it is more stable in a small class. It should broadly agree with D.
- 6 Look at KR-20 for the whole test. Above 0.8 is good for a classroom test; below 0.7 suggests the items are not measuring one consistent thing.
Interpreting the indices
Conventional bands from classical test theory. They are guidance for reviewing items, not rules for discarding them automatically.
| Value | Difficulty (p) | Discrimination (D) |
|---|---|---|
| 0.90 and above | Very easy — may be a warm-up item | Excellent |
| 0.70 – 0.89 | Easy | Excellent |
| 0.40 – 0.69 | Moderate — best for discriminating | Excellent |
| 0.30 – 0.39 | Moderate to hard | Good |
| 0.20 – 0.29 | Hard | Acceptable |
| Below 0.20 | Very hard — check it is answerable | Poor — revise the item |
| Negative | Not possible | Review immediately — likely miskeyed |
Difficulty and discrimination are linked
An item everyone answers correctly cannot discriminate, because there is no variation to correlate with anything — its D is necessarily 0. The same is true of an item nobody answers correctly. Discrimination is mathematically greatest when difficulty is near 0.5, which is why assessment texts recommend aiming there on average.
That does not mean every question should sit at 0.5. A test needs some easy items so weaker students can demonstrate what they know, and some hard ones to separate the top of the class. What it means is that a paper made entirely of very easy or very hard items will fail to rank students reliably regardless of how many questions it has.
Two cautions on sample size. The 27% split is a convention chosen to balance group size against how extreme the groups are, but in a class of twelve it means groups of three, and D computed from three students is noisy. The point-biserial correlation uses the whole class and is the more trustworthy figure in small groups. And KR-20 is Cronbach's alpha for right/wrong scoring — the same measure, specialised to dichotomous items.