Skip to content
K Knidox Search…
Statistics · Reliability

Split-Half Reliability

Correct a half-test correlation back up to full length with the Spearman-Brown formula.

Pearson correlation between the odd- and even-numbered halves of the test.
How many times longer the new test is. 2 rebuilds the full test from one half.
Split-half reliability (Spearman-Brown)
0.788rAcceptable

Correcting the half-test correlation of 0.65 back up to full test length.

Split the test in half, correlate the two halves, then correct for length: r_sb = 2r ÷ (1 + r). A correlation of .65 between halves gives 2 × .65 ÷ 1.65 = .79 for the full-length test. The correction is needed because each half is only half as long.

Why the raw correlation understates reliability

Split-half reliability estimates consistency by treating one test as two — typically odd-numbered items against even-numbered ones — and correlating the two scores. But that correlation describes the reliability of a half-length test, and shorter tests are always less reliable than longer ones because a single unlucky item carries more weight.

The Spearman-Brown formula corrects for this, projecting what the correlation would be if each half were as long as the whole. That is why a raw half-correlation of .65 becomes a corrected reliability of .79: the same items, measured at their true length.

How to split

The odd-even split is standard because it distributes item difficulty and any fatigue effect evenly across both halves. Splitting first-half against second-half is a poor choice on a timed test, where later items may be rushed or unreached — the two halves then differ systematically, and the correlation reflects that rather than reliability.

Any given split is one of many possible splits, and different splits give different answers. This is the known weakness of the method, and it is precisely what Cronbach’s alpha resolves — alpha is mathematically the average of all possible split-half coefficients.

Split-half: r_sb = 2r ÷ (1 + r) General prophecy: r* = k·r ÷ (1 + (k − 1)·r)

r is the observed correlation between halves. In the general form, k is how many times longer the new test is — k = 2 recovers the split-half formula.

Worked example: halves correlating at .65

Correlate, then correct:

  1. 1
    Split the test into two halves. Odd-numbered items in one half, even-numbered in the other, so difficulty is balanced across both.
  2. 2
    Score each half separately. Every respondent now has two scores, one per half.
  3. 3
    Correlate the two halves. The Pearson correlation across respondents comes out at r = .65.
  4. 4
    Apply Spearman-Brown. 2 × .65 ÷ (1 + .65) = 1.30 ÷ 1.65 = .79.
  5. 5
    Report the corrected value. The full-length test has an estimated reliability of .79 — the uncorrected .65 would understate it.

Corrected reliability at different half-correlations

Spearman-Brown correction, r_sb = 2r ÷ (1 + r). The correction matters most in the middle of the range.

Correlation between halvesCorrected reliabilityGain from correction
.40.57+.17
.50.67+.17
.60.75+.15
.65.79+.14
.70.82+.12
.80.89+.09
.90.95+.05

Using the prophecy formula to plan a test

The general form answers a design question rather than a measurement one: how long would this test need to be to reach a target reliability? A 20-item test with a reliability of .70 doubled to 40 comparable items would reach 2 × .70 ÷ (1 + .70) = .82. Tripled, .88.

Two cautions on that projection. It assumes the added items are of comparable quality to the existing ones, which is optimistic — the best items usually go in first. And it says nothing about respondent fatigue, which can degrade data quality on a long instrument faster than the extra items improve it.

Why correct the correlation at all?
Because the correlation describes two half-length tests, and shorter tests are inherently less reliable. Spearman-Brown projects the value back to the full length, which is what you actually administer.
How should I split the test?
Odd items against even items is standard, because it balances difficulty and any fatigue effect. A first-half against second-half split is a poor choice on timed tests, where later items are often rushed.
Different splits give different answers — is that a problem?
It is the known weakness of the method. Cronbach’s alpha solves it directly, since alpha is mathematically the mean of every possible split-half coefficient for that scale.
What is the Spearman-Brown prophecy formula?
The general version, r* = k·r ÷ (1 + (k − 1)·r), which predicts reliability at any change in test length. Setting k = 2 gives the ordinary split-half correction.
How many items would I need for a reliability of .90?
It depends on where you start. A 20-item test at .70 would need roughly four times the length — about 80 comparable items — to approach .90, which is often impractical.
Does the prophecy formula account for fatigue?
No. It assumes added items are as good as existing ones and that respondents answer them as carefully. In practice both assumptions weaken as a questionnaire grows.
Should I use split-half or alpha?
Alpha in almost all cases, since it avoids the arbitrary choice of split. Split-half remains useful when you can only compute one correlation, or when a specific split is theoretically meaningful.