What is the difference between KR-20 and Cronbach's alpha?
KR-20 (Kuder & Richardson, 1937) estimates internal-consistency reliability for tests scored dichotomously — right or wrong. Cronbach's alpha (1951) generalizes the same formula to items with any scoring scale, including partial credit. On an all-dichotomous exam the two produce identical values; alpha is simply the more general statistic.
KR-20 and Cronbach's alpha are not competing statistics — one is a special case of the other. Both estimate internal-consistency reliability: the degree to which the items on a test hang together and measure the same underlying construct.
The historical relationship
Kuder and Richardson (1937) derived Formula 20 for tests in which every item is scored dichotomously — correct or incorrect, 1 or 0. Cronbach (1951) later generalized the derivation to items scored on any scale, and named the result coefficient alpha. Mathematically, when every item is dichotomous, alpha reduces exactly to KR-20. They are the same number in that case; the labels differ by tradition, not substance.
When each applies
Situation | Appropriate statistic |
|---|---|
MCQ exam, each item scored 1/0 | KR-20 (equivalently, alpha) |
Items with partial credit (e.g., multi-part questions) | Cronbach's alpha |
Rating scales, OSCE checklists, Likert-style instruments | Cronbach's alpha |
Mixed exam (some dichotomous, some polytomous items) | Cronbach's alpha |
This is why software output varies: some systems label the reliability of an MCQ exam "KR-20," others "Cronbach's alpha," and for a fully dichotomous exam both labels are correct. If your platform reports only alpha, you are not missing anything.
A related statistic you may encounter is KR-21, a simplified approximation from the same 1937 paper that assumes all items are equally difficult. Because that assumption rarely holds, KR-21 typically underestimates reliability and is mainly of historical interest; prefer KR-20/alpha when item-level data are available.
What both statistics share — including their limits
Because they are the same family, KR-20 and alpha share interpretation and caveats:
- Both range in practice from 0 to 1, with the usual benchmarks applying equally (see what counts as a good KR-20 value).
- Both increase with test length and with the average discrimination of the items (see point-biserial correlation).
- Both are depressed by restricted score variance — a very homogeneous cohort lowers the estimate even when items are sound.
- Neither is a measure of validity or unidimensionality. A high alpha does not prove the exam measures one coherent construct, and it says nothing about whether the content sampled matches the course blueprint.
Practical takeaway
For a standard single-best-answer MCQ exam, do not spend committee time on the KR-20-versus-alpha question — the values coincide. Spend it instead on what moves the number: identifying and repairing the low- and negative-discrimination items that the same item-analysis report will flag alongside the reliability coefficient.