What is a good KR-20 value for an exam?
For high-stakes exams, KR-20 values of 0.80 or higher are generally considered good, and 0.90 or higher is expected for licensure-level decisions. Values of 0.70-0.79 are acceptable for lower-stakes classroom tests. KR-20 depends heavily on test length, so short quizzes naturally produce lower values.
KR-20 (Kuder-Richardson Formula 20) estimates the internal-consistency reliability of an exam scored right/wrong: the extent to which its items measure the same underlying ability rather than random noise. It ranges from 0 to 1, and the acceptable value depends on the stakes of the decision the exam supports.
Commonly used benchmarks
KR-20 | Interpretation | Typical use |
|---|---|---|
≥ 0.90 | Excellent | Licensure, certification, high-stakes pass/fail decisions |
0.80 - 0.89 | Good | Summative course exams, progression decisions |
0.70 - 0.79 | Acceptable | Lower-stakes summative and formative testing |
< 0.70 | Weak | Interpret individual scores with caution |
These bands trace back to widely cited measurement texts — Nunnally's guidance that around 0.70 is a minimum for early-stage measurement while higher-stakes decisions demand more, and standard educational-measurement references such as Ebel and Frisbie (1991). They are conventions, not laws: what matters is whether the reliability is adequate for the decision being made about each examinee.
What drives KR-20 up or down
Three factors dominate:
- Test length. Reliability rises with the number of items, all else equal (the Spearman-Brown relationship). A 15-item quiz with KR-20 of 0.60 is not necessarily written worse than an 80-item exam at 0.88 — it is mostly shorter. Judge short tests accordingly.
- Item discrimination. Items with strong point-biserial correlations pull reliability up; items with near-zero or negative discrimination drag it down. Removing or fixing a handful of poorly discriminating items is usually the fastest way to raise KR-20.
- Score spread in the cohort. KR-20 is computed from score variance, so a homogeneous group (for example, a small cohort of uniformly strong students) yields lower values even when the items are sound. A low KR-20 in a mastery-oriented course where nearly everyone scores high is expected, not alarming.
Practical guidance
Do not treat a single low KR-20 as a verdict on the exam. First check the item-level statistics: one or two miskeyed items with negative discrimination can noticeably depress reliability on a medium-length test. Second, consider cohort size — reliability estimates from very small groups are themselves unstable. Third, look at the trend across administrations rather than any single value.
Also remember what KR-20 does not tell you. It says nothing about validity: an exam can reliably measure the wrong thing. Blueprint coverage (see what an exam blueprint is) and content review remain separate, necessary checks.
Most modern delivery systems report KR-20 automatically with each administration, so the real work is interpretive: matching the observed value to the stakes of the exam and acting on the item-level diagnostics behind it. For how KR-20 fits into the full post-exam review workflow, see the complete guide to item analysis.