Item Analysis
How to read item difficulty, discrimination, KR-20, point-biserial, and distractor statistics to improve exam quality.
Item analysis is the statistical review of how individual exam questions performed after an administration. For every item it answers three practical questions: how hard was it (difficulty index), did it separate stronger examinees from weaker ones (discrimination, usually a point-biserial correlation), and did the wrong answers do any work (distractor analysis)? At the whole-exam level, reliability statistics such as KR-20 and Cronbach's alpha summarize how consistently the test measured.
Assessment teams should care because item flaws are common and consequential. Tarrant, Knierim, Hayes and Ware (2006) found that nearly half of the high-stakes nursing MCQs they reviewed contained at least one item-writing flaw, and Tarrant and Ware (2008) showed that flawed items disproportionately penalize high-achieving students. Item analysis is the cheapest quality gate available: it flags miskeyed answers, ambiguous stems, and implausible distractors before scores are released, and over successive administrations it turns a question bank into a progressively better-calibrated measurement instrument. The guides in this topic explain each statistic, the thresholds commonly used in medical education, and what to do when an item fails them.
Questions & answers
What is the difference between KR-20 and Cronbach's alpha?
KR-20 (Kuder & Richardson, 1937) estimates internal-consistency reliability for tests scored dichotomously — right or wrong. Cronbach's alpha (1951) generalizes the same formula to items with any scoring scale, including partial credit. On an all-dichotomous exam the two produce identical values; alpha is simply the more general statistic.
Read the full answer →What does a negative point-biserial correlation mean?
A negative point-biserial means examinees who scored well overall answered that item incorrectly more often than weaker examinees. It almost always signals a problem — a miskeyed answer, an ambiguous or flawed stem, or a plausible distractor that misleads knowledgeable students — and the item should be reviewed before scores are released.
Read the full answer →What is a good KR-20 value for an exam?
For high-stakes exams, KR-20 values of 0.80 or higher are generally considered good, and 0.90 or higher is expected for licensure-level decisions. Values of 0.70-0.79 are acceptable for lower-stakes classroom tests. KR-20 depends heavily on test length, so short quizzes naturally produce lower values.
Read the full answer →What is distractor analysis, and what is a non-functioning distractor?
Distractor analysis examines how many examinees selected each incorrect option on a multiple-choice item. A non-functioning distractor is one chosen by very few examinees — commonly fewer than 5% — or one that attracts more high scorers than low scorers. Non-functioning distractors add length without adding measurement value.
Read the full answer →What is item analysis in exam development?
Item analysis is the statistical review of how each exam question performed after administration. It examines difficulty (the proportion answering correctly), discrimination (how well the item separates strong from weak examinees), and distractor performance, so flawed items can be revised or removed before scores are finalized.
Read the full answer →What is a point-biserial correlation, and what value is acceptable?
The point-biserial correlation measures how strongly performance on one item correlates with total exam score. Positive values mean stronger examinees answered correctly more often. Values above roughly 0.20 are generally acceptable, 0.30 and above is good, and items near zero or negative deserve review before scoring is finalized.
Read the full answer →