Skip to content

Item Analysis

How to read item difficulty, discrimination, KR-20, point-biserial, and distractor statistics to improve exam quality.

Item analysis is the statistical review of how individual exam questions performed after an administration. For every item it answers three practical questions: how hard was it (difficulty index), did it separate stronger examinees from weaker ones (discrimination, usually a point-biserial correlation), and did the wrong answers do any work (distractor analysis)? At the whole-exam level, reliability statistics such as KR-20 and Cronbach's alpha summarize how consistently the test measured.

Assessment teams should care because item flaws are common and consequential. Tarrant, Knierim, Hayes and Ware (2006) found that nearly half of the high-stakes nursing MCQs they reviewed contained at least one item-writing flaw, and Tarrant and Ware (2008) showed that flawed items disproportionately penalize high-achieving students. Item analysis is the cheapest quality gate available: it flags miskeyed answers, ambiguous stems, and implausible distractors before scores are released, and over successive administrations it turns a question bank into a progressively better-calibrated measurement instrument. The guides in this topic explain each statistic, the thresholds commonly used in medical education, and what to do when an item fails them.

No questions published under this topic yet.