What is item analysis in exam development?
Item analysis is the statistical review of how each exam question performed after administration. It examines difficulty (the proportion answering correctly), discrimination (how well the item separates strong from weak examinees), and distractor performance, so flawed items can be revised or removed before scores are finalized.
Item analysis is the post-administration quality check of an exam. After students take a test, each question is examined statistically to answer a simple question: did this item measure what we intended, or did it add noise? In medical education, where multiple-choice exams feed progression and remediation decisions, item analysis is the standard mechanism for catching problems before scores are released.
The core statistics
Classical item analysis reports three things for every question:
Statistic | What it tells you | Typical concern threshold |
|---|---|---|
Difficulty index (p-value) | Proportion of examinees answering correctly | Below ~0.30 (very hard) or above ~0.90 (very easy) |
Discrimination (point-biserial) | Correlation between getting the item right and total score | Below ~0.20, and especially negative values |
Distractor performance | How many examinees chose each wrong option | Options selected by almost no one |
At the whole-exam level, a reliability coefficient — usually KR-20 or Cronbach's alpha — summarizes how consistently the full set of items measured.
Why it matters
Item-writing flaws are not rare. Tarrant, Knierim, Hayes and Ware (2006) reviewed 2,770 high-stakes nursing MCQs and found that 46.2% contained at least one item-writing flaw, with ambiguous stems, negatively worded stems, and implausible distractors among the most frequent. Tarrant and Ware (2008) further showed that flawed items tend to disadvantage high-achieving students — the opposite of what a fair exam should do. Item analysis is how those flaws surface in practice: a miskeyed answer shows up as a negative point-biserial, and an implausible distractor shows up as an option nobody selects.
What assessment teams do with it
A typical post-exam workflow looks like this:
- Run the statistics immediately after the administration (assessment platforms such as StudyDrome compute item difficulty, discrimination, distractor counts, and reliability automatically after each exam).
- Flag items with negative or near-zero discrimination and verify the answer key first — miskeys are the most common cause.
- Review flagged items with content experts and decide whether to score, revise, or drop each one before releasing results.
- Record the statistics in the question bank so that item history accumulates across administrations and future exams draw on questions with a proven track record.
Item analysis does not replace good item writing — following an established guideline set such as Haladyna, Downing and Rodriguez (2002) prevents many flaws from being written at all — but it is the empirical backstop that catches what review committees miss. For a full walkthrough of each statistic with worked examples, see the complete guide to item analysis.