Skip to content

The Complete Guide to Item Analysis in Medical Education

How to read difficulty, discrimination, reliability, and distractor statistics — and turn every exam into better questions.

item analysis

Every written exam you deliver produces two things: scores for your students, and data about your questions. Most programs use the first and ignore the second. That is a waste because the second is how exams improve.

Item analysis is the systematic review of how each question performed after an exam: how hard it was, how well it separated stronger from weaker candidates, and whether its wrong answers did any work. Downing (2005) has shown that flawed items are not a cosmetic problem: they can distort scores and pass/fail decisions in ways that disadvantage capable students. Item analysis is your primary instrument for finding those flaws after the fact, and validity frameworks in the tradition of Kane treat this kind of internal-structure evidence as a core part of the argument that your exam scores mean what you claim they mean.

This guide walks through each statistic, what its values mean in practice, and a concrete review workflow you can run after every exam. If you want the short version first, start with what item analysis is and come back.

Why item analysis matters

Three reasons, in increasing order of importance:

  1. Quality control. Miskeyed answers, ambiguous stems, and two-defensible-options items are invisible during writing and glaring in the data. Item analysis catches them before scores are finalized.
  2. Fairness. Downing (2004, 2005) argued that item flaws function as construct-irrelevant variance: they measure test-taking savvy or tolerance for ambiguity rather than medical knowledge, and they can penalize precisely the students who know the material.
  3. Bank improvement. Over successive administrations, item statistics let you retire weak items, refine mediocre ones, and build a question bank whose properties you actually know. This is the difference between a bank and a pile.

Item difficulty: the P-value

The difficulty index (usually written P) is simply the proportion of examinees who answered the item correctly. An item with P = 0.85 was answered correctly by 85% of the cohort. Note the naming quirk: a higher P-value means an easier item.

There is no single "correct" difficulty, because the right target depends on the exam's purpose:

P-value

Interpretation

Typical action

0.90 – 1.00

Very easy

Fine for criterion-referenced mastery content; contributes little to discrimination

0.70 – 0.90

Easy to moderate

Usually acceptable

0.30 – 0.70

Moderate to hard

Best range for spreading out scores; check discrimination

< 0.30

Very hard

Review carefully: hard, flawed, miskeyed, or untaught?

Ebel & Frisbie's classic treatment notes that items of moderate difficulty carry the most measurement information for ranking examinees. But medical school exams are mostly criterion-referenced (you care whether students meet a standard, not how they rank), so a clutch of very easy items covering must-know material is not a defect. What deserves scrutiny is the extremes: an item nearly everyone gets right tells you little about individuals, and an item nearly everyone gets wrong is more often a flawed item or a curriculum gap than a legitimately fiendish question.

One caution: P-values are sample-dependent. The same item is "easier" for a strong cohort. Never compare raw P-values across very different groups without acknowledging that.

Discrimination: the point-biserial correlation

Difficulty tells you how many got the item right; discrimination tells you who. The point-biserial correlation (r-pbis) relates performance on one item (right/wrong) with total score on the rest of the exam. A well-functioning item is answered correctly more often by students who did well overall.

Working interpretation, consistent with guidance in standard measurement texts (Ebel & Frisbie) and the NBME's item-writing tradition (Case & Swanson):

  • ≥ 0.30: good discriminator; keep.
  • 0.20 – 0.29: acceptable; usually fine as-is.
  • 0.10 – 0.19: marginal; review the item, especially its distractors.
  • ≈ 0 or negative: problem item; investigate before counting it.

A negative point-biserial means stronger students got the item wrong more often than weaker students. In practice this almost always signals one of three things: a miskeyed answer (the most common and most fixable), a genuinely ambiguous item where the sophisticated reading leads away from the keyed answer, or content taught inconsistently across groups. A negative discrimination on an item is a fire alarm, not a curiosity. Pull the item, read it fresh, and check the key against the source material before releasing scores.

Two structural caveats. First, very easy and very hard items are arithmetically constrained to low discrimination (an item everyone answers correctly cannot discriminate), so do not condemn a P = 0.95 must-know item for its r-pbis of 0.08. Second, in small cohorts the point-biserial is noisy; more on that below.

Reliability: KR-20, alpha, and the SEM

Item statistics describe questions; reliability describes the exam. KR-20 (for dichotomously scored items) and Cronbach's alpha (its generalization to polytomous scoring) estimate internal consistency: roughly, how reproducible the rank-ordering of examinees would be with a parallel set of items. The two coincide for right/wrong MCQ exams; the practical differences are covered in KR-20 vs Cronbach's alpha.

What counts as a good KR-20 value depends on stakes. Downing (2004) frames reliability as a necessary condition in the validity argument, and the widely used rule of thumb in health-professions assessment is that high-stakes decisions warrant coefficients around 0.80 or higher, while lower-stakes in-course tests can tolerate less. Reliability rises with test length and with average item discrimination, so the two levers you control are: write more good items, and fix or drop the non-discriminating ones.

Report the standard error of measurement (SEM) alongside reliability. The SEM translates the reliability coefficient into score points: an observed score is best interpreted as a band of roughly ±1 SEM around the true score, not a point. This matters most at the cut score: a student one SEM below the pass mark is not confidently distinguishable from one at the mark, which is why many programs pair item analysis with explicit standard-setting and borderline review policies.

Distractor analysis

Every option in an MCQ is supposed to earn its place. Distractor analysis examines the wrong options: how many examinees chose each, and whether choosers were weaker overall on the exam.

A functioning distractor is typically defined as one selected by at least 5% of examinees and chosen more often by low scorers than by high scorers. Tarrant and Ware's work on MCQ quality in nursing and health-professions exams found that non-functioning distractors are extremely common in teacher-written items. In practice, many four- and five-option items have only one or two options doing any work. This dovetails with Rodriguez's (2005) synthesis of decades of research, which concludes that three options (one key, two distractors) are generally sufficient because writers rarely produce more than two plausible wrong answers anyway.

What to look for in the option table:

  • Distractors chosen by almost no one: dead weight. Replace or delete them (and consider whether the item needs fewer options).
  • A distractor attracting high scorers: the tell-tale sign of an ambiguous or arguably-correct option. This is where negative point-biserials come from.
  • One distractor that absorbs nearly all errors: often fine, but check whether it reflects a specific teachable misconception (useful feedback for the course) or a flaw, such as implausible remaining options.

Distractor analysis is also your best feedback loop to item writers: it shows, option by option, which of their crafted misconceptions actually landed. Many of the patterns it exposes trace back to the most common item-writing flaws cataloged by Haladyna, Downing & Rodriguez (2002).

A post-exam review workflow

Here is a workflow that a course team can run in under an hour for a typical exam, ideally before scores are released. Platforms with automatic item analysis (StudyDrome generates these statistics per item as soon as an exam closes) remove the spreadsheet step entirely, but the review logic is the same, regardless of how you compute the numbers.

  1. Run the numbers. For every item: P-value, point-biserial, and option-level choice frequencies split by top and bottom scorers. Compute KR-20/alpha and SEM for the whole exam.
  2. Triage flagged items. Flag anything with negative or near-zero discrimination, P below ~0.30, or a distractor outperforming the key among high scorers.
  3. Check keys first. For each flagged item, verify the keyed answer against the source. Miskeys are common and fixing them often resolves the statistics entirely.
  4. Adjudicate with content experts. For surviving flags, have two content experts read the item cold. Decide per item: keep as-is (statistics reflect a hard but fair item), accept multiple keys, or remove from scoring.
  5. Re-score and re-check reliability. After removals, recompute scores, KR-20, and the SEM. Confirm pass/fail decisions near the cut score with the new numbers.
  6. Feed the bank. Record the statistics against each item in your bank. Tag items to revise (weak distractors, marginal discrimination) and route them back to their authors with the data attached.
  7. Close the loop with writers. Aggregate recurring flaw patterns and use them in your next item-writing workshop. Statistics without feedback change nothing.

This workflow sits inside a larger governance cycle. For which of these numbers a committee should actually be shown, and when, see what reports a board of examiners needs.

Common pitfalls

Small cohorts. With 30–60 examinees, a typical module cohort, every statistic in this guide is noisy. A point-biserial of 0.05 in a class of 40 may be sampling error, not a flawed item. Treat statistics from small groups as screening signals that trigger reading the item, never as automatic delete rules. Accumulate statistics across administrations before making permanent bank decisions.

Over-deleting items. Removing every item below an arbitrary discrimination threshold shortens the exam, narrows blueprint coverage, and can perversely lower the validity of decisions even if alpha ticks up. Deletion is for defective items, not merely unflattering statistics. Downing's flaw research justifies removing broken items; it does not justify purging easy items on essential content.

Judging items on one number. A P-value without discrimination, or discrimination without reading the item, misleads. Always look at difficulty, discrimination, and the option table together — and then at the item text.

Chasing reliability with homogeneous content. You can inflate alpha by writing many near-clone items on the same narrow topic. That is redundancy, not measurement quality. Blueprint coverage comes first.

Comparing across cohorts naively. P-values and discrimination indices are population statistics. An item that behaved well with last year's cohort may behave differently after a curriculum change — that is information, not malfunction.

Takeaways

Written by Dimitri · Jul 5, 2026

Put this into practice with Exam Manager

Run a real exam with your own questions and see the results analysis on your own data — guided setup, no commitment.

Book a pilot

Share this post

Get the next article by email

Assessment and edtech articles, straight to your inbox. Double opt-in, unsubscribe anytime.