Skip to content

How to Write Flaw-Free MCQs: What the Evidence Says

The item-writing rules with actual research behind them, and the review loop that keeps flawed questions off your exams.

MCQ writing

Multiple-choice questions look easy to write. That is precisely the problem. A plausible-looking MCQ can be assembled in ten minutes; a sound one (an unambiguous stem, one defensible key, distractors that work, alignment with what was taught) takes real craft. And the gap between the two is not cosmetic. It shows up in your students' scores and your pass lists.

The good news is that MCQ writing is one of the best-researched corners of assessment. This post distills what the evidence actually supports and lays out a review process that catches flaws before an exam rather than after.

What flawed items cost

Item-writing flaws are not neutral noise. Downing's studies of exams in medical education (2004, 2005) found that flawed items behave as construct-irrelevant variance: they make questions harder or easier for reasons unrelated to the knowledge being tested. Two findings deserve particular attention. First, flawed items can disadvantage precisely the able students. The careful reader is the one who notices the ambiguity, the second defensible option, the trick in the negative phrasing. Second, because flawed items distort scores, they can affect pass/fail decisions near the cut score, meaning students may fail (or pass) partly on the quality of your item writing rather than their knowledge.

Tarrant and Ware, examining nursing and health-science examinations, reached compatible conclusions: flawed items were widespread in teacher-written tests and had measurable consequences for examinees, with high-achieving students disproportionately affected. In short: every flaw you fail to catch is a small tax levied on your best students. That is the case for taking the rules seriously.

The taxonomy worth knowing

The reference point for the whole field is the taxonomy of item-writing guidelines assembled by Haladyna, Downing & Rodriguez (2002), who reviewed the measurement literature and textbook consensus to produce a validated set of rules. You do not need to memorize all of it. The rules that do the most work in medical education, and the flaws they prevent, cluster into a shortlist, and they map directly onto the most common MCQ writing flaws found in real exam banks:

  • Test important content, not trivia. Each item should map to a blueprint objective (more below). If the item exists because the fact was easy to write about, cut it.
  • One clear problem in the stem. The stem should pose a complete question. In clinical items, the vignette-plus-lead-in format of the NBME's item-writing guide (Case & Swanson) is the model: a focused clinical scenario followed by a single direct question.
  • One unambiguously best answer, defensible against the literature — not merely the "most expected" answer.
  • No cueing. Grammatical mismatches between the stem and options, the longest option being correct, word repeats between the stem and the key, and convergence cues all allow testwise students to answer without knowledge.
  • No trick items. Difficulty should come from the content, not from wording designed to mislead.

Writing the stem

Two rules cover most stem problems.

Pass the cover-the-options test. A well-built item can be answered, at least provisionally, with the options hidden. If a knowledgeable reader covers the options and cannot even begin to answer, the stem is not posing a question; it is a fragment waiting for the options to define it ("Which of the following is true about warfarin?"). The cover-the-options rule is the fastest single screen you can apply in peer review, and it naturally pushes writers toward focused lead-ins like "What is the most appropriate next step in management?"

Avoid negatively worded stems. "Each of the following is correct EXCEPT" and "Which is NOT a feature of…" items are consistently discouraged in the Haladyna, Downing & Rodriguez guidelines and in the NBME guide. Negatives add a comprehension hurdle unrelated to the content, are error-prone under time pressure, even for students who know the material, and typically arise when the writer has several true statements and no single best answer. That is a sign the item is testing an ill-posed question. The full argument is in why to avoid negatively worded stems. If a negative is truly unavoidable, emphasize it typographically and never stack a negative stem with negative options.

Prefer clinical vignettes where the objective allows: Case & Swanson's guidance is that application-level items (patient presentation, then a decision) assess the reasoning medical schools actually care about, while still following every rule above.

Options: how many, and which

Three options are usually enough. The reflexive standard of four or five options has little evidence behind it. Rodriguez (2005) synthesized 80 years of research and concluded that three options (one key, two distractors) is generally the sweet spot: most fourth and fifth distractors attract almost no responses, so removing them barely changes item difficulty or discrimination, while the saved writing and reading time lets you ask more questions, which does improve reliability. Empirical distractor studies (Tarrant and Ware among them) confirm that non-functioning distractors are rife in real exams. The practical rule in how many distractors an MCQ needs: write as many plausible distractors as the content supplies, and stop. Two good ones beat four fillers.

Make distractors work for a living. A plausible distractor is a statement a student with a specific misconception, or incomplete knowledge, would actually choose: a drug from the same class, the diagnosis one step earlier in the reasoning chain, the value off by the classic error. The best sources are real student errors: short-answer responses, tutorial misunderstandings, and the option-level data from previous exams. Homogeneity matters too: options should be parallel in content, grammar, and length, or the odd one out becomes a cue.

Retire "all of the above" and "none of the above." All-of-the-above is broken by construction: a student who can verify any two options can deduce the answer, and a student who spots one false option eliminates it. Partial knowledge is rewarded either way. None-of-the-above has narrower legitimate uses (notably calculation items), but as a keyed answer it verifies only that students rejected your distractors, not that they know the right answer. The Haladyna, Downing & Rodriguez review recommends against both as routine practice.

The review process: where flaws actually get caught

Rules make individual writers better; process makes exams better. Flaw-free exams come from a loop with three checkpoints.

1. Blueprint first. Every item should be commissioned against an exam blueprint: a grid that maps content areas and competency levels to numbers of items. Blueprinting is a validity requirement in its own right (it is the content-evidence backbone in the frameworks of Downing and of Yudkowsky, Park & Downing's assessment text), but it also prevents a specific writing failure: orphan items written about whatever the author found interesting, which is where trivia and trick questions breed.

2. Peer review before administration. No item goes on an exam read only by its author. A structured review (two reviewers, working from a short checklist like the one below) catches most flaws cheaply. Reviewers should answer the item cold before seeing the key; disagreement between reviewer and key is the single most informative red flag. The review also checks blueprint fit, appropriateness of difficulty, and consistency of vignette details (age, sex, and findings that don't contradict each other).

3. Item analysis after administration. Some flaws only surface in data. Post-exam item analysis (difficulty, discrimination, and especially distractor analysis) closes the loop: a negative point-biserial exposes the miskeyed or ambiguous item, and option-level frequencies show which distractors earn their place and which are dead weight to prune before the item returns to the bank. If your delivery platform computes these automatically (StudyDrome's item analysis produces the per-option statistics as soon as an exam closes), route each writer's item statistics back to them; nothing improves item writers faster than seeing their own distractors ignored by 200 students.

Programs that run this loop consistently accumulate a bank of items with known statistics and documented review history. That is also exactly the evidence trail you want when a scoring decision is challenged.

Item-writing checklist

Run every item through this before it enters the bank:

Content

Stem

Options

Process

Takeaways

Written by Dimitri · Jul 5, 2026

Put this into practice with Exam Manager

Run a real exam with your own questions and see the results analysis on your own data — guided setup, no commitment.

Book a pilot

Share this post

Get the next article by email

Assessment and edtech articles, straight to your inbox. Double opt-in, unsubscribe anytime.