What does a negative point-biserial correlation mean?
A negative point-biserial means examinees who scored well overall answered that item incorrectly more often than weaker examinees. It almost always signals a problem — a miskeyed answer, an ambiguous or flawed stem, or a plausible distractor that misleads knowledgeable students — and the item should be reviewed before scores are released.
A negative point-biserial correlation inverts what an exam item is supposed to do. Instead of rewarding knowledge, the item is answered correctly more often by low-scoring examinees than by high-scoring ones. Statistically it means getting the item right is associated with a lower total score. On a functioning exam this should essentially never happen for legitimate reasons, so a negative value is the highest-priority flag in any item-analysis report.
The usual causes, in order of likelihood
- A miskeyed answer. The most common and most fixable cause. If the recorded key is wrong, the students who actually know the material select the true correct answer, get marked wrong, and the point-biserial flips negative. Always verify the key first — the fix (rekey and rescore) takes minutes.
- An ambiguous or flawed item. Two defensible options, an imprecise stem, an unnoticed "except"/negative phrasing, or outdated content can split knowledgeable examinees between the keyed answer and a defensible alternative. This is consistent with the broader finding that item flaws systematically disadvantage high achievers (Tarrant & Ware, 2008): strong students read more into a flawed stem than the writer intended, while weaker students pick the superficially obvious keyed option.
- An over-plausible distractor. A distractor that is nearly correct — or that reflects a more sophisticated but non-keyed line of reasoning — can pull in exactly the students with the deepest understanding. Distractor analysis makes this visible: the tell is a wrong option whose selectors have a higher mean total score than the selectors of the key.
- Chance, on small cohorts. With few examinees, sampling noise alone can produce a mildly negative value on an otherwise sound item. Treat small-cohort statistics as a prompt for review, not proof of a flaw, and check the item's history across administrations.
What to do when you see one
- Before releasing scores: verify the key against the source material; have a content expert re-read the item cold; inspect which option the top quintile chose.
- If the item is miskeyed: correct the key and rescore.
- If the item is flawed: the defensible options are to remove it from scoring or credit multiple responses, then revise or retire it in the bank. Leaving a confirmed negative-discrimination item in the score actively subtracts measurement value — it adds error variance and drags down reliability.
- Either way: record the outcome against the item so the bank accumulates evidence.
Because a miskey caught after results are published is far more damaging than one caught before, most programs make a negative point-biserial an automatic hold in the post-exam workflow; platforms such as StudyDrome flag these items automatically as soon as an exam closes. For where this check sits in the overall review sequence, see the complete guide to item analysis.