What is a point-biserial correlation, and what value is acceptable?
The point-biserial correlation measures how strongly performance on one item correlates with total exam score. Positive values mean stronger examinees answered correctly more often. Values above roughly 0.20 are generally acceptable, 0.30 and above is good, and items near zero or negative deserve review before scoring is finalized.
The point-biserial correlation (often written r-pbis) is the standard item-discrimination statistic in classical test theory. For each question it correlates a dichotomous variable — did the examinee get this item right? — with a continuous one, the examinee's total exam score. It answers the question every item should pass: do the people who know more do better on this item?
How to read the value
The point-biserial ranges from -1 to +1. In practice, values for reasonable items fall between about 0.10 and 0.50.
Point-biserial | Interpretation | Action |
|---|---|---|
≥ 0.30 | Good discriminator | Keep; bank as a proven item |
0.20 - 0.29 | Acceptable | Keep; consider minor revision |
0.10 - 0.19 | Marginal | Review stem and options |
0.00 - 0.09 | Not discriminating | Revise or replace |
Negative | Problem item | Check the key immediately |
Bands in this vicinity appear throughout the educational-measurement literature — Ebel and Frisbie (1991) is a commonly cited source for discrimination benchmarks — but they are guidance, not statute. Two caveats matter when applying them:
- Difficulty constrains discrimination. Items that nearly everyone answers correctly (or incorrectly) cannot discriminate much, because there is almost no variance to correlate. A p-value of 0.95 with a point-biserial of 0.12 may be a perfectly good easy item placed deliberately at the start of an exam.
- Small cohorts produce noisy estimates. With 30 examinees, a point-biserial can swing substantially between administrations by chance alone. Treat single-administration values from small groups as indicative, not definitive, and rely on item history across uses.
Many systems report the corrected point-biserial, which excludes the item itself from the total score before correlating. This removes a small built-in inflation (an item always correlates somewhat with a total that contains it) and matters most on short tests; on an 80-item exam the correction is minor.
Why assessment teams watch it
Discrimination is the single most actionable item statistic. Items with strong point-biserials pull exam reliability up; weak items dilute it (see what makes a good KR-20). And a negative point-biserial is the classic signature of a miskeyed or ambiguous item — the highest-priority flag in any post-exam review, since flawed items have been shown to penalize precisely the strongest students (Tarrant & Ware, 2008).
In routine practice the workflow is simple: sort the post-exam item report by point-biserial ascending, verify keys on anything negative, review anything under about 0.20 alongside its distractor statistics, and let well-discriminating items accumulate a track record in the bank.