What is the borderline regression method in OSCE standard setting?
The borderline regression method sets an OSCE station's pass mark by regressing candidates' checklist scores onto examiners' global ratings, then reading off the predicted score at the borderline grade. Because it uses data from every candidate rather than only borderline ones, it yields stable, defensible cut scores for typical cohort sizes.
Borderline regression (BRM) is currently the most widely used standard-setting method for OSCEs in medical education, and the method most post-exam quality frameworks — notably Pell, Fuller, Homer & Roberts (2010) — are built around.
How it works
At each station, the examiner records two things for every candidate:
- A checklist or domain score — the detailed marks for the performance.
- A global rating — an overall judgment on a short ordinal scale, typically something like fail / borderline / pass / good / excellent.
After the exam, a linear regression is fitted for each station with the checklist score as the outcome and the global rating as the predictor, using all candidates who sat that station. The station's pass mark is the predicted checklist score at the "borderline" point of the global scale. Station pass marks are then combined (usually summed or averaged) to produce the exam-level cut score, often with an adjustment such as a standard error of measurement.
Why it is popular
- It is examinee-centered and data-rich. The standard reflects examiners' holistic judgments of actual performances, not hypothetical ones — unlike test-centered methods such as Angoff, where judges estimate how an imaginary borderline candidate would perform.
- It uses every data point. The regression line is anchored by the whole cohort, so the cut score is far less sensitive to sampling noise than the borderline group method, which discards everyone not rated borderline. See borderline group vs borderline regression.
- No extra judging panel. The standard is a by-product of scoring the exam itself, which keeps operational cost low.
- It generates quality metrics for free. The R-squared of each station's regression indicates how well checklist scores track global judgments; a low value flags a station where the checklist and the examiners' overall impression disagree — a station worth reviewing.
Requirements and caveats
BRM has assumptions worth respecting:
- The global rating must be independent and meaningful. Examiners need training so the global grade is a genuine holistic judgment, not a mechanical restatement of the checklist total.
- Cohort size matters. With very small cohorts (roughly a few dozen candidates per station or fewer), regression estimates become unstable; small-cohort programs sometimes pool data across administrations or prefer other methods.
- Watch for range restriction and outliers. If almost no candidates receive low global grades, the borderline point is an extrapolation. Routine plots and diagnostics per station are essential — this is exactly the monitoring the Pell et al. metrics formalize.
For how BRM compares with Angoff, Ebel, borderline group, and other approaches — and how to choose among them — see OSCE standard-setting methods compared and the primer on what standard setting is.