Four methods, chosen per exam: Custom, Angoff, Borderline Group, and Borderline Regression.
Borderline regression fits checklist scores against examiner Global Ratings per station using ordinary least squares. It needs at least 3 valid candidates. It returns slope, intercept, and R² to four decimal places, then clamps the derived mark into the achievable range. The exam pass mark is the sum of the station pass marks.
Four warnings surface during calculation, so whoever is about to publish a cut score sees why it might not hold:
Warning | What it means |
|---|
Same rating for all | Every candidate received the same Global Rating, so the regression is undefined. Borderline Group is the method to use. |
Negative slope | Checklist scores fall as Global Ratings rise. Review the station and the examiner. |
Low R² | Below 0.5, checklist and examiner judgment agree poorly. |
Small cohort | Fewer than 20 completed candidates, so the result may be unstable. |
The R² threshold follows Pell and Roberts, AMEE Guide 49, which treats a correlation above 0.5 as a reasonable relationship between checklist and examiner judgment. The cohort figure is a conservative floor: the resampling literature puts serious instability higher, and the spread of the global grades predicts it better than the headcount. Our guide to borderline regression sample size sets out the evidence.
The warnings are advisory. The exam board decides, and the record shows what it was told.
Every station carries its own controls: lock the mark against recalculation, mark it must-pass, or withdraw it from the calculation. An exam can also require a minimum number of stations passed.