Every OSCE ends with the same question: where is the line between pass and fail? The answer cannot be a tradition ("60% because it has always been 60%") or a quota ("the bottom decile fails"). A pass mark is a claim that candidates above it are competent enough to progress, and it requires a defensible method to support it.
This post compares the standard-setting methods actually used for OSCEs, with enough operational detail to run them. It assumes you know what standard setting is in general; the focus here is the specific problem of performance exams, where the OSCE format Harden introduced in 1975 gives you many short stations, examiner judgments, and checklist-plus-rating score data that written-exam methods were never designed for.
Why norm-referenced passing is indefensible
Norm-referenced standards (fail the bottom N%, or pass everyone above one standard deviation below the mean) are seductively easy and fundamentally wrong for competence decisions. The pass mark moves with the cohort: the identical performance passes in a weak year and fails in a strong one. That is incompatible with the purpose of a licensure-style examination, which is to certify that a candidate can do the job, regardless of who happened to sit alongside them.
Norcini's AMEE Guide No. 18 on standard setting makes the core point plainly: for tests of competence, the standard should be criterion-referenced, anchored to a conception of adequate performance rather than to the distribution of scores. Cizek & Bunch's standard-setting text adds the procedural corollary: since all standards are ultimately judgments, what makes one defensible is the method — who judged, using what process, with what evidence of internal consistency. Norm-referencing fails on both counts. Every method below is criterion-referenced; they differ in whose judgment is captured and when.
Angoff and modified Angoff
The Angoff method (Angoff 1971) is the workhorse of written-exam standard setting. A panel of judges first builds a shared mental model of the borderline candidate, one who just barely deserves to pass. Then, for each item, each judge estimates the probability that this borderline candidate would answer correctly. Sum a judge's estimates to get that judge's cut score; average across judges for the panel's standard. The mechanics of the Angoff method include a definition-building discussion, independent first-round ratings, and usually a second round after judges see each other's estimates and, in some variants, real difficulty data.
"Modified Angoff" is an umbrella for the many practical adaptations: yes/no judgments per item instead of probabilities, feeding performance data back between rounds, and so on. The differences between Angoff and modified Angoff matter less than running whichever variant you choose with discipline: trained judges, a written borderline-candidate definition, and documented rounds. Panel composition matters too: how many judges an Angoff panel needs is a trade-off between cut-score stability and faculty time constraints, with more (and more diverse) judges yielding a more reproducible standard.
Why Angoff is awkward for OSCEs. Angoff asks judges to imagine performance on discrete items. OSCE stations are not items. They are five- to ten-minute performances scored using multi-point checklists and rating scales. Asking a judge to estimate the probability that a borderline candidate scores each individual checklist point produces enormous rating burdens and strains judgment: "What is the chance the borderline candidate palpates the abdomen in all four quadrants?" is a much less natural judgment than "Would a minimally competent student pass this item?" Angoff-style methods also take place away from the exam, in a conference room, whereas the OSCE places expert judges in front of every real performance. The examiner-based methods below exploit that.
Borderline group method
The borderline group method moves the judgment into the exam itself. At each station, alongside the checklist, the examiner gives a global judgment of overall performance, typically on a scale such as fail/borderline/pass/good pass. After the exam, you collect all candidates rated borderline at a station and set that station's pass mark at the mean (or median) of their checklist scores. The borderline candidates define the standard directly: the cut is literally the typical score of performances that expert observers judged to be borderline.
The method's appeal is its face validity and simplicity. Its weakness is statistical: only the borderline-rated candidates contribute to the cut score. At a single station in a modest cohort, there may be only a handful of performances, making the standard unstable and vulnerable to idiosyncratic examiner judgments. It also discards the information in every non-borderline performance.
Borderline regression: the de facto OSCE standard
Borderline regression fixes exactly that flaw, which is why it has become the default in OSCE practice, particularly in UK and European medical schools. The borderline regression method uses the same data collection (a checklist score and a global rating per candidate per station), but instead of averaging only the borderline group, it regresses checklist scores on global ratings across all candidates at the station. The fitted regression line is then read off at the borderline point of the global scale: the predicted checklist score at "borderline" becomes the station's pass mark. Sum station cut scores (or average, depending on your decision rule) for the exam-level standard.
The comparison between borderline group and borderline regression comes down to data efficiency: regression uses every candidate's data, so the standard is more stable, especially where few candidates are rated borderline. Pell, Fuller, Homer and Roberts (2010), in their widely used AMEE guide to OSCE quality metrics, work within exactly this framework and treat the checklist–global-rating relationship as a central quality signal.
Operational requirements. Borderline regression is not free. It requires:
- Two scores per candidate per station: a checklist (or domain) score and an independent global rating. If your examiners only complete checklists, you cannot run it. The design question of checklists versus global rating scales is therefore not either/or for BRM: you need both, and the global rating should be a holistic judgment, not a mechanical restatement of the checklist.
- Examiner training so that global ratings are made independently, and the "borderline" anchor means something shared.
- Adequate cohort size per station. A regression fitted on very few candidates is fragile; small-cohort schools (or stations seen by few candidates in multi-track circuits) should check the fit carefully and may prefer to pool data across administrations or fall back on other methods. There is no universally agreed-upon minimum, and what the evidence actually says about borderline regression sample size is that the spread of the global grades predicts stability better than headcount does.
- Sensible score distributions. Range restriction (everyone scores high), clustering of global ratings at one point, or a handful of extreme outliers can distort the fitted line.
Station count matters too: the reliability of the overall OSCE decision depends far more on having enough stations than on polishing any single station's cut score, because case specificity makes performance vary across content.
Contrasting groups, briefly
In the contrasting-groups method, judges classify candidates into two groups, competent and not competent, based on holistic review of their performance, independent of scores. You then plot the score distributions of the two groups and set the cut where they intersect (or shift to weight false positives versus false negatives according to the stakes). It is conceptually clean and forces an explicit conversation about which classification error is worse. In OSCE practice, it is used less than the borderline methods, mainly because it needs a credible external classification of whole candidates, which is harder to obtain than station-level global ratings. Cizek & Bunch cover it as one of the standard examinee-centered methods; it is worth knowing for cross-checking.
How to choose
Method | Judgment happens | Data required | Strengths | Watch out for |
|---|---|---|---|---|
Angoff / modified Angoff | Before the exam, panel | Item/checklist content; trained panel | Well-researched; works without performance data | Awkward fit to performance stations; heavy panel workload |
Borderline group | During the exam, examiners | Checklist + global rating | Simple, direct, high face validity | Unstable with few borderline candidates |
Borderline regression | During the exam, examiners | Checklist + global rating, all candidates | Uses all data; stable; the OSCE default | Needs adequate n, independent global ratings, sane distributions |
Contrasting groups | After the exam, judges | Independent competent/not classification | Explicit error trade-off | Credible classifications hard to obtain |
Practical guidance:
- Running a standard multi-station OSCE with checklist + global scoring? Borderline regression is the default choice, with borderline group as the simpler fallback.
- Very small cohorts or new stations with no data? Consider an Angoff-style panel standard, or pool borderline-regression data across sittings before relying on it.
- Whatever you choose, document it: the borderline-candidate definition, judge/examiner training, the computation, and the decision rules. Following Kane's argument-based view of validity, the defensibility of a cut score rests on the documented process and evidence, not on the method's name.
After the cut score: quality metrics
Setting the standard is not the end of the job. Pell et al. (2010) describe a set of post-hoc metrics that have become the common quality dashboard for OSCEs run with borderline regression. Check them every cycle, per station:
- R² of the regression: how much of the checklist-score variance the global ratings explain. A low R² at a station means the two judgments are disconnected: examiners' holistic view does not track the checklist, so the derived cut is on shaky ground. Investigate the station and examiner behavior rather than just accepting the number.
- Between-group variation in checklist scores across global-rating groups: candidates rated higher globally should score meaningfully higher on the checklist; weak separation signals a station or scoring problem.
- Overall reliability (alpha across stations) and inter-grade discrimination: how much checklist scores rise per step of the global scale.
- Failure-rate anomalies: a station failing a wildly different proportion than its peers deserves review, as do large differences between parallel circuits or sites.
Treat these as screening tools, exactly like item analysis for written exams: a flagged station triggers human review of the station, its checklist, and its examiners — not automatic exclusion. If your OSCE management platform computes station-level regression statistics automatically (StudyDrome's exam statistics do this alongside its item analysis for written papers), run the dashboard the same day scores close, while examiner memory is fresh.