OSCE Assessment
How OSCEs work and how to assure their quality — station design, reliability, standard setting, scoring instruments, and examiner training.
The Objective Structured Clinical Examination (OSCE), introduced by Harden et al. (1975), is the dominant format for assessing clinical skills in medical education. Candidates rotate through a circuit of timed stations, each testing a defined skill — history taking, physical examination, communication, procedures — scored by an examiner against a predetermined instrument. Because every candidate faces the same tasks under the same conditions, the OSCE is far more standardized than traditional long-case or viva examinations.
Standardization alone does not guarantee quality. An OSCE's defensibility rests on evidence: enough stations to sample broadly across cases, a blueprint tied to learning outcomes, credible standard-setting methods such as borderline regression, trained and calibrated examiners, and routine psychometric monitoring using metrics like station-level reliability and R-squared (Pell, Fuller, Homer & Roberts, 2010). Weakness in any of these can produce passing decisions that are hard to defend to students, regulators, or the public.
This topic covers the practical questions assessment leads and OSCE coordinators face: how many stations are enough, which standard-setting method to use, checklists versus global ratings, and how to train examiners.
Questions & answers
What is the difference between the borderline group and borderline regression methods?
The borderline group method sets the pass mark as the mean checklist score of candidates rated borderline by examiners. Borderline regression instead fits a regression of checklist scores on global ratings across all candidates and predicts the score at the borderline point. Regression uses the full dataset, so it is more stable with small cohorts.
Read the full answer →Should OSCEs use checklists or global rating scales?
Neither is universally superior. Checklists suit junior learners and procedural tasks where completeness matters; global rating scales capture expertise, sequencing, and judgment, and in experienced examiners' hands they are at least as reliable. Many OSCEs combine both — checklists for feedback and station scores, a global rating for standard setting.
Read the full answer →How many stations does an OSCE need to be reliable?
Most OSCEs need roughly 14–18 stations, totaling three to four hours of testing time, to reach a generalizability coefficient near 0.8. Reliability depends mainly on total testing time and station count, not extra examiners per station, because candidate performance is case-specific and varies substantially from station to station.
Read the full answer →How do you train OSCE examiners?
Effective OSCE examiner training combines orientation to the exam blueprint and scoring instruments, calibration exercises using recorded or live performances with group discussion of discrepant scores, clear behavioral anchors for global ratings, and post-exam feedback using examiner metrics such as stringency. Brief one-off briefings alone do little to change scoring behavior.
Read the full answer →What is an OSCE in medical education?
An OSCE (Objective Structured Clinical Examination) is a performance-based exam in which candidates rotate through a series of timed stations, each assessing a specific clinical skill against a predetermined marking scheme. Introduced by Harden et al. (1975), it standardizes clinical assessment by exposing every candidate to the same tasks, examiners, and scoring criteria.
Read the full answer →What is the borderline regression method in OSCE standard setting?
The borderline regression method sets an OSCE station's pass mark by regressing candidates' checklist scores onto examiners' global ratings, then reading off the predicted score at the borderline grade. Because it uses data from every candidate rather than only borderline ones, it yields stable, defensible cut scores for typical cohort sizes.
Read the full answer →