Skip to content

How do you train OSCE examiners?

Effective OSCE examiner training combines orientation to the exam blueprint and scoring instruments, calibration exercises using recorded or live performances with group discussion of discrepant scores, clear behavioral anchors for global ratings, and post-exam feedback using examiner metrics such as stringency. Brief one-off briefings alone do little to change scoring behavior.

Examiners are the measurement instrument in an OSCE. Untrained or uncalibrated examiners introduce systematic error that no amount of station design can fix afterwards — hawks and doves, halo effects, checklist-global conflation. AMEE Guide No. 81 (Khan et al., Med Teach 2013) makes the underlying point directly: untrained examiners affect the reliability of the OSCE, so examiner training is a core organizational requirement of a defensible exam, not a courtesy briefing.

What a sound training program includes

1. Orientation. Before examining, every examiner should understand the exam's purpose and stakes, the blueprint, the station they will mark (including the expected performance and common candidate errors), and the mechanics: timing, permitted prompts, and how to interact (or not) with the simulated patient.

2. Instrument training. Walk through the checklist item by item, agreeing on what counts as "done." For global rating scales, provide behavioral anchors for each grade and stress that the global judgment is an independent holistic rating — critical if the exam uses borderline regression standard setting, which depends on that judgment being genuine.

3. Calibration practice. The core of training: examiners score the same performances — video recordings or live role-plays at multiple ability levels, ideally including a borderline performance — then compare scores and discuss discrepancies until the group converges on a shared standard. Discussing why scores differed is where calibration actually happens; simply watching videos without discussion adds little.

4. Bias awareness. Cover known rater effects: stringency/leniency (hawk–dove), halo (letting one strong dimension color everything), central tendency, contrast effects from the preceding candidate, and fatigue across a long circuit. Awareness does not eliminate these effects, but it supports the monitoring that does.

5. Post-exam feedback and re-training. Training is a cycle, not an event. After each administration, analyze examiner-level data — mean scores and global grades by examiner, flagging outliers relative to peers marking the same station across parallel circuits. Pell, Fuller, Homer & Roberts (2010) show how between-group variation and related metrics expose examiner effects; assessment platforms with automatic exam statistics, such as StudyDrome, make producing these examiner reports routine. Feed results back individually and target refresher training where the data warrant it.

Practical points

  • Keep sessions short but active. A 60–90 minute workshop built around scoring and discussing 2–3 recorded performances outperforms a long lecture.
  • Train for the specific exam. Generic examiner training transfers poorly; calibrate on the actual stations and instruments in use.
  • Maintain an examiner register. Track who has been trained, when, and on what, with periodic renewal — regulators and appeals panels ask.
  • Set expectations honestly. The literature on rater training (summarized in Yudkowsky, Park & Downing's Assessment in Health Professions Education) is sobering: training reduces but does not remove examiner variance. Design compensates for the remainder — more stations, so no single examiner dominates a candidate's result (see how many stations does an OSCE need?).

Train, measure, feed back, repeat: examiner quality is a monitored system, not a one-off induction.

Updated

← All OSCE Assessment questions