Should OSCEs use checklists or global rating scales?
Neither is universally superior. Checklists suit junior learners and procedural tasks where completeness matters; global rating scales capture expertise, sequencing, and judgment, and in experienced examiners' hands they are at least as reliable. Many OSCEs combine both — checklists for feedback and station scores, a global rating for standard setting.
This is one of the longest-running debates in performance assessment, and the honest answer from the literature is: it depends on what the station is measuring and who is examining.
What each instrument does well
Checklists decompose a task into observable items ("washed hands," "asked about drug allergies," "checked distal pulses"), each scored done/not done or on a short scale. Their strengths:
- Transparency and perceived objectivity — useful when examiners are less experienced or non-expert.
- Rich, item-level feedback for learners.
- Good fit for procedural and safety-critical tasks where completeness genuinely is the construct.
Their well-documented weakness is that thoroughness is not expertise. Research on expert performance (notably work by Hodges and colleagues) showed that experienced clinicians can score lower on checklists than trainees, because experts take efficient, hypothesis-driven shortcuts that skip checklist items. Checklists can reward exhaustive, scripted behavior over clinical judgment.
Global rating scales ask the examiner for holistic judgments on a small number of dimensions (e.g., overall performance, communication, organization) using anchored ordinal scales. Their strengths:
- They capture qualities checklists miss: sequencing, prioritization, rapport, and judgment.
- In the hands of expert examiners, they show inter-station reliability at least as high as checklists — a finding replicated repeatedly since the 1990s (e.g., work by Regehr, Norman and colleagues).
- They discriminate better at the expert end of the performance spectrum.
Their cost is dependence on examiner expertise and calibration, and less granular feedback.
The pragmatic answer: use both
Most modern high-stakes OSCEs use a hybrid: a checklist or domain-based score sheet as the station score, plus an independent global rating. This is not just hedging — the pairing is what makes borderline regression standard setting possible, since the method regresses checklist scores on global grades. It also generates a built-in quality check: stations where the two instruments disagree badly (low regression R-squared, in the Pell, Fuller, Homer & Roberts (2010) framework) are flagged for review.
Design guidance consistent with the literature and AMEE Guide No. 81 (Khan et al.):
- Match the instrument to the construct. Procedural completeness → checklist-weighted; clinical reasoning, communication, professionalism → rating-scale-weighted or domain-based rubrics.
- Match it to your examiners. Expert clinician examiners can be trusted with global judgments; lay or less experienced raters need more structured instruments.
- Consider learner level. Checklists work well for early learners acquiring routines; global ratings differentiate better among senior students and postgraduates.
- Keep the global rating independent. Examiner training should stress that the global grade is a separate holistic judgment, not the checklist total restated.
In short: don't frame it as either/or. Use the checklist to score and teach, the global rating to judge and set standards, and the disagreement between them to audit your stations.