Skip to content

How many stations does an OSCE need to be reliable?

Most OSCEs need roughly 14–18 stations, totaling three to four hours of testing time, to reach a generalizability coefficient near 0.8. Reliability depends mainly on total testing time and station count, not extra examiners per station, because candidate performance is case-specific and varies substantially from station to station.

There is no single magic number, but the literature converges on a clear range: for a high-stakes pass/fail decision, an OSCE generally needs on the order of 14–18 stations and three to four hours of total testing time to reach a reliability (generalizability) coefficient around 0.8. Shorter circuits — say 6–8 stations — are common for formative or lower-stakes exams but rarely support defensible high-stakes decisions on their own.

Why station count dominates

The main threat to OSCE reliability is not examiner disagreement but content specificity: a candidate who performs well on a cardiology examination station may perform poorly on a communication station. Because performance varies so much from case to case, the only way to get a stable estimate of overall competence is to sample many cases. This is why generalizability studies (summarized in van der Vleuten and Swanson's classic work on performance-based assessment, and in Yudkowsky, Park & Downing's Assessment in Health Professions Education) consistently show that adding stations improves reliability far more than adding a second examiner to existing stations.

Practical implications:

Design choice

Effect on reliability

More stations (broader case sampling)

Large improvement — the dominant factor

Longer total testing time

Improves reliability; how it is divided into stations matters less

Second examiner per station

Small improvement; rarely worth the cost

Better checklists at few stations

Cannot compensate for narrow sampling

What reliability target to aim for

A common convention is a generalizability coefficient (or Cronbach's alpha across stations) of at least 0.7 for moderate-stakes exams and 0.8 or higher for high-stakes licensure-style decisions. Note that inter-station alpha is only one lens: Pell, Fuller, Homer & Roberts (2010) recommend a battery of post-hoc metrics — including R-squared from borderline regression, failure rates per station, and between-group variation — because a respectable alpha can coexist with badly behaving individual stations.

Practical advice

  • Blueprint first. Station count should also be driven by content coverage: enough stations to sample each domain in the blueprint, not just to hit a reliability number.
  • Prefer more short stations over fewer long ones, within reason — sampling breadth beats depth for score reliability.
  • Measure, don't assume. Run the reliability analysis after every administration and track it over time; platforms with built-in item and station analysis, such as StudyDrome, make this routine rather than an annual statistics project.
  • Combine evidence across the program. If a single sitting cannot reach the target, sequential or programmatic designs that combine multiple assessment points can.

Reliability is necessary but not sufficient: the pass mark itself must also be set defensibly. See what is the borderline regression method? and the broader comparison in OSCE standard-setting methods compared.

Updated

← All OSCE Assessment questions