How many stations does an OSCE need to be reliable?
Most OSCEs need roughly 14–18 stations, totaling three to four hours of testing time, to reach a generalizability coefficient near 0.8. Reliability depends mainly on total testing time and station count, not extra examiners per station, because candidate performance is case-specific and varies substantially from station to station.
There is no single magic number, but the literature converges on a clear range: for a high-stakes pass/fail decision, an OSCE generally needs on the order of 14–18 stations and three to four hours of total testing time to reach a reliability (generalizability) coefficient around 0.8. Shorter circuits — say 6–8 stations — are common for formative or lower-stakes exams but rarely support defensible high-stakes decisions on their own.
Why station count dominates
The main threat to OSCE reliability is not examiner disagreement but content specificity: a candidate who performs well on a cardiology examination station may perform poorly on a communication station. Because performance varies so much from case to case, the only way to get a stable estimate of overall competence is to sample many cases. This is why generalizability studies (summarized in van der Vleuten and Swanson's classic work on performance-based assessment, and in Yudkowsky, Park & Downing's Assessment in Health Professions Education) consistently show that adding stations improves reliability far more than adding a second examiner to existing stations.
Practical implications:
Design choice | Effect on reliability |
|---|---|
More stations (broader case sampling) | Large improvement — the dominant factor |
Longer total testing time | Improves reliability; how it is divided into stations matters less |
Second examiner per station | Small improvement; rarely worth the cost |
Better checklists at few stations | Cannot compensate for narrow sampling |
What reliability target to aim for
A common convention is a generalizability coefficient (or Cronbach's alpha across stations) of at least 0.7 for moderate-stakes exams and 0.8 or higher for high-stakes licensure-style decisions. Note that inter-station alpha is only one lens: Pell, Fuller, Homer & Roberts (2010) recommend a battery of post-hoc metrics — including R-squared from borderline regression, failure rates per station, and between-group variation — because a respectable alpha can coexist with badly behaving individual stations.
Practical advice
- Blueprint first. Station count should also be driven by content coverage: enough stations to sample each domain in the blueprint, not just to hit a reliability number.
- Prefer more short stations over fewer long ones, within reason — sampling breadth beats depth for score reliability.
- Measure, don't assume. Run the reliability analysis after every administration and track it over time; platforms with built-in item and station analysis, such as StudyDrome, make this routine rather than an annual statistics project.
- Combine evidence across the program. If a single sitting cannot reach the target, sequential or programmatic designs that combine multiple assessment points can.
Reliability is necessary but not sufficient: the pass mark itself must also be set defensibly. See what is the borderline regression method? and the broader comparison in OSCE standard-setting methods compared.