A reliable OSCE reports a Cronbach's alpha or a generalizability coefficient of 0.7 or above across stations. A valid OSCE can show, with evidence, that its scores support the decision made from them. Reliability is one part of that evidence. This page explains both the thresholds set by the literature and how many stations it takes to reach them.
Why an OSCE lead needs these two words
Every OSCE result is a claim. The claim is that a candidate who passed can do the thing the exam tested. A candidate who failed cannot yet. Validity is the question of whether that claim holds. Reliability is a narrower question. Would the same candidate get the same result on another day, with other examiners, on another set of stations from the same blueprint?
A lead who cannot answer both questions cannot defend a result. An appeal will ask them. An accreditation visit will ask them. And the numbers that answer them are cheap to compute, because every OSCE produces the data.
What does validity mean for an OSCE?
Validity is not a property of the exam. It is a property of the interpretation you place on its scores. The modern framing comes from Kane. Cook and colleagues summarized it for health professions. It treats validity as an argument in four steps. Each step is an inference, and each inference needs evidence.
Inference | The claim | What supports it in an OSCE |
|---|---|---|
Scoring | The examiner's marks capture the performance | A clear scoring scheme, trained examiners, agreement between examiners, no missed items |
Generalization | The stations sampled represent the whole domain | A blueprint that maps the domain, enough stations, an acceptable reliability coefficient |
Extrapolation | Performance in the exam reflects performance in practice | Stations built on real tasks, realistic simulated patients, scores that relate to later performance |
Implications | The pass or fail decision leads to the right outcome | A defensible standard, a sensible remediation route, consequences that fit the stakes |
Reliability sits in the second row. It is necessary for validity, not sufficient. An exam can be perfectly consistent and measure the wrong thing.
The AMEE guide on the OSCE's theoretical basis places the exam at the "shows how" level of Miller's pyramid. That is the extrapolation claim in one phrase: a candidate showed, under observation, that they can perform the task. An OSCE cannot show what a candidate does unsupervised in a clinic. Any validity argument should say so.
The twelve tips of Daniels and Pugh apply Kane's framework in the order of development: decide what the scores will be used for, then build each inference into the design. That order matters. Validity evidence gathered after the exam can only confirm or refute a design. It cannot repair one.
What does reliability mean, and which coefficient should you report?
Reliability is the share of score variation attributable to real differences between candidates. The rest comes from which stations they drew, which examiners they met, or which circuit they sat. Three coefficients appear in OSCE reports.
Cronbach's alpha treats the stations as items on a single test and measures their internal consistency. It is the most common coefficient because it needs only the station scores. AMEE Guide 49 states that a value of 0.7 or higher is usually considered acceptable. A value above 0.9 may indicate redundancy between stations. The AMEE organization guide gives the same working range: an alpha or G coefficient between 0.7 and 0.8 reflects acceptable reliability for a high-stakes exam.
The generalizability coefficient comes from a G-study. It separates the sources of error. It can tell you how much error comes from stations and how much from examiners. It needs a crossed or partially crossed design, which many OSCEs lack. When you can compute it, it is the more informative number.
Inter-rater agreement measures whether two examiners watching the same performance give the same marks. It is reported as a correlation, an intraclass coefficient, or kappa for categorical grades. It supports scoring inference, not generalization, and says nothing about how many stations the exam needs.
Report alpha every time. Report a G coefficient when the design allows it. Report inter-rater agreement when stations were double-marked, and say how many were.
How many stations and how much testing time does a reliable OSCE need?
The honest answer is that the evidence points in two directions, and a lead should know both.
The 2020 Ottawa Consensus Statement on Performance Assessment sets a floor. It states that an OSCE with fewer than 12 stations will not yield reliable results. The same holds for testing times below 150 minutes. That is the position a regulator or an external examiner will most often quote.
A 2025 meta-analysis by Peng and colleagues pooled 26 alpha values from 23 studies of medical student OSCEs. The median exam had 9 stations, each lasting 9 minutes. Exams with fewer than 5 stations averaged an alpha of 0.75. Exams with 5 to 10 stations averaged 0.83. Exams with more than 10 stations averaged 0.77. Within the 5-to-10-station group, stations under 10 minutes averaged 0.88 and stations of 10 minutes or more averaged 0.64. The authors read this as support for more, shorter stations. Two caveats apply. Heterogeneity between studies was very high, and the funnel plot suggested publication bias.
A 2011 systematic review by Brannick and colleagues provides the baseline the two positions dispute. Across 39 studies, the mean alpha across stations was 0.66. Reliability within a station across its items was 0.78. The mean generalizability coefficient was 0.49. In other words, the average published OSCE fell short of 0.7 at the exam level.
The practical reading is this. Reliability rises with the number of stations, because more stations sample more of the domain. It rises faster with more short stations than with fewer long ones. The main threat is content specificity: a candidate who is strong on one case is not reliably strong on the next. The AMEE organization guide calls content specificity "a major contributor to poor reliability". It offers 14 to 18 stations of 5 to 10 minutes as a general recommendation. Treat the Ottawa floor as the minimum for a high-stakes exam. Treat the meta-analysis as evidence that adding shorter stations is the better lever once you are past it.
What lowers OSCE reliability, and what fixes it?
Four sources of error account for most low coefficients. Each has a fix, and the fix lives in a different section of this guide.
Source of error | What it looks like in the data | Fix |
|---|---|---|
Too few stations | Alpha below 0.7 with no single weak station | Add stations; shorten them if the blueprint allows |
Content specificity | Low correlations between stations that were meant to measure the same skill | Sample the domain more widely; do not expect one case to stand for a competency |
Examiner variation | Circuits or examiners with different mean scores; hawks and doves | Training and calibration; scoring schemes with anchors; compare examiners after the exam |
Weak scoring schemes | Missed items, checklists that reward the wrong behaviors, global ratings without descriptors | Rewrite the scheme; pilot it; use electronic capture to stop missed items |
Examiner variation deserves a number. A study covered 442 sittings of a UK licensing OSCE. Examiner stringency accounted for 15.9 percent of the variance in domain scores. Training reduces that share but does not remove it. The examiner section of this guide covers what training can and cannot do.
How do you write a validity argument for your OSCE?
A validity argument is a short document the committee can keep with the exam's records. It states the decision the scores support, then lists the evidence for each inference. Most schools already hold the evidence; they have not written it down. The list below is the outline.
The limits of the numbers
Two limits matter. First, alpha depends on both the number of stations and their quality. A long exam of weak stations can post a respectable coefficient. Read alpha together with the station-level metrics in the quality section. Second, a coefficient describes the cohort that sat the exam. A small cohort gives a wide confidence interval. An exam with 30 candidates can swing from 0.6 to 0.8 between years with no change in the stations. Report the interval, not just the point estimate, and compare years with that in mind.
Key terms
Frequently asked questions
What is an acceptable reliability for an OSCE?
An alpha or generalizability coefficient of 0.7 or above is the widely used threshold for a summative OSCE. The AMEE guides describe 0.7 to 0.8 as acceptable for high-stakes. Licensing bodies sometimes aim higher. A value above 0.9 in a local exam is unusual and may mean stations are measuring the same thing twice.
Is a longer OSCE always more reliable?
More stations raise reliability because they sample more of the domain. Longer stations do not. A 2025 meta-analysis looked at exams of 5 to 10 stations. Stations under 10 minutes had higher average reliability than stations of 10 minutes or more. The reason is content specificity. Breadth of sampling matters more than depth on any one case.
Does inter-rater reliability tell me my OSCE is reliable?
No. Inter-rater agreement shows that two examiners marked the same performance the same way. It supports the scoring step of the validity argument. Exam-level reliability, the alpha or G coefficient, asks whether the set of stations gives a stable picture of the candidate. An exam can have excellent examiner agreement and too few stations.
What is the difference between validity and reliability?
Reliability asks whether the scores are consistent. Validity asks whether the scores mean what you say they mean. A consistent exam can still measure the wrong thing. Reliability is necessary for validity but never enough on its own. The validity argument adds evidence for scoring, generalization, extrapolation, and implications.
Do I need a psychometrician to compute these?
No. Cronbach's alpha needs only the station scores and a spreadsheet or any statistics package. A generalizability study needs more design work and someone who has run one before. Station-level metrics and examiner comparisons are arithmetic. What a school needs is someone who can read the output and explain it to the committee.