An OSCE works by three design ideas. Every candidate meets the same task under the same conditions, so the exam is standardized. Each candidate rotates through many short stations, so the exam samples widely. Trained examiners score what they observe against a shared scheme, so the judgment is structured. Harden's original 1975 design used these three ideas to address the long case's unreliability.
Why the design matters to the people who run it
An OSCE is easy to describe and hard to run well. The description is a circuit of timed stations. The difficulty is that each of the three design ideas must be protected at every station, for every candidate, on every circuit. Lose standardization and two candidates sit different exams. Lose sampling and one lucky case decides a result. Lose structure, and the score reflects the examiner more than the candidate.
Understanding why each idea is there makes the protection easier. It also shows where an OSCE stops. The exam observes performance under controlled conditions. It does not observe practice.
What problem was the OSCE built to solve?
Before 1975, clinical competence was assessed with the long case and the short case. A candidate saw one or two real patients, then discussed them with one or two examiners. Harden and colleagues at Dundee identified the problem in their original paper. The patient varied, the examiner varied, and the questions varied, so the result depended heavily on luck. A candidate's score said as much about which patient and which examiner they drew as about their ability.
The OSCE answered each source of luck with a design rule. The same patients and tasks for everyone removed the patient lottery. Many short stations replaced the single case, so no one patient could decide the result. A structured scoring sheet, marked by an examiner who observed the performance, replaced the open discussion. In the original format, a procedure station was paired with a follow-up station where the candidate interpreted what they had found.
Harden revisited the design forty years later and placed it in context. The OSCE is one instrument in an assessment program, not the whole program. That framing matters for what follows: the design solves the reliability problem in the long case and creates trade-offs of its own.
What does standardization buy, and what does it cost?
Standardization means every candidate meets the same station under the same conditions. The task is the same. The simulated patient plays the same case the same way. The equipment, the time, and the instructions are identical. The examiner scores against the same scheme.
The purpose is fairness and comparability. Two candidates' scores on a station can be compared because they attempted the same thing. A pass mark can be set for the station because the difficulty is fixed. An appeal can be answered because the record shows what was asked of everyone.
The cost is realism. A real patient does not present the same way twice. A real consultation is not eight minutes long with a bell at the end. Standardizing a task narrows it, and a narrowed task can reward the candidate who has learned the format rather than the skill. The design section of this guide covers how to write stations that remain realistic within the constraints.
Why does an OSCE need so many stations?
Sampling is the second idea, and it is the one most often under-delivered. A candidate who performs well on a chest examination is not reliably good at a knee examination or at explaining a diagnosis. Performance is specific to the case. The research literature refers to this as content specificity, and it is the main reason a single long case was unreliable.
The circuit answers it by sampling many cases. Each station is one sample of the candidate's competence. The more samples, the more the total score reflects the candidate rather than the cases. Reliability rises with the number of stations for exactly this reason, and it rises faster with more short stations than with fewer long ones. The figures, including the 2020 Ottawa floor of 12 stations and 150 minutes of testing time, are on the validity and reliability page.
Sampling has a second dimension. The stations must cover the domain the exam claims to assess, not just be numerous. That is the blueprint's job: a grid of competencies against clinical problems, with each station placed in a cell. A circuit of fourteen history-taking stations is well sampled in one sense and badly sampled in another.
How does structured observation turn performance into a score?
The third idea is that a trained observer scores what they see against an agreed scheme. Three parts make it work.
The examiner observes. At most stations, the examiner observes the entire encounter and does not interact with the candidate. The candidate's performance, not their account of it, is what gets scored. That is the "shows how" level of Miller's pyramid, and the AMEE guide on the OSCE's theoretical basis places the exam at that level. Written papers assess "knows" and "knows how". Workplace assessment reaches for "does".
The scheme structures the judgment. A checklist lists the actions that should happen. A global rating asks the examiner to provide an overall judgment of performance, usually on a scale with descriptors. Many stations use both. The choice between them and the evidence for each are covered in the design section. The point here is that either one gives the examiner a shared frame, so that two examiners looking at the same performance are more likely to agree.
Scores are aggregated across stations. The exam result is not any one station's score. It is the combination of all of them, weighted or unweighted, compared with a standard set by a defensible method. Aggregation is what turns many small samples into one reliable measurement. It also means a single poor station rarely fails a good candidate, and a single strong station rarely rescues a weak one.
The anatomy of a station
Every station, whatever it tests, has the same parts. Knowing them makes station writing and quality review faster.
Part | What it contains | Who uses it |
|---|---|---|
Candidate instructions | The scenario, the setting, what has already happened, the task, and what not to do | The candidate, read in the minute before entering |
Examiner information | The purpose of the station, what the examiner may say, and the scoring scheme | The examiner, before and during the station |
Simulated patient brief | Identity, history, presenting complaint, what to volunteer and what to hold back, standard answers | The simulated patient, in training and rehearsal |
Equipment and set-up | The room layout, the equipment, the manikin or props, the spare of each | The site team |
Scoring scheme | The checklist items, the global rating scale, and any critical action | The examiner; the analyst afterward |
Timing | Reading time, station time, and any changeover allowance | The floor manager and the bell |
A station missing any of these parts fails standardization. The most common gap is the simulated patient brief. A patient who improvises changes the station for everyone who follows.
What an OSCE cannot do
The design has limits, and a lead should state them rather than discover them in an appeal.
- It observes a performance, not a practice. An OSCE shows what a candidate can do when asked and watched. It cannot show what they choose to do on a ward at three in the morning.
- It narrows tasks to fit a station. Some competencies do not fit eight minutes. Longitudinal judgment, teamwork over a shift, and follow-up decisions need other instruments.
- It is expensive in people. Every station needs an examiner, and many need a simulated patient. The cost is what drives standardization, and there is no cheap version that maintains it.
- It can be gamed by format. A candidate who has learned to perform for checklists may score well without understanding. Global ratings and well-written stations reduce this, but do not remove it.
- One exam is one sample. However many stations it has, an OSCE is a single sitting. The program's judgment should rest on more than one instrument, as Harden's retrospective argued.
None of these limits argues against the OSCE. They argue for using it for what it does well, which is a fair,wide, observed sample of clinical performance, and for pairing it with other assessments for the rest.
Test yourself
Frequently asked questions
What does OSCE stand for?
OSCE stands for Objective Structured Clinical Examination. Objective refers to the shared scoring scheme, structured to a fixed sequence of stations and tasks, and clinical to the skills assessed. The term was introduced by Harden and colleagues in 1975 in a paper describing the format at the University of Dundee.
How is an OSCE different from a long case?
In a long case a candidate saw one real patient and discussed the case with examiners, so the result depended on which patient and examiner they drew. In an OSCE every candidate rotates through the same set of short, timed stations and is scored on a shared scheme. The OSCE trades depth on one case for breadth and comparability.
Why is an OSCE more reliable than a single clinical case?
Because clinical performance is case-specific. A candidate who is strong on one case is not reliably strong on the next, so any single case is a noisy sample. Many stations average that noise out. Reliability rises with the number of stations, which is why guidelines set floors on station count and total testing time.
Does an OSCE measure what a clinician does in practice?
It measures what a clinician can do under observation, which Miller's pyramid calls "shows how". That is a strong predictor but not the same as unsupervised practice. Programs pair the OSCE with workplace-based assessments that observe real practice, and with written papers for knowledge.
Can an OSCE be run with real patients?
Yes, and some programs use real patients with stable clinical signs for examination stations. Simulated patients are preferred for history and communication stations because they can be trained to present the same case the same way to every candidate. Manikins and part-task trainers cover procedures that cannot be performed on a person.