A sequential OSCE tests everyone on a short screening circuit. It then tests again only the candidates who did not clearly pass it. The decision for those tested twice rests on the full sequence, not on the second test alone. It buys more evidence about the candidates the exam is least sure of, and spends examiner time only on them. The design questions are how long the screen is, where its cut sits, and how the two parts combine.
The variants page covers what a sequential OSCE is and how it compares with other formats. This page is about building one.
Why the format exists
A pass or fail decision is least reliable near the cut score. That is where the exam's error matters most and where an extra station can change the result. Adding stations for everyone fixes it, but it costs a great deal. A sequential design puts the extra stations only where the uncertainty is.
The published work behind it comes from a final-year medical OSCE at Leeds. Its authors describe the aim plainly. For those tested twice, the pass-or-fail decision rests on the full sequence of tests.
How long should the screen be?
Long enough to sort candidates reliably, short enough to save something. Work at Aberdeen answered this by simulation. Two cohorts of final-year students, 388 in all, were resampled 5,000 times. Screen lengths ran from 6 to 14 stations. Each outcome was compared with the decision from the full 15-station exam.
Across that range, sensitivity stayed at or above 87 percent in one year and 84 percent in the other. Specificity ran from 60 to 100 percent, and rose as the screen got longer. Eight stations came out as a reasonable compromise. Sensitivity was 88 to 89 percent, and specificity was 83 to 86 percent.
Two cautions come with those numbers. They are a simulation on one school's data, so the shape of the answer travels further than the value does. And the screen still has to sample the blueprint. A screen of eight stations that leaves out a whole domain is not a screen. It is a smaller exam about something else.
Where does the screening cut sit?
Above the final standard, not at it. The screen is a filter. It should send through anyone who might fail, and it will also send some who would pass.
The Aberdeen study set that margin in standard errors of measurement above the cut. Results were similar at one, two, and three of them. In practice, the margin decides your workload. A wider margin recalls more candidates, costs more, and misses fewer. Choose it on purpose. Record the choice in the decision record and publish it with the other rules.
What did it deliver?
At Leeds, in a cohort of 228 students, 13.5 percent were required to sit the sequential OSCE. The full sequence gave an overall reliability estimate of 0.79.
Two result groups are worth naming. One student, 0.4 percent of the cohort, would have passed under the previous system but failed on the extended testing. Nine students, 3.9 percent, would have needed a retest under the old system, passed the full sequence, and graduated at the normal time. The authors report significant savings in delivery costs.
Read those two numbers together. The design is not mainly a way to fail more people. It is a way to stop sending borderline candidates into a retest cycle when more evidence shows they are competent.
When does it pay?
Condition | Why it matters |
|---|---|
A large cohort | Savings scale with the candidates who sit only the screen |
A deep station bank | The second test needs stations of the same quality, not leftovers |
Stable station quality | The authors are explicit that station-level quality must be high enough for the reliability gain to appear |
Space in the calendar | The second test has to run before the board meets |
Staff who can be recalled | A second circuit needs examiners and simulated patients twice |
If any of the middle three fails, the format costs more than it saves. A weak second circuit produces a less reliable decision about precisely the candidates who most need a sound one.
What the standard-setting record must add
A sequential design adds four lines to the record that a single-sitting exam does not need. What the screening cut is, and how far above the standard it sits. Which candidates are recalled. How the two tests combine into one score and one decision. And what the extra rules do across a sequence. A minimum number of stations passed has to be defined over the full set a candidate sat.
Candidates need the same four facts before they sit. Being recalled is stressful. A cohort told in advance that a recall is not a failure takes it better than one told afterward.
Limits
The evidence base is small and comes from a handful of schools, mostly in the United Kingdom. The reliability gain depends on station quality, which is a claim about your exam rather than about the format. And the second test flags a group of candidates as at risk. The timetable should soften that rather than advertise it.
Frequently asked questions
How does a sequential OSCE decide who passes?
Everyone sits a screening circuit. Candidates who clearly clear it are done. The rest sit a second circuit, and their pass-or-fail decision is based on the full sequence of tests, not on the second one alone.
How many stations should a sequential OSCE screening test have?
A simulation study at Aberdeen compared screens of 6 to 14 stations against a 15-station exam and found 8 to be a reasonable compromise, with sensitivity of 88 to 89 percent and specificity of 83 to 86 percent. The screen must still cover the blueprint.
Where should the screening pass mark be set?
Above the final standard, so that anyone who might fail is sent through to the second test. Aberdeen set the margin at one, two and three standard errors of measurement and found similar results. A wider margin recalls more candidates and misses fewer.
How many candidates sit the second test?
At Leeds, 13.5 percent of a cohort of 228. The share depends on where the screening cut is set and on how the cohort performs, so it should be modeled on previous sittings before the timetable is fixed.
Does a sequential OSCE fail more candidates?
Not mainly. In the Leeds cohort one student, 0.4 percent, would have passed under the old system and failed on extended testing. Nine students, 3.9 percent, would have needed a retest under the old system and instead passed the full sequence.