Every OSCE station is reviewed twice before it counts. A peer-review workshop reads the draft against the blueprint cell, the marking scheme, and the time, and sends it back or forward. A pilot then runs the station with real candidates, examiners, and a simulated patient, identified to candidates but not named, and the station's scores are analyzed before the station enters the bank.
Why review is a stage, not a courtesy
The AMEE guide on OSCE organization draws the station bank as a flow: topics, writers, station types, template, writing, marking guidance, peer-review workshops, piloting, psychometric analysis, bank. Review and piloting are two of its ten steps. The guide also lists peer review of stations among the six elements of an OSCE quality assurance system, alongside standardization, external examiners, psychometrics, evaluation, and examiner training.
A South African study of OSCE planning reached the same place from the other direction. Its seven recommendations for strengthening OSCE design include a station bank built with peer-review workshops and a separate step for station piloting. Both sources treat a station that has not been reviewed as a draft, regardless of what its author thinks.
Who sits in the review workshop?
A workshop of four to six people, for a batch of stations, half a day. The seats:
- Two clinicians at the candidate's level of practice who did not write the station. They judge whether the task is real and the time is fair.
- An assessment lead or someone who knows the blueprint. They check the cell.
- A simulated patient trainer. They read the script for playability and standardization.
- An examiner trainer or experienced examiner. They apply the marking scheme to an imagined performance and identify the items that cannot be observed.
- The exam office. They read the equipment list and the room requirements.
The author attends to answer questions and to take the notes. The author does not defend the station. The workshop's output is a decision for each station: accept, accept with listed changes, or return.
What does the workshop check?
The AMEE guide's second appendix is a station QA questionnaire with feasibility items and validity items. The checklist below follows that division. It is the form the workshop completes for each station.
Twelve items are enough. A longer form turns the workshop into a clerical exercise. The judgment the workshop is there for is in the first four items: the right cell, a feasible task, honest instructions, and an observable scheme.
What sends a station back?
The workshop returns a station when a fault cannot be fixed at the table. The common ones:
- Wrong cell. The scheme scores examination items in a history station. The blueprint claims coverage the station does not deliver.
- Impossible time. Two tasks in one station. The fix is a split, which means two stations and a second review.
- Unobservable items. "Demonstrates empathy" is a judgment. It becomes a global-rating anchor or an observable behavior, or it goes.
- A script with holes. The scheme rewards a question the script does not answer.
- Unplayable role. A simulated patient cannot present the sign or present it 40 times.
Accept-with-changes is for wording, quantities, and layout. Anything that changes what the station measures goes back to the author and returns to the next workshop.
What are the pilot rules?
A pilot puts the reviewed station in front of candidates under exam conditions. The AMEE guide states the two rules. Pilot stations placed inside a live circuit are declared to candidates, and they are not identified. Candidates know that one or more stations do not count. They do not know which.
Four more rules follow from those two.
- A pilot station does not count toward the candidate's result. Its scores are analyzed, not aggregated. The candidate handbook says so.
- A pilot runs in a real slot with a real examiner and a real simulated patient. A rehearsal with faculty playing candidates is useful but is not a pilot.
- The pilot's data are read before the station enters the bank. Facility, discrimination, the checklist-global agreement, and examiner comments. The quality section sets out the metrics and the thresholds.
- A pilot that fails is revised and piloted again. It does not enter the bank on the strength of a rewrite.
The 2018 "twelve tips" paper on OSCE development frames the whole sequence as a validity argument built in development order, from the decision the exam supports back to the station. Piloting is where the argument meets data for the first time.
What happens after the pilot?
The pilot analysis produces one of three outcomes. The station enters the bank with its metrics and its version recorded. The station is revised, with the change logged, and piloted again. Or the station is retired before it ever counts, which is a good outcome for a station that would have unfairly failed candidates.
Metrics from a single pilot are noisy. One circuit of 20 candidates is a small sample, and the reliability page explains why. The pilot answers the coarse questions: did the task work, did the examiners agree on what they saw, did the simulated patient hold the role. The fine questions await the station's first live run and the bank's record of it over time.
Limits
Review and piloting cost time that the planning timeline must accommodate: a workshop months before the exam, a pilot slot in a live circuit or a formative sitting, and analysis in between. A school that writes stations in the last six weeks skips both. Its stations are drafts on the day, and the results show it.
Frequently asked questions
What is a peer-review workshop for OSCE stations?
A half-day meeting where a small group reads a batch of draft stations against the blueprint, the marking scheme, the time, and the simulated patient script. Clinicians who did not write the station, an assessment lead, a simulated patient trainer and an examiner trainer sit in. Each station is accepted, accepted with changes or returned.
Should pilot OSCE stations count toward the result?
No. A pilot station runs in a live circuit, so its data are real, but its scores are analyzed rather than added to the candidate's total. The AMEE organization guide states that pilot stations are declared to candidates and not identified. Candidates know a station does not count, but not which one.
How many candidates does an OSCE station pilot need?
One circuit is usually what you get, so the pilot answers coarse questions: did the task fit the time, did examiners agree, did the role hold. Item statistics from 20 candidates are noisy. The fine metrics come from the station's first live run and from its record in the bank over several sittings.
What does a station QA questionnaire cover?
Feasibility and validity. Feasibility: can the task be done in the time, can the role be played, is the equipment available, are the instructions complete. Validity: does the station fill its blueprint cell, are the items observable, does the scheme measure the intended competency? The AMEE guide's second appendix is the model.
Who has the final say on a reviewed station?
The workshop decides to accept, accept with changes, or return. The author makes the changes but does not overrule the decision. The assessment committee approves the set of stations for the sitting along with the blueprint, and the bank owner records the version that was approved.