A large language model can draft the parts of an OSCE station: the scenario, the candidate instructions, a first marking scheme, and a simulated patient script. It cannot validate them. A drafted station enters the same workflow as a human-written one, with two extra steps: a clinical accuracy check against named guidelines and a bias audit. The blueprint cell, the critical items, and the pass decisions stay with people.
Why the bar is the point
Station writing is the slowest part of the design cycle, and the temptation is to let a model do it. A 2026 paper in Frontiers in Medicine outlines 10 tips for generating OSCE stations using AI. Its central warning is that AI-generated cases are not self-validating. They should not be adopted without a defined validation and quality assurance process. The tips are practical because they assume the model will be used. The workflow on this page is that process, written for an assessment committee.
The evidence for AI in clinical assessment is early. A 2026 scoping review of large-language-model virtual patients identified 21 studies, mostly pilot studies, with a median of 21 participants. Nothing in that literature supports letting a model make an assessment decision. Everything in it supports using a model to produce drafts that experts then test.
What may the model draft, and what may it not decide?
Part of the station | Model may draft | People decide |
|---|---|---|
Blueprint cell, level, duration | No | The working group assigns the cell before any prompt is written |
Scenario and patient background | Yes, from a structured brief | Clinical accuracy, local epidemiology, cultural fit |
Candidate instructions | Yes | That the task matches the cell and the time |
Marking scheme items | Yes, as a first list | Which items stay, the anchors, the weights |
Critical items | No | The critical-error policy and the committee |
Examiner information | Yes | What the examiner may say and release |
Simulated patient script | Yes, including variations | Playability, standardization, what the patient must not say |
Supporting materials: results, charts, images | Yes, as mock-ups | Every value checked against a real reference range |
Case variations for the bank | Yes | That the variations still fill the same cell |
Pass decisions, scoring, feedback to candidates | No | Examiners, the standard-setting method, the board |
The rule under the table: the model produces text, and a named person is accountable for every line of it before it reaches the review workshop.
The validation workflow
Six steps, in order. A station that has not completed all six is a draft.
- Brief before prompt. Write the blueprint cell, the level, the duration, the station type, and the competency the cell samples. Name the guideline the case must follow. The Frontiers paper's first tips are to blueprint first and to design before prompting: the model fills a specification; it does not invent one.
- Structured prompt. Give the model the station template's five parts as headings, the brief, and the local context: the health system, the drug names in use, the population the candidates will treat. Ask for the parts separately, not the whole station in one pass.
- Clinical accuracy check. A clinician at the candidate's level reads every clinical statement against the named guideline. Every number in a results sheet is checked against a reference range. Anything the model invented that the brief did not ask for is removed.
- Bias and fairness audit. A second reader checks names, ages, occupations, the framing of the presenting complaint and the assumptions in the script. They look for stereotypes, and for a case that reads differently to candidates from different backgrounds. The audit is recorded with the station's version.
- The standard review workshop. The station goes to the peer-review workshop as any station does, with the two audits attached. The workshop is told the station was AI-drafted, so it reads with the right suspicion.
- Pilot and monitor. The station is piloted under the pilot rules and its metrics are read. The Frontiers paper's last tip is psychometric monitoring; the quality section sets the thresholds. An AI-drafted station that drifts is retired like any other.
Steps 3 and 4 are the additions. Steps 1, 2, 5, and 6 are part of the normal workflow; the point is that the model changes the drafting stage and nothing after it.
What goes into the prompt, and what must not?
Into the prompt: the brief, the template headings, the guideline name, the local context, and an instruction to state uncertainty rather than fill gaps.
Never into the prompt: real patient data. A station drafted from a real case is a data protection event before it is an assessment event, and the consent and data page explains why. Use invented patients built to the brief. Also, never into a public model: existing bank stations, marking schemes, or scripts. A bank station pasted into a third-party tool has been exposed, and the bank's exposure rules apply.
What does an AI-drafted station look like in the bank?
Like any other, with one more field. The record says the station was AI-drafted, which model, and who ran the accuracy check and the bias audit. That field lets the bank owner compare AI-drafted and human-written stations on their metrics after a few sittings. That comparison is the school's own evidence, and it is the only evidence that answers the question for that school's candidates and examiners.
Limits
The published guidance is expert consensus and experience, not trial evidence. The scoping review literature is small and mostly focuses on virtual patients for learning, not on station drafting for assessment. Model outputs change between versions, so a prompt that produced a good station in one month may not in the next. The workflow is designed for that uncertainty: it trusts nothing the model writes until a person has checked it, and it keeps the record that lets the school find out whether the checking worked.
Frequently asked questions
Can AI write OSCE stations?
It can draft them: the scenario, candidate instructions, a first marking scheme, and a simulated patient script, from a structured brief. It cannot validate them. A 2026 set of ten tips for AI-generated stations warns that outputs are not self-validating and must go through a defined validation and quality assurance process before use.
What should be checked in an AI-drafted OSCE station?
Clinical accuracy against a named guideline, every number against a reference range, invented details the brief did not ask for, and bias in names, ages, occupations, and framing. Then the station goes through the normal peer-review workshop and a pilot, with the two audits attached to its version record.
Should critical items be set by AI?
No. A critical item fails a station regardless of the total, and the policy behind it is approved by the assessment committee before the exam. The model can propose a list of first items. Which items are critical and at what level they fail are decisions people make and record.
Is it safe to put real patient cases into an AI tool to draft stations?
No. Real patient data in a prompt is a data protection event, and existing bank statements pasted into a public tool are exposed content. Draft from invented patients built to the blueprint brief, and keep bank material out of third-party tools unless counsel and the bank owner have approved the tool.
Is there evidence that AI-drafted stations work as well as human-written ones?
Not yet at the level a committee should rely on. The published guidance is expert consensus, and the scoping-review literature on large-language-model virtual patients is limited, consisting mostly of pilots with a median of 21 participants. Record which stations were AI-drafted and compare their metrics with the rest of the bank over several sittings.