Skip to content

Piloting an OSCE on StudyDrome: what a pilot covers and what you get back

Last updated · 5 min read

A pilot is one small circuit run end to end on real data: your station, your examiners, your candidates. It exists because a demonstration answers the vendor's questions and a pilot answers yours. Three stations and eight candidates produce every report a twelve-station exam produces, at a scale where you can check each one by hand.

This page describes the scope, the timeline, and the exports you keep. The wider product picture is on OSCE exams in StudyDrome, and the reports themselves on which report answers your question.

What a pilot covers

It walks the whole cycle once. Six steps, in order.

Getting a station in. You bring a station file. It imports, or the checklist builder produces one, or it gets typed in. Sections, items, the scoring scale, a critical item with its threshold, and a global rating all arrive together. Blueprinting and station authoring cover the model.

Building the circuit. The stations form a circuit with rest stations where you want them, and the rotation timetable is issued to candidates, examiners and rooms.

Printing the pack. The exam-day documents are printed from the same data: run sheet, rotation, signage, and marking sheets. Running exam day covers what the pack contains and what the live view shows.

Marking. Examiners score one station on one screen, with comments alongside the item scores, and the app keeps working through a connection drop. The examiner app is the page for this step.

Setting the standard. Four methods run on the marks you just collected, with the per-station controls and the regression output visible rather than hidden behind a single number. See setting the standard.

Reading the results. Item analysis, examiner reports, and a candidate record are released the same day, and the release controls determine who sees what and when. See OSCE analytics and releasing results and feedback.

What it asks from you

What

How much

Why

Stations

Three, one of them with a critical item

Three is enough to produce a circuit and to show a station comparison

Candidates

Eight to twelve

Below eight, the statistics stop computing, which is itself worth seeing

Examiners

Two, ideally marking the same station

Two examiners on one station is what makes the leniency report possible

Time

One afternoon for the circuit, two meetings around it

The setup meeting and the results meeting are where the decisions happen

A pilot works just as well on a formative sitting you were running anyway. Most schools prefer that: the exam is real, the stakes are not, and nobody has to explain a new system to candidates whose progression depends on it.

What you get back

Everything the pilot produces is yours to keep, in a spreadsheet and in PDF.

The station item analysis includes facility, discrimination, and the corrected item-total correlation for each item, plus reliability for each station and the exam. The examiner activity report covers stations staffed, candidates marked, and the hours each session spanned. The examiner leniency and agreement figures for cases where two examiners shared a station. A per-candidate record showing each item score, the global rating, the comments, and the pass/fail status, with any critical failures marked. And the standard-setting output: the pass mark under each method, with the fitted line for borderline regression rather than only the number it produced.

Those exports are the evidence file a quality review asks for. Evidence for accreditors describes what a reviewer looks for in them.

The timeline

When

What happens

Week 1

Setup meeting. Accounts, the station file, the scale, and the critical rule.

Week 2

The circuit is built, the pack is printed, and examiners get fifteen minutes on the app.

Week 3, exam day

The circuit runs. Marking closes the same afternoon.

Week 3, next day

Standard setting is run with your committee, and the reports are exported.

Week 4

Results meeting. You keep the exports and decide what a full sitting would need.

Three weeks is the comfortable version. Schools that already have a station file and a date have done it in ten days.

Frequently asked questions

Does a pilot commit us to anything?

No. It is a working session with your data, and the exports are yours at the end. Most schools use it as the last step in a tender because it addresses the requirements list with evidence rather than a vendor's response.

Can we pilot with a summative exam?

You can, and some schools do once their regulations name the method. The safer first run is formative. The circuit is real, the marking is real, and a surprise in the results costs nobody a progression decision.

What if we have no station file yet?

Then the pilot starts a step earlier. Bring the skill you want to assess and the people who teach it. The station template and the checklist builder turn an afternoon of their time into an importable station.

Do our examiners need training on the app?

Fifteen minutes, on the day or the day before. One station on one screen is the whole interface. The training that matters is the assessment training covered in examiner training, which is about the scale, not the software.

How small can the cohort be before the reports stop working?

The item analysis needs at least three candidates for discrimination, and the corrected correlation and reliability need at least two complete cases. Below those numbers, the reports say so rather than printing a figure. Eight candidates give you something to read.

Was this page helpful?