OSCE exams in StudyDrome

Station-based clinical examination, end to end

StudyDrome Exam Manager runs station-based clinical examinations end to end. It builds the circuit and generates seven exam-day documents. Examiners score stations in a dedicated app. Four standard-setting methods derive the pass mark. Borderline regression returns slope, intercept and R², plus four named warnings when the data will not carry a cut score.

How does StudyDrome model an OSCE circuit?

Candidates rotate forward one station per rotation. Examiners stay put. That is the whole model, and it is a tested, pure function rather than a spreadsheet.

Each slot is reading time plus station duration. The waiting-room gap sits only between slots, so the final slot ends exactly at the event window end.

Setting

Range

Station duration

1 to 99,999 minutes, excluding reading time

Candidate reading time

0 to 99,999 minutes, optional

Station waiting room

0 to 999 seconds, between slots

A circuit slot need not be scored. A rest station is a break with no question attached. It occupies a rotation position, collects no scores, and is permanently excluded from the Global Rating.

Stations can be chained into a linked group. A candidate completes them back to back, in an explicit order, with no rotation break between.

How do you blueprint an OSCE?

The blueprint matrix is a two-axis coverage grid. Each axis is set independently to a tag dimension, to Skills, or to nothing.

  • Sparse by design: Every cell lists the questions at that intersection. Empty intersections are omitted, so the grid stays readable.
  • History per question: Each question carries the observed facility across all past administrations and how many times it has been administered.
  • Bank or exam: The matrix covers either a whole question bank or one OSCE exam.
  • Excel export: Each cell takes the colour of its facility band, so coverage survives being emailed to a committee.
  • A coverage instrument: The blueprint is evidence of what your circuit assesses. You build the exam; the matrix shows you the gaps and the concentrations.

How are stations designed and scored?

A station is a question whose body is an ordered list of items. Each row is a scorable checklist item or a section header.

Scoring scales attach either to the whole station or per item. Any item can be marked critical, with its own failure threshold. The examiner sees the rule stated plainly: "Critical: Fail if ≤ 3".

One shared evaluator decides critical failure, so the candidate-centric and station-centric reports can never disagree. The comparison is inclusive: a score at or below the threshold fails the station. An unassessed critical item does not trigger a failure.

Critical failure fails that station. It only fails the exam if the station is also marked must-pass.

What does the examiner use on exam day?

A dedicated screen, opened for one station at one event, with no left menu and no chrome. Nothing competes with the checklist during a five-minute station.

  • The examiner works candidate by candidate with explicit Previous and Next controls.
  • The screen renders the checklist, the configured scoring scale, the critical-item thresholds, a Global Rating and a comment box. A minimum comment length can be set per exam.
  • Three examiner roles exist: Examiner, Simulated Patient and Observer.
  • A scoring action that cannot reach the server goes into a browser-storage queue. The queue deduplicates it, then replays it when the network recovers.
  • The queue holds up to 500 operations and drops anything older than 24 hours. A banner shows the count.
  • Finishing is blocked while anything is still pending, at three separate call sites. A circuit cannot be closed on scores that exist only on one laptop.
  • It is a queue-and-replay layer. The exam loads over the network at the start of the session. Marking then survives a dropped connection.

Which exam-day documents does StudyDrome generate?

Seven documents come out of the published exam, including the station list, doorway sheets, examiner checklists and the rotation timetable.

The station list reports total stations, active stations, rest stations and linked-group count. Rest slots render as "Rest Station", typed "Rest" rather than "Scored".

The rotation timetable is a rotation-by-station grid with slot times per event, derived from the timing model above.

How does StudyDrome set the pass mark?

Four methods, chosen per exam.

Method

How the station mark is derived

Custom

A typed preset pass mark.

Angoff

Expert-judgment pass marks per station. Exam mark is the sum or the average, defaulting to average.

Borderline Group

Mean or median checklist score of candidates rated at the borderline grade. Defaults to mean.

Borderline Regression

Checklist score regressed on Global Rating. Pass mark is the fitted value at the borderline grade.

Borderline regression is ordinary least squares, fitted per station. It needs at least 3 valid candidates. It returns slope, intercept and R² to four decimal places, then clamps the derived mark into the achievable range. The exam pass mark is the sum of station marks.

For the methods themselves, read OSCE standard-setting methods compared.

Which warnings stop you publishing a cut score?

Four warnings are raised during calculation.

  • Same rating for all: Every candidate got the same Global Rating, so the regression is undefined. Use Borderline Group instead.
  • Negative slope: Checklist scores fall as Global Ratings rise. Review the station and the examiner.
  • Low R²: Below 0.5, checklist and examiner judgment agree poorly.
  • Small cohort: Fewer than 20 completed candidates, so the result may be unstable.

Those two thresholds are not invented. The R² threshold cites Pell and Roberts, AMEE Guide 49. The cohort threshold cites Homer et al. 2015 and Moreno-López et al. 2022.

The warnings are advisory. None blocks calculation, publication or result release. That is a deliberate choice: the panel decides, not the software.

Every station carries its own controls. Lock the mark against recalculation, mark it must-pass, or withdraw it from the calculation entirely. An exam can also require a minimum number of stations passed. Total score, minimum stations and must-pass are all ANDed.

Station-level item analysis uses top and bottom 33% groups, not the 27% used for MCQs. Facility index, item-rest correlation and SEM are computed per station, and SEM again on candidate totals. Examiner agreement is reported as a mean absolute percentage difference. It is descriptive and unadjusted, and the product says so.

How is the circuit model scoped?

  • One circuit, rotating forward. Candidates move one station per rotation; examiners stay at their station.
  • The examiner app is the responsive web app, so an examiner uses the tablet or laptop they already have.
  • Angoff pass marks are entered per station by whoever holds the judgment. The exam mark is their sum or average.
  • Examiner agreement is a mean absolute percentage difference from the station mean. The product labels it as descriptive.
  • Borderline regression is ordinary least squares on one predictor. The output is a slope, an intercept and an R² you can read directly.
  • Stations use checklists and scoring scales, with per-item or per-station scales and critical items carrying their own threshold.

OSCE specification

Area

Detail

Circuit

Forward rotation, one station per rotation. Examiners fixed to a station.

Timing

Station 1–99,999 min. Reading 0–99,999 min. Waiting room 0–999 s.

Stations

Checklist items and section headers. Per-station or per-item scoring scales.

Critical items

Per-item threshold. Inclusive comparison. Fails the station, not the exam.

Rest stations

No question attached. No scores. Always excluded from Global Rating.

Linked stations

Chained group with explicit order, completed without a rotation break.

Blueprint

Two axes over a bank or one exam. Colour-banded Excel export.

Examiner roles

Examiner, Simulated Patient, Observer.

Offline

500-operation queue, 24-hour expiry, auto-replay, finish blocked while pending.

Standard setting

Custom, Angoff, Borderline Group, Borderline Regression.

Pass rules

Exam pass mark, minimum stations passed, must-pass stations. All ANDed.

Statistics

Facility index, 33% discrimination, item-rest correlation, SEM, OLS with R².

Frequently asked questions

Does StudyDrome support borderline regression standard setting?

Yes. Checklist scores are regressed on examiner Global Ratings using ordinary least squares, fitted per station. The station pass mark is the fitted value at the borderline grade, clamped into the achievable range. Slope, intercept and R² are returned to four decimal places. The exam pass mark is the sum of the station pass marks.

How many candidates do we need for a defensible cut score?

StudyDrome flags any borderline calculation with fewer than 20 completed candidates as potentially unstable, citing Homer et al. 2015 and Moreno-López et al. 2022. Below 3 valid candidates the regression will not run at all. The warning is advisory rather than blocking, so the exam board keeps the decision.

What happens if an examiner's tablet loses connection?

Scoring actions are queued in browser storage, deduplicated, and replayed automatically when the network returns. The queue holds 500 operations and expires anything older than 24 hours. A banner shows the pending count. Finishing is blocked while operations remain unsynced, so a circuit is never closed on marks held on one laptop.

Can a candidate fail one station and still pass the OSCE?

That depends on your rules. A critical failure fails that station, not the exam. It only fails the exam if the station is also marked must-pass. Separately, an exam can require a minimum number of stations passed. Total score, minimum stations and must-pass conditions are evaluated together.

Which documents does StudyDrome produce for exam day?

Seven, generated from the published exam. They include the station list, doorway sheets, examiner checklists and the rotation timetable. The station list reports total, active, rest and linked-group counts. The rotation timetable is a rotation-by-station grid with slot times per event, derived from the configured timing model.

Can we blueprint stations against our competency framework?

Yes. The blueprint matrix maps two axes, each set to a tag dimension or to clinical Skills. Cells carry observed facility and administration counts per question. It covers a whole bank or one exam. It exports to a colour-banded spreadsheet you can send to a committee as evidence.

Book a pilot