Exam Manager for medical and health-sciences schools

Written finals and OSCE circuits, one programme

StudyDrome Exam Manager runs written finals and OSCE circuits for the same cohort, from one item bank. A station is a question type, so one blueprint covers both formats. Borderline regression derives OSCE pass marks against cited thresholds. One statistics engine reports on every exam, whichever format produced it.

How does one bank serve both exam types?

An OSCE station is a question type. It lives in the same bank as an MCQ. It carries the same typed tags and the same clinical skills. That is what lets one blueprint cover a whole programme, rather than one exam format.

The blueprint matrix is a two-axis coverage grid. Each axis is set on its own, to a tag dimension or to Skills.

Each cell lists the questions at that intersection. Per question it carries the observed facility and the administration count. The matrix runs over a whole bank, or over one OSCE exam.

Empty intersections are omitted, so the grid stays readable at scale.

It exports to a spreadsheet. Each cell takes the colour of its facility band. That is a coverage picture that survives being emailed to a committee.

How are stations designed?

A station is a question whose body is an ordered list of items. Each row is a scorable checklist item or a section header.

  • Scoring scales: Attach one scale to the whole station, or a different scale per item. Item analysis resolves the item's own scale first, then the station's.
  • Critical items: Any item can be marked critical with its own failure threshold. The examiner sees the rule stated plainly: "Critical: Fail if ≤ 3". One shared evaluator decides critical failure, so candidate-centric and station-centric reports can never disagree.
  • Rest stations: A break with no question attached. It occupies a rotation position, collects no scores, and is permanently excluded from the Global Rating.
  • Linked stations: Chain stations into a group. A candidate completes them back to back, in an explicit order, with no rotation break.
  • Examiner roles: Examiner, Simulated Patient and Observer, assigned per station per event.

What does exam day look like?

  1. 1
    The circuit is derived, not typed. Candidates rotate forward one station per rotation; examiners stay put. A slot is reading time plus station duration, and the waiting-room gap sits only between slots.
  2. 2
    Seven documents come out of the published exam, including station lists, doorway sheets, examiner checklists and the rotation timetable.
  3. 3
    Examiners score in a dedicated screen with no menu and no chrome. They work candidate by candidate, with explicit Previous and Next controls.
  4. 4
    Marking survives a dropped connection. Scoring actions queue in browser storage and replay when the network returns. The queue holds 500 operations for 24 hours.
  5. 5
    Finishing is blocked while anything is still pending, at three separate call sites. A circuit is never closed on marks held on one laptop.

How is the pass mark set?

Four methods, chosen per exam: Custom, Angoff, Borderline Group and Borderline Regression.

Borderline regression fits checklist scores against examiner Global Ratings using ordinary least squares, per station. It needs at least 3 valid candidates. It returns slope, intercept and R² to four decimal places, then clamps the derived mark into the achievable range. The exam pass mark is the sum of the station pass marks.

Four warnings surface during calculation, so whoever is about to publish a cut score sees why it might not hold:

Warning

What it means

Same rating for all

Every candidate received the same Global Rating, so the regression is undefined. Borderline Group is the method to use.

Negative slope

Checklist scores fall as Global Ratings rise. Review the station and the examiner.

Low R²

Below 0.5, checklist and examiner judgment agree poorly.

Small cohort

Fewer than 20 completed candidates, so the result may be unstable.

Both thresholds are cited, not invented. The R² threshold follows Pell and Roberts, AMEE Guide 49. The cohort threshold follows Homer and colleagues in 2015 and Moreno-López and colleagues in 2022.

The warnings are advisory. The exam board decides, and the record shows what it was told.

Every station carries its own controls: lock the mark against recalculation, mark it must-pass, or withdraw it from the calculation. An exam can also require a minimum number of stations passed.

What evidence comes out afterwards?

Both exam types feed one statistics engine.

For written papers, reliability is chosen for you. KR-20 runs when every scorable item is multiple choice or true/false, and no item type allows partial credit. Cronbach's alpha runs otherwise. The report names the method it used, with the standard error of measurement beside it.

For OSCE stations, item analysis uses top and bottom 33% groups, not the 27% used for MCQs. Facility, polytomous discrimination and item-rest correlation are computed per item. Station reliability is Cronbach's alpha across items; exam reliability is alpha across station totals.

Item-level quality flags run on both. Negative discrimination is flagged Critical, with a plain recommendation attached. Items with fewer than 10 responses are marked "Insufficient Data" rather than rated. Below that, the top and bottom groups are one or two people.

Read the complete guide to item analysis for the method. Read OSCE standard-setting methods compared for the choice between the four.

Who is this for?

A medical, dental, nursing or allied-health school. One that runs written finals and OSCE circuits for the same cohort, in the same year, against the same competency framework.

The practical test is whether your blueprint has to span both formats. Does a station and an MCQ map to the same clinical skill? Do you want one coverage view over both? That is the case this is built for.

Programmes running only written exams are served by written exams. Programmes running only OSCEs are served by OSCE exams.

Frequently asked questions

Can one blueprint cover both written exams and OSCEs?

Yes. An OSCE station is a question type. Stations and MCQs sit in the same bank, under the same tags and clinical skills. The blueprint matrix maps two axes over that bank, or over one exam. Each cell carries the observed facility and administration count for its questions. It exports to a colour-banded spreadsheet.

How does StudyDrome derive an OSCE pass mark?

By one of four methods, chosen per exam: Custom, Angoff, Borderline Group or Borderline Regression. Borderline regression fits checklist scores against examiner Global Ratings, per station, using ordinary least squares. The pass mark is the fitted value at the borderline grade. The exam pass mark is the sum of the station marks.

What happens if an examiner's tablet loses connection mid-station?

Scoring actions queue in browser storage, deduplicate, and replay automatically when the network returns. The queue holds 500 operations and expires anything older than 24 hours. A banner shows the pending count, and finishing the exam is blocked while operations remain unsynced.

How many candidates does borderline regression need?

At least 3 with valid Global Ratings, or the regression will not run. Below 20 completed candidates the result is flagged as potentially unstable. That threshold cites Homer and colleagues in 2015, and Moreno-López and colleagues in 2022. The warning is advisory, so the exam board keeps the decision.

Can a candidate fail one station and still pass the OSCE?

That depends on the rules you set. A critical failure fails that station, and it fails the exam only if the station is also marked must-pass. Separately, an exam can require a minimum number of stations passed. Total score, minimum stations and must-pass conditions are evaluated together.

How does reliability differ between written exams and OSCEs?

Written exams get KR-20 or Cronbach's alpha, selected automatically by item type, with the standard error of measurement beside it. OSCE stations get Cronbach's alpha across items, and the exam gets alpha across station totals. OSCE item analysis also uses top and bottom 33% groups, where written exams use 27%.

Book a pilot