Skip to content

Exam Manager for medical and health-sciences schools

Written finals and OSCE circuits, one programme

StudyDrome Exam Manager runs written finals and OSCE circuits for the same cohort using a single item bank. A station is a question type, so one blueprint covers both formats. Borderline regression maps OSCE pass marks to the cited thresholds. One statistics engine reports on every exam, regardless of the format that produced it.

The pilot is free and time-boxed. Keep it if it works for your team; walk away if it doesn't.

How does one bank serve both exam types?

An OSCE station is a question type. It lives in the same bank as an MCQ. It carries the same typed tags and the same clinical skills. That is what allows a single blueprint to cover a whole program, rather than a single exam format.

The blueprint matrix is a two-axis coverage grid. Each axis is set on its own, to a tag dimension or to Skills.

Each cell lists the questions at that intersection. Per question, it carries the observed facility and the administration count. The matrix runs across an entire bank or a single OSCE exam.

Empty intersections are omitted to keep the grid readable at scale.

It exports to a spreadsheet. Each cell takes the color of its facility band. That is a coverage picture that survives being emailed to a committee.

How are stations designed?

A station is a question whose body is an ordered list of items. Each row is a scorable checklist item or a section header.

  • Scoring scales:Attach one scale to the whole station, or a different scale per item. Item analysis resolves the item's own scale first, then the station's.

  • Critical items:Any item can be marked critical with its own failure threshold. The examiner sees the rule stated plainly: "Critical: Fail if ≤ 3". One shared evaluator decides critical failure, so candidate-centric and station-centric reports can never disagree.

  • Rest stations:A break with no question attached. It occupies a rotation position, does not collect scores, and is permanently excluded from the Global Rating.

  • Linked stations:Chain stations into a group. A candidate completes them back to back, in an explicit order, with no rotation break.

  • Examiner roles:Examiner, Simulated Patient, and Observer are assigned per station per event.

What does exam day look like?

  1. 1

    The circuit is derived, not typed. Candidates rotate forward one station per rotation; examiners stay put. A slot is reading time plus station duration, and the waiting-room gap sits only between slots.

  2. 2

    Seven documents come out of the published exam, including station lists, doorway sheets, examiner checklists, and the rotation timetable.

  3. 3

    Examiners score on a dedicated screen with no menu and no chrome. They work on a candidate-by-candidate basis, with explicit Previous and Next controls.

  4. 4

    Marking survives a dropped connection. Scoring actions queue in browser storage and replay when the network returns. The queue holds 500 operations for 24 hours.

  5. 5

    Finishing is blocked while anything is still pending at three separate call sites. A circuit is never closed on marks held on one laptop.

How is the pass mark set?

Four methods, chosen per exam: Custom, Angoff, Borderline Group, and Borderline Regression.

Borderline regression fits checklist scores against examiner Global Ratings per station using ordinary least squares. It needs at least 3 valid candidates. It returns slope, intercept, and R² to four decimal places, then clamps the derived mark into the achievable range. The exam pass mark is the sum of the station pass marks.

Four warnings surface during calculation, so whoever is about to publish a cut score sees why it might not hold:

Warning

What it means

Same rating for all

Every candidate received the same Global Rating, so the regression is undefined. Borderline Group is the method to use.

Negative slope

Checklist scores fall as Global Ratings rise. Review the station and the examiner.

Low R²

Below 0.5, checklist and examiner judgment agree poorly.

Small cohort

Fewer than 20 completed candidates, so the result may be unstable.

The R² threshold follows Pell and Roberts, AMEE Guide 49, which treats a correlation above 0.5 as a reasonable relationship between checklist and examiner judgment. The cohort figure is a conservative floor: the resampling literature puts serious instability higher, and the spread of the global grades predicts it better than the headcount. Our guide to borderline regression sample size sets out the evidence.

The warnings are advisory. The exam board decides, and the record shows what it was told.

Every station carries its own controls: lock the mark against recalculation, mark it must-pass, or withdraw it from the calculation. An exam can also require a minimum number of stations passed.

What evidence comes out afterward?

Both exam types feed one statistics engine.

For written papers, reliability is chosen for you. KR-20 runs when every scorable item is multiple choice or true/false, and no item type allows partial credit. Cronbach's alpha runs otherwise. The report names the method it used, with the standard error of measurement beside it.

For OSCE stations, item analysis uses top and bottom 33% groups, not the 27% used for MCQs. Facility, polytomous discrimination, and item-rest correlation are computed per item. Station reliability is Cronbach's alpha across items; exam reliability is alpha across station totals.

Item-level quality flags run on both. Negative discrimination is flagged Critical, with a plain recommendation attached. Items with fewer than 10 responses are marked "Insufficient Data" rather than rated. Below that, the top and bottom groups are one or two people.

Read the complete guide to item analysis for the method. Read OSCE standard-setting methods compared for the choice between the four.

Who is this for?

A medical, dental, nursing, or allied-health school. One that runs written finals and OSCE circuits for the same cohort, in the same year, against the same competency framework.

The practical test is whether your blueprint has to span both formats. Does a station and an MCQ map to the same clinical skill? Do you want a single coverage view for both? That is the case this is built for.

Programs running only written exams are served by written exams. Programs running only OSCEs are served by OSCE exams.

Frequently asked questions

Can one blueprint cover both written exams and OSCEs?

Yes. An OSCE station is a question type. Stations and MCQs sit in the same bank, under the same tags and clinical skills. The blueprint matrix maps two axes over that bank, or over one exam. Each cell carries the observed facility and administration count for its questions. It exports to a color-banded spreadsheet.

How does StudyDrome derive an OSCE pass mark?

By one of four methods, chosen per exam: Custom, Angoff, Borderline Group, or Borderline Regression. Borderline regression fits checklist scores against examiner Global Ratings, per station, using ordinary least squares. The pass mark is the fitted value at the borderline grade. The exam pass mark is the sum of the station marks.

What happens if an examiner's tablet loses connection mid-station?

Scoring actions are queued in browser storage, deduplicated, and replayed automatically when the network returns. The queue holds 500 operations and expires anything older than 24 hours. A banner shows the pending count, and finishing the exam is blocked while operations remain unsynced.

What happens when a cohort is small?

At least 3 candidates with valid Global Ratings, or the regression will not run. Below 20 completed candidates, the result is flagged as potentially unstable. The warning is advisory, so the exam board keeps the decision. A flag-free small cohort is not a clean bill of health: check the spread of the global grades and the R² for each station.

Can a candidate fail one station and still pass the OSCE?

That depends on the rules you set. A critical failure at that station fails the exam only if the station is also marked as must-pass. Separately, an exam may require a minimum number of stations to be passed. Total score, minimum stations, and must-pass conditions are evaluated together.

How does reliability differ between written exams and OSCEs?

Written exams get KR-20 or Cronbach's alpha, selected automatically by item type, with the standard error of measurement beside it. OSCE stations get Cronbach's alpha across items, and the exam gets alpha across station totals. OSCE item analysis also uses top and bottom 33% groups, whereas written exams use 27%.

Book a pilot

Bring one station set and one past paper. We will blueprint both against your competency framework and show you the statistics they produce.