Skip to content

Setting the standard in StudyDrome: four methods, station controls and the pass decision

Last updated · 6 min read

The scoring section asks a committee to decide four things. A method, an anchor, an aggregation rule, and a set of extra rules. Then to record all of it. StudyDrome Exam Manager covers the arithmetic part of that. It computes station pass marks by the method you choose. It applies the exam-level rules in a fixed order. And it shows the diagnostics that say whether a computed mark can be trusted.

The product page standard setting in Exam Manager covers the feature. The free standard-setting calculator does the same arithmetic on a spreadsheet of scores, without an account. This page connects the section's decisions to the screen the exam office uses.

The four methods

Standard setting is step 9 of the OSCE exam wizard. One method applies to the whole exam. All four share the same per-station statistics, the same withdrawn handling, and the same pass-or-fail recount.

Method

What it needs

How the station mark is produced

How the exam mark is formed

Custom

A typed pass mark per station

Used as typed

Sum of station marks

Angoff

A typed expert-judgment mark per station

Used as typed

Sum or Average, set by you, default Average

Borderline Group

Global ratings, and candidates at the borderline grade

Mean or median of their checklist scores, default Mean

Sum of station marks

Borderline Regression

Global ratings across all candidates

The fitted value at the borderline grade

Sum of station marks

Two boundaries are worth stating plainly. The Angoff option takes marks that a panel agreed elsewhere. It does not collect judgments from several judges and average them, so the panel meeting stays a meeting. The method list is exactly these four, and all four are for OSCEs. Any other approach your committee uses, such as a panel exercise held during a meeting, falls outside the scope of the software. Written exams here use grading scales instead.

Both borderline methods need a Global Rating Scale on the exam. Without one, the calculation stops and says why. No global ratings were collected. A scale has to be attached, and the exam republished first. The scale itself, and the wording of its borderline point, come from the marking scheme.

What the borderline regression returns

For each station, checklist scores are regressed on the examiners' global ratings. The fit is ordinary least squares. The station pass mark is the fitted value at the borderline grade. It is clamped into the achievable range and rounded to two decimals. Slope, intercept, and R squared are reported to four decimals.

Alongside them comes the block the board will want. The number of candidates, the lowest and highest score, the mean, and the standard deviation. The report page draws the scatter, with global rating on one axis and checklist score on the other. The fitted line sits over it. The whole page exports as a PDF.

It is simple regression on one predictor. There is no weighting and no multilevel model. There is no confidence band on the line and no error band on the derived mark. A calculation needs at least three candidates with a complete station and a global rating.

The four warnings

Four advisory warnings are raised during the calculation and displayed for each station. Whoever is about to publish a cut score sees why it might not hold.

  1. Same rating for all. Every candidate at that station received the same global rating, so the regression is undefined. The borderline group method is the suggested alternative.
  2. Negative slope. Checklist scores fall as global ratings rise. That points at the station or at its examiners, not at the candidates.
  3. Low R squared. The value is below 0.5, so the checklist and the examiners' holistic judgment agree poorly at that station. The threshold is read from AMEE Guide 49. It treats a correlation above 0.5 as a reasonable indication of a relationship between the two.
  4. Small cohort. Fewer than 20 candidates completed the station, so the result may be unstable. This applies to both borderline methods.

None of the four blocks a calculation, a publication, or a release. They are there for the conversation the board has, not to stop it.

One honest note on the fourth. Twenty is a conservative floor chosen by the product. No paper sets it as the point of sufficiency. The evidence says stability depends more on the spread of global grades at a station than on headcount. So the first warning matters more than the fourth. The cohort-size evidence sets out why.

The per-station controls

Every station carries its own row. It holds a borderline-grade anchor that can override the exam-level one, a pass mark, and three switches.

  • Lock keeps a mark against recalculation. A locked station with a preset mark uses that mark exactly as entered.
  • Must Pass: a candidate who scores below that station's mark fails the exam, regardless of their total.
  • Withdrawn removes the station from every exam-mark sum and from the calculation.

A fourth flag, exclude from Global Rating, lives on the station itself and is always on for rest stations. An excluded station produces no borderline pass mark, and the row says so.

Those three switches are how the section's conjunctive rules and a board's withdrawal decision look in software. Resetting all stations unlocks them and clears the marks.

How the pass decision is made

Three conditions are evaluated in a fixed order, and all must hold.

  1. The candidate's total reaches the exam pass mark.
  2. If the exam requires a minimum number of stations to be passed, the candidate meets it. A station counts toward that total only when it has a pass mark and is not withdrawn.
  3. No must-pass station was failed.

The exam pass mark is derived from the stations by the method's rule. It can also be overridden outright with a typed value. The pass rate is reported as a percentage to two decimals. Completed candidates are those with a score for every station.

What the platform does not do

It does not write the decision record. There is no audit trail of who locked or withdrew a station and when, and no approval step before a computed cut score goes live. Both of those belong to the minute and the ratification, which happen outside any calculation screen.

It does not choose the method or define the borderline anchor. It does not train the examiners who apply it or justify a minimum number of stations. The software makes the arithmetic fast, repeatable, and visible. The defense is still a document your committee writes.

Was this page helpful?