---
title: "Choosing and defending your pass standard: the decision record"
description: "What an OSCE standard-setting decision record must contain, who decides each item, when it is written, and how it answers an appeal a year later."
canonical: https://studydrome.com/docs/osce/scoring/choosing-and-defending-a-pass-standard/
updated: 2026-10-09
---

# Choosing and defending your pass standard: the decision record

What an OSCE standard-setting decision record must contain, who decides each item, when it is written, and how it answers an appeal a year later.

A pass standard is defended by its record, not by the name of its method. The record says who chose the method, what the borderline grade means, how station marks add up, which extra rules apply, and who signed the result. Most of it is written before the exam runs. An appeal a year later is answered from that file, or it is not answered at all.

## Why the method name is not a defense

Two schools can both say "borderline regression" and mean different exams. One trained its examiners on the borderline anchor, and one did not. One sums the station pass marks and the other averages them. One also requires a minimum number of stations passed. The name is the same. The standard is not.

Norcini sets out what actually makes a method choice credible. It should suit the purpose of the test, rest on expert judgment, be informed by data, be supported by research, be transparent, and show due diligence. He adds that the credibility of a standard rests largely on who the standard setters are. A broad and knowledgeable group is essential. In an OSCE with a borderline method, the judges are the examiners at the stations, so examiner training is part of the standard, not a separate task.

One choice is ruled out first. The AMEE guide on OSCE organization records that norm referencing is usually deemed unacceptable for clinical competency licensing tests. A pass mark that moves with the cohort cannot certify that a candidate is safe to progress. The [methods themselves are compared elsewhere](/blog/osce-standard-setting-methods-compared/). The 2020 Ottawa consensus statement states that the borderline regression method is now regarded as the standard-setting method of choice for OSCEs. This page is about the file that method choice sits in.

## What has to be in the record

| Decision | The choice to make | Who decides | Where it is recorded |
| --- | --- | --- | --- |
| Type of standard | Criterion referenced, never a quota | Assessment committee | Assessment policy |
| Method | Angoff, borderline group, borderline regression, or a set mark | Assessment committee, on advice | Standard-setting minute |
| The global scale | Its points and their wording, and which point is the borderline | Station writers and the committee | Marking scheme |
| Examiner training on the anchor | What examiners are told the borderline point means | Assessment lead | Briefing pack and register |
| Level of the cut | Per station, or one mark for the whole exam | Assessment committee | Standard-setting minute |
| How marks combine | Sum of station marks, or their average | Assessment committee | Standard-setting minute |
| Extra rules | Minimum stations passed, must-pass stations, critical errors | Assessment committee | Assessment policy |
| Rounding | The number of decimals, and which way a half goes | Assessment committee | Standard-setting minute |
| Faulty stations | What happens to a station withdrawn after the day | Exam board | Board minute |
| The candidate at the line | Whether a borderline result is resolved, and how | Exam board | Board minute |
| Who ratifies | The people who turn a computed mark into a result | Exam board | Terms of reference |
| What candidates are told | Method, rules and appeal grounds, published before the exam | Assessment lead | Candidate handbook |

Most rows are settled long before exam day. Only two are filled in afterward: the number the method produces, and any station the board withdraws. That split is the point of the record. The rules are fixed while nobody knows who they will fail.

## When is the standard actually set?

The method, the scale, the aggregation rule, and the extra rules are set before the exam. With Angoff or a set mark, the number is set before it too. With a borderline method, the number can only come from the exam itself, because it is read off the examiners' own judgments. That is not a late decision. It is a fixed procedure that produces a number.

A change made after the pass list is visible is the one move that cannot be defended. If a standard has to move, the reason is recorded before anyone sees who is affected, the board agrees it, and the minute names the evidence. A station with a broken item is a reason. A pass rate the faculty dislikes is not.

Norcini also asks for a check after the fact. Do the pass rates sit in a sensible relationship with other markers of competence? A cut score that fails a third of a cohort that passed everything else deserves a second look. Check the exam before applying it to the candidates.

## What about the candidate exactly on the line?

Every exam has candidates whose scores fall within the exam's error. Saying so in the record is better than pretending the mark is exact. Schools take three routes. Some pass on the mark as computed. Some widen the borderline grade and resolve it with a rule. Some send the candidate to further testing, which is what a [sequential design](/docs/osce/scoring/sequential-osce/) does by design.

The middle route has evidence to back it up. A medical school in Sydney replaced a "borderline pass" grade with a plain "borderline" grade on a four-point scale of distinction, pass, borderline and fail. Each borderline grade was then resolved to a pass or a fail by a formula. It used the station's difficulty and how that candidate performed in the same domain at other stations. Of 14,634 observations, 4.69 percent were borderline. Mean scores by domain did not move. The failure rate for the exam rose from 0.7 percent to 4.1 percent. Examiners and students were satisfied with the outcome. The lesson is not that every school needs that formula. It is that a stated rule for the line found failures; an unstated one was hiding.

## Does a small cohort change the record?

It changes what the record has to show, not what it has to contain. The stability of a computed cut score depends more on the spread of the global grades at each station than on the headcount. A station where everyone received the same grade cannot produce a regression line at all, regardless of the cohort size. The evidence on [cohort size and borderline regression](/blog/borderline-regression-sample-size/) covers the arithmetic and the per-station check.

The same logic applies to a design with one examiner per station. The global rating that anchors the standard is, in effect, one person's judgment, and nothing in the data will tell you it was harsh. That is an argument for training on the anchor and for a recorded fallback, not for a different method.

For a small program, the practical answer is to record a fallback in advance. Name the check that decides whether a station's computed mark is used. Name what is used instead when it fails: the borderline group mean, a mark carried from a previous sitting, or a committee judgment. A fallback procedure is chosen in advance. The same fallback chosen after results are out is an adjustment.

## Limits

A record does not make a weak exam defensible. A standard set well on eight stations is still a standard on eight stations. The Ottawa consensus is blunt about the floor. Exams with fewer than 12 stations or under 150 minutes of testing time will not yield reliable results. The record also cannot fix examiners who were never trained on the anchor. In a borderline method, their judgment is the standard. And it is only evidence if it exists in one place. A method in a handbook, a rule in an email, and a decision in someone's memory are not a record.

## Frequently asked questions

### What makes an OSCE pass mark defensible?

A written record showing the method, who chose it, what the borderline point means, how station marks combine, the extra rules, and who ratified the result. Norcini's criteria are that the method suits the purpose of the test, rests on expert judgment, is informed by data, is supported by research, is transparent, and shows due diligence.

### Should the pass mark be set before or after the OSCE?

The method, the scale, and the rules are set beforehand. With a borderline method, the number itself can only come from the exam, because it is read from the examiners' own judgments. That is a fixed procedure, not a late decision. Nothing changes after the pass list is visible.

### Can an OSCE pass mark be changed after the exam?

Only for a recorded reason agreed by the exam board before anyone sees who is affected. A station found to be faulty is a reason. A pass rate the faculty dislikes is not. The minute names the evidence and the decision.

### Why is a norm-referenced OSCE pass mark not acceptable?

The mark moves with the cohort, so the same performance passes in a weak year and fails in a strong one. The AMEE guide on OSCE organization records that norm referencing is usually deemed unacceptable for clinical competency licensing tests.

### What should candidates be told about the pass standard?

The method used, whether a minimum number of stations must be passed, whether any station is must-pass, what a critical error does, and the grounds on which a result can be appealed. All of it is published before the exam, not after the results.

## Sources

**Sources**

1. [Norcini JJ. Setting standards on educational tests](https://doi.org/10.1046/j.1365-2923.2003.01495.x) (Medical Education, 2003)
2. [Boursicot K, Kemp S, Wilkinson T, et al. Performance assessment: consensus statement and recommendations from the 2020 Ottawa Conference](https://doi.org/10.1080/0142159X.2020.1830052) (Medical Teacher, 2021)
3. [Shulruf B, Adelstein BA, Damodaran A, et al. Borderline grades in high stakes clinical examinations: resolving examiner uncertainty](https://doi.org/10.1186/s12909-018-1382-0) (BMC Medical Education, 2018)
4. [Pell G, Fuller R, Homer M, Roberts TE. How to measure the quality of the OSCE: a review of metrics. AMEE Guide No. 49](https://doi.org/10.3109/0142159X.2010.507716) (Medical Teacher, 2010)
5. [Khan KZ, Gaunt K, Ramachandran S, Pushkar P. The Objective Structured Clinical Examination (OSCE): AMEE Guide No. 81. Part II: organisation and administration](https://doi.org/10.3109/0142159X.2013.818635) (Medical Teacher, 2013)

> [!TIP]
> **In practice**
> The record is a document your software will not write for you. To see which parts of it a platform can hold, and which stay with the committee, read /docs/osce/scoring/in-studydrome-standard-setting/.
