Borderline regression needs enough candidates to fit a stable line, and enough spread in their grades to fit it through. The figure you will see quoted is 50. It comes from a single resampling study published in 2016. Later work found that spread matters more than headcount, and that cohorts of 20 to 30 can still set defensible marks.
That second finding is the one that rarely travels. The number gets repeated; the condition attached to it does not. So exam offices with 25 candidates abandon a method that would have worked, and exam offices with 200 assume theirs is fine when one thin station is quietly unstable.
How much does cohort size change the pass mark?
Homer, Pell, Fuller and Patterson answered this in 2016 by resampling real data. They collected four OSCE datasets from two medical schools, each with 17-21 stations. They drew 1,000 samples for each of a range of cohort sizes, from 15 to 250 candidates. For every sample, they recalculated the pass mark. Then they measured how much it moved.
The result is a clean relationship. The standard error of the pass mark shrinks with the square root of the cohort size. Quadruple your candidates and the error halves. It never reaches zero.
How far the pass mark can move
Standard error of the borderline-regression pass mark at the two extremes of the cohort sizes tested.
Whole-exam pass mark
One station's pass mark
Read the two measures separately. The whole-exam mark is the stable one, because summing across stations averages out much of the noise. A single station's mark is far more fragile. At 15 candidates, it can move by up to 10% of the available marks, purely from which candidates happened to sit it.
Two other numbers moved even more. The reliability coefficient and the R² of each regression both became unstable in small samples. That matters, because those are the statistics you would use to judge whether the cut score was trustworthy. In a small cohort, your quality checks get noisy exactly when you need them most.
Where does the 50-candidate figure come from?
From that same 2016 paper. Its conclusion is that cohorts with fewer than 50 candidates show an unacceptably large standard error for the pass mark.
That sentence is the source of nearly every cohort-size rule in circulation. It is a reasonable finding, honestly reported. But it has hardened into a threshold that the paper never framed as a pass-or-fail test, and most guidance stops there.
Why has that advice changed?
Because the same research group went back and tested it against small cohorts in the field.
Homer and colleagues published that work in 2020. They looked at three real high-stakes assessments, all below the 50 line. The exam for international medical graduates seeking UK registration runs 30 to 35 candidates per administration. A sequential undergraduate OSCE recalls 20 to 50 students for a second round. A physician associate program examines cohorts of 20 to 30.
They then asked how often borderline regression actually failed. The answer was: rarely. Problematic stations accounted for 5%, 14%, and 12% of the three exams. In the large majority of stations, the method produced defensible marks.
How often the method actually failed
Share of stations judged problematic in three real high-stakes exams, all below the 50-candidate line.
IMG registration exam
30–35 candidates per sitting
Sequential undergraduate OSCE
20–50 recalled candidates
Physician associate programme
cohorts of 20–30
More useful still is what explained the failures. They were not caused by headcount. They came from a weak relationship between global grades and checklist scores, or from a lack of spread in the scores. Where a station had a real range of performance and grades that tracked it, the regression worked at 25 candidates.
Moreno-López and colleagues reached a compatible conclusion in dental education in 2022, comparing borderline regression against a modified Angoff standard. They found it an acceptable replacement for Angoff in small cohorts — with conditions attached: a genuine range of candidate ability, well-designed stations, and trained examiners.
Note the direction of that paper. It argues that borderline regression can be used in small cohorts. It is often cited in support of the opposite claim.
What matters more than headcount?
Spread. The regression fits a line to the relationship between the checklist score and the global grade. If either one barely varies, there is no relationship to fit.
Three distributions break it, and only the third has anything to do with cohort size.
- No variation in the grades. Every candidate is rated the same. The line is undefined, and no amount of candidates will fix it. A hundred candidates, all graded "pass," give you nothing.
- Range restriction. Everybody does well. Nobody is rated at or near borderline. The fitted line still exists, but the borderline point sits outside the data you actually observed, so the mark is an extrapolation rather than a reading.
- Too few points to fit through. The genuine small-sample problem, and the one the 2016 paper measured.
So the useful question is not "how many candidates do I have". It is "how many distinct global grades did my examiners use, and are there candidates in the lower ones?" Fifteen candidates spread across a four-point grade scale carry more information about where the borderline sits than sixty candidates who were all rated "good pass".
Is your n per station, or per exam?
Per station. This is the trap that catches large programs, and it is the reason a 120-candidate exam can contain a station as thin as a 20-candidate one.
In a single circuit run in waves, every candidate passes through every station, so the two numbers are the same. In parallel circuits, they are not. Run 120 candidates through six parallel circuits, with each station replicated six times, each with its own examiner, who sees 20 candidates.
Whether that is one regression on 120 candidates or six regressions on 20 depends on a decision you have already made without necessarily noticing: are the six copies one station or six? If your examiners differ in stringency, pooling them hides that. If you split them, each regression runs on 20. Neither is wrong, but the choice changes what your cohort size actually is, and any per-station report will show you the n it used.
How do you check whether your own cut score is stable?
You do not need a large cohort to know whether a particular mark held up. You need to look at the station, one at a time.
- Plot the scatter. Checklist score against global grade, one plot per station. AMEE Guide 49 recommends this as routine practice regardless of the other metrics, and it, in a second, exposes what a table of numbers hides.
- Read the R². Pell and colleagues consider a correlation above 0.5 to indicate a reasonable relationship between checklist scores and global grades. Below that, the two judgments are disconnected, and the derived mark rests on a weak link.
- Check the grade distribution. Count how many candidates sit in each global grade. Empty or near-empty categories at the bottom are the warning that the borderline point is being extrapolated.
- Drop one candidate and recompute. If removing a single performance moves the station mark appreciably, then the mark is not stable, regardless of the cohort size.
- Compare against something. The borderline group mark on the same data, or the same station's mark from a previous sitting. Two methods that disagree sharply on one station is a finding about that station.
None of this requires new software. It requires the station-level data you already collected and someone to review it before the exam board meets.
What should you do when the cohort is always small?
Some programs never reach 50 and never will. Dental schools, physician associate programs and postgraduate specialties sit permanently in that range. They still have to publish a defensible mark.
- Borrow the mark from a bigger sitting. Where a station has been used before in a larger cohort with sound metrics, use that existing pass mark instead of refitting on the small group. This is the remedy Homer and colleagues recommend in the 2020 paper, and it is the practice already in use in the sequential exams they studied.
- Use a panel method for new stations. An Angoff-style standard needs no candidate data at all, which is exactly the case where regression has least to work with. It costs panel time instead. How many judges an Angoff panel needs is the equivalent question on that side.
- Pool across sittings, carefully. Combining administrations grows the data. It works only if the station and its checklist are unchanged. You are also averaging across cohorts that may differ. Record what you pooled and why.
- Fix the grading, not the arithmetic. The problem is that examiners use two of the five grades; examiner training moves the number more than another 20 candidates would.
Whichever route you take, minute it. A cut score derived from 22 candidates with a documented rationale is defensible. The same mark with no record is not.
A checklist before you sign off a cut score
- Every station has a scatter plot someone has actually looked at.
- R² is above 0.5, or the station is on the review list.
- Each global grade includes candidates, including those below the borderline.
- You know whether your n is per station or per circuit.
- Stations that failed a check have a named action: borrowed mark, panel standard, or review.
- The exam board minutes record the method, the checks, and the exceptions.
Where do software thresholds sit?
Most systems that compute borderline regression carry two limits. There is a hard floor below which they will not fit a line at all. There is also a soft warning that flags a small cohort. StudyDrome Exam Manager fits from three valid candidates per station upward. A valid candidate is one who completed the station and carries a global rating. Below 20 completed candidates, it raises a small-cohort advisory flag. The flag never blocks anything. The exam board decides.
Be clear about what a threshold like that is. It is a floor, not a certificate. The evidence above puts the onset of serious instability higher than 20, so a station that passes without a warning at 25 candidates has not been declared sound. It has only avoided the worst case. The checks in the previous section tell you whether the mark holds, and they apply to every station, whether it was flagged or not.
Frequently asked questions
Is borderline regression better than Angoff for OSCEs?
For most OSCEs, yes, on two grounds. It uses the actual performance of each candidate rather than a panel's prediction of a hypothetical one, and it costs no panel time. Angoff has two real advantages: it requires no candidate data, so it works for a brand-new station, and it can be run before the exam. Small programs often run both and compare.
Does borderline regression need candidates at the borderline grade?
No, and this is a common confusion with the borderline group method. Borderline group averages the checklist scores of candidates rated borderline, so it needs enough of them. Regression fits a line through every candidate at the station and reads the line at the borderline point, so it requires no borderline-rated candidates at all. AMEE Guide 49 states this difference directly.
What happens if every candidate gets the same global rating?
The regression cannot be computed. A line needs variation on both axes, and identical grades mean there is no variation to fit against. Software will normally refuse to produce a mark rather than return a misleading one. The fix is not more candidates. It is examiner training, or a grade scale that examiners are willing to use across its full range.
What are global ratings and how do they map to numbers?
A global rating is the examiner's holistic judgment of the overall performance, made independently of the checklist. Scales usually run four or five points, such as fail, borderline, pass, good pass and excellent. Those points are numbered in order for the regression, and the cut score is the fitted value at the borderline number. The judgment should be independent, not a total of the checklist.
Should you pool OSCE data across sittings to grow the cohort?
Sometimes, with care. Pooling helps only where the station, its checklist and its grade scale are unchanged between sittings. You are also assuming the cohorts are comparable, which is weaker for a resit group than for successive year groups. Record what you combined and why, so the exam board can see the assumption it is accepting.
How do you make a standard-setting decision defensible?
By documenting the process rather than by picking a method name. Record the borderline-candidate definition, the examiner training, the computation, the quality checks you ran, and every station where you departed from the default and why. A mark from a small cohort with that record behind it will survive scrutiny. A mark from a large cohort without one will not.
Where to go next
If you are still choosing between methods, our comparison of OSCE standard-setting methods works through Angoff, borderline group, borderline regression, and contrasting groups side by side. For the mechanics of the method itself, see what the borderline regression method is.
To run the checks in this article on your own data, the free OSCE standard-setting calculator computes station marks, R² and the scatter plots, with no signup. Our standard-setting page covers the four methods. It also shows the checks that run against each one.
References
- Pell G, Fuller R, Homer M, Roberts T (2010). How to measure the quality of the OSCE: a review of metrics — AMEE Guide No. 49. Medical Teacher 32(10):802–811.
- Homer M, Pell G, Fuller R, Patterson J (2016). Quantifying error in OSCE standard setting for varying cohort sizes: a resampling approach to measuring assessment quality. Medical Teacher 38(2):181–188.
- Homer M, Fuller R, Hallam J, Pell G (2020). Setting defensible standards in small cohort OSCEs: understanding better when borderline regression can 'work'. Medical Teacher 42(3):306–315.
- Moreno-López R, Sinclair S, Wright J, Rowland C (2022). Can borderline regression method be used to standard set OSCEs in small cohorts? European Journal of Dental Education 26(4):686–691.