What is standard setting in assessment?
Standard setting is the structured, judgment-based process of deciding how much is enough on an examination — translating a performance expectation into a cut score that separates passing from failing. Methods such as Angoff, Ebel, and borderline regression combine expert judgment with test data to produce defensible, criterion-referenced standards.
Standard setting is how an examination's pass mark gets decided. Scoring a test tells you how many points each candidate earned; standard setting answers the separate question of how many points are enough. The output is a cut score — the boundary between pass and fail, or between adjacent grade categories.
Why a procedure is needed at all
No cut score is "true" in the way a measurement is true. As Cizek & Bunch (2007) put it, standard setting is fundamentally a matter of informed policy, not statistical discovery. What distinguishes a defensible standard from an arbitrary one is the process: qualified judges, a shared definition of the minimally competent (borderline) candidate, a recognized method applied systematically, and documentation of every step. Norcini (2003) and Ben-David's AMEE Guide No. 18 (2000) both emphasize that the credibility of a pass/fail decision rests on the credibility of the procedure that produced it.
A fixed institutional pass mark — 50%, 60%, 70% — fails this test. Item difficulty varies from one exam form to the next, so a constant percentage cut means the real standard silently drifts every administration.
The main families of methods
Family | Judges examine | Examples |
|---|---|---|
Test-centered | The items themselves | Angoff, Ebel, Nedelsky, Bookmark |
Examinee-centered | Actual candidate performances | Borderline group, borderline regression, contrasting groups |
Compromise | Both judgment and pass-rate acceptability | Hofstee, Beuk |
Test-centered methods such as the Angoff method dominate written examinations in medical education, while examinee-centered methods like borderline regression are the norm for OSCEs, where examiners already observe each performance directly.
What a defensible process looks like
Most published guidance converges on the same elements: select a panel of judges who know both the content and the candidate population; train them on the concept of the borderline candidate; apply the chosen method, usually over more than one round with discussion; and evaluate the result — the standard error of the cut score, the plausibility of the resulting pass rate, and consistency with other evidence of candidate ability. All of these steps are strengthened by good item-level data, which is one reason routine item analysis and standard setting tend to mature together in an assessment program.
Standards should also be revisited, not fossilized. Curricula, cohorts, and blueprints change; a cut score set once and reused for years gradually loses whatever defensibility it had at the start.