Skip to content

Standard Setting

How to set defensible cut scores for medical exams — Angoff, modified Angoff, borderline methods, and criterion- vs norm-referenced standards.

Standard setting is the process of deciding where the pass mark on an examination should sit — and being able to defend that decision. Every cut score is ultimately a policy judgment, but the judgment can be made systematically: structured methods such as Angoff, Ebel, Hofstee, borderline group, and borderline regression gather expert opinion about what a minimally competent candidate should be able to do and translate it into a score on the test.

The stakes are real. A cut score set too high fails competent students; one set too low certifies candidates who are not ready for supervised practice. When a failing student appeals, "the pass mark has always been 60%" is not a defense. A documented, criterion-referenced procedure — qualified judges, a clear definition of the borderline candidate, a replicable method, and evidence that the resulting standard behaves sensibly — is what makes a pass/fail decision hold up to academic and legal scrutiny (Cizek & Bunch, 2007; Norcini, 2003).

This topic covers the major methods, how they differ, and the practical decisions assessment teams face when implementing them.

Questions & answers

What is the difference between criterion-referenced and norm-referenced cut scores?

Criterion-referenced cut scores are set against a fixed performance standard — what a competent candidate must know — so all examinees can pass or fail. Norm-referenced cut scores depend on group performance (for example, failing the bottom decile), so a candidate's result varies with cohort strength. Licensure and medical education strongly favor criterion-referenced standards.

Read the full answer →

How does the Angoff method work?

In the Angoff method, a panel of subject-matter experts imagines a minimally competent (borderline) candidate and estimates, for each item, the probability that this candidate would answer correctly. Each judge's estimates are summed across items, and the panel's ratings are averaged to yield the examination's cut score.

Read the full answer →

How many judges does an Angoff panel need?

There is no single required number, but published guidance generally recommends panels of roughly 8-12 judges, and generalizability studies such as Hurtz and Hertz (1999) suggest 10-15 for stable cut scores. Fewer judges increase the standard error of the cut score; diversity of expertise matters as much as headcount.

Read the full answer →

What is the difference between the Angoff and modified Angoff methods?

Angoff's original proposal asked judges for yes/no judgments about whether a borderline candidate would answer each item correctly. Modified Angoff procedures — now the norm — use percentage or probability estimates instead, and typically add multiple rounds, discussion, and performance data between rounds. In practice, "modified Angoff" describes almost every Angoff study conducted today.

Read the full answer →

What is standard setting in assessment?

Standard setting is the structured, judgment-based process of deciding how much is enough on an examination — translating a performance expectation into a cut score that separates passing from failing. Methods such as Angoff, Ebel, and borderline regression combine expert judgment with test data to produce defensible, criterion-referenced standards.

Read the full answer →