What is the difference between criterion-referenced and norm-referenced cut scores?
Criterion-referenced cut scores are set against a fixed performance standard — what a competent candidate must know — so all examinees can pass or fail. Norm-referenced cut scores depend on group performance (for example, failing the bottom decile), so a candidate's result varies with cohort strength. Licensure and medical education strongly favor criterion-referenced standards.
The distinction is about what the cut score is anchored to: a standard of performance, or the performance of other candidates.
Norm-referenced standards
A norm-referenced (relative) cut score is defined by the score distribution of the cohort: fail the lowest 5%, pass anyone above one standard deviation below the mean, and so on. The standard travels with the group.
The convenience is obvious — no judgment panel needed, pass rates are predictable by construction. The problems are equally obvious. A candidate's fate depends on who else sat the exam: the same performance passes in a weak cohort and fails in a strong one. A fixed failure quota fails someone even if every candidate is demonstrably competent, and passes the top of a cohort in which no one is. For decisions that certify readiness to treat patients, that is very hard to defend, which is why guidance in health-professions assessment (Norcini, 2003; Ben-David's AMEE Guide No. 18, 2000) consistently recommends against relative standards for pass/fail decisions.
Criterion-referenced standards
A criterion-referenced (absolute) cut score is anchored to a defined level of performance — what a minimally competent candidate should be able to do — independent of how the cohort happens to score. In principle, everyone can pass; in principle, everyone can fail. The standard is set through a structured judgment process such as the Angoff method for written exams or borderline regression for OSCEs.
Note that "criterion-referenced" does not mean "a fixed percentage." An institutional pass mark of 60% on every paper looks absolute but is anchored to nothing: because item difficulty changes between forms, the effective standard drifts with the test rather than with the criterion. A genuine criterion-referenced standard is re-derived for each exam form from judgments about its items or performances.
Side by side
Criterion-referenced | Norm-referenced | |
|---|---|---|
Anchor | Defined performance standard | Cohort score distribution |
Same performance, different cohort | Same outcome | Outcome can flip |
Pass rate | Free to vary | Fixed by design |
Effort to set | Judgment panel, documentation | Trivial |
Defensibility for competence decisions | Strong when well run | Weak |
Where relative information still helps
Cohort data is not useless. Compromise methods such as Hofstee deliberately blend absolute judgment with acceptable pass-rate bounds, and reviewing the pass rate implied by a criterion-referenced cut is a standard reality check (Cizek & Bunch, 2007) — a wildly implausible failure rate usually signals a problem with the exercise, not with the cohort. Norm-referenced comparisons also remain appropriate for genuinely selective decisions, such as ranking applicants for a limited number of places. But for competence decisions — progression, graduation, licensure — the cut score should be criterion-referenced, documented, and set by a recognized method (see what standard setting is).