How many judges does an Angoff panel need?
There is no single required number, but published guidance generally recommends panels of roughly 8-12 judges, and generalizability studies such as Hurtz and Hertz (1999) suggest 10-15 for stable cut scores. Fewer judges increase the standard error of the cut score; diversity of expertise matters as much as headcount.
Panel size is one of the first practical questions in planning an Angoff exercise, and the honest answer is that the literature offers ranges and reasoning, not a magic number.
What the literature says
The cut score from an Angoff panel is an average of judge estimates, so it carries sampling error: a different draw of judges would give a somewhat different standard. More judges shrink that error, with diminishing returns. Hurtz & Hertz (1999) examined this directly with generalizability theory and recommended on the order of 10-15 raters for dependable Angoff cut scores. Brandon's (2004) review of Angoff research and general treatments such as Cizek & Bunch (2007) land in a similar neighborhood, with common operational guidance in the 8-12 range and smaller panels treated as a recognized compromise rather than best practice.
Very small panels — three or four judges — are common in resource-constrained settings and are not automatically indefensible, but they demand honesty about the consequence: the standard error of the cut score is larger, and a single idiosyncratic judge moves the pass mark noticeably. If a small panel is unavoidable, report the variability across judges alongside the cut score, and consider how the standard error compares with the measurement error of the exam itself when candidates sit near the boundary.
Composition matters as much as count
Ten clones of the same course director are worth less than seven judges who differ in useful ways. Aim for:
- Content coverage. Judges collectively spanning the disciplines on the exam blueprint — a panel rating items outside anyone's expertise produces noise, and mapping panel expertise against the exam blueprint is a quick check.
- Familiarity with the candidates. Judges must know what learners at this stage can actually do; this is what keeps borderline-candidate estimates realistic.
- Mixed perspectives. Combining basic scientists and clinicians, or junior and senior faculty, guards against a shared blind spot shifting the standard.
Practical checklist
Decision | Reasonable default |
|---|---|
Panel size | 8-12; 10-15 where feasible |
Minimum to avoid | Fewer than 5-6 without reporting cut-score error |
Selection | Cover blueprint content areas; mix seniority and discipline |
Training | Borderline-candidate discussion plus practice ratings before live ratings |
Rounds | Two, with discussion and item-difficulty feedback between them |
One last reframing: judges are the expensive part of standard setting, so protect their time for judgment. Rating logistics, collating estimates, and producing the item statistics used as round-two feedback are all mechanical, and modern assessment platforms generate the item analysis automatically — meaning the panel's hours go into the discussion that actually improves the standard. For performance examinations, where panel-based item judgments give way to examiner-based methods entirely, see the comparison of OSCE standard-setting methods.