How does the Angoff method work?
In the Angoff method, a panel of subject-matter experts imagines a minimally competent (borderline) candidate and estimates, for each item, the probability that this candidate would answer correctly. Each judge's estimates are summed across items, and the panel's ratings are averaged to yield the examination's cut score.
The Angoff method, named after William Angoff's brief but enormously influential description (Angoff, 1971), is the most widely used test-centered standard-setting procedure in medical education and credentialing. It sets the cut score from expert judgments about the test items rather than from candidate score distributions.
The procedure, step by step
- Convene a panel. Recruit judges who know the content and, critically, know the candidate population — faculty who teach and examine at the relevant stage. Panel size matters for the stability of the result; see how many judges an Angoff panel needs.
- Define the borderline candidate. Before any ratings, the panel discusses and agrees on a shared picture of the minimally competent candidate — one who just barely deserves to pass. This conceptualization step is the hardest part of the method and the most common source of trouble; unrealistic borderline definitions produce unrealistic standards.
- Rate each item. For every item, each judge independently estimates the probability that the borderline candidate would answer it correctly (e.g., 0.65, or "65 out of 100 borderline candidates"). Very easy items attract high estimates; hard or peripheral items attract low ones.
- Aggregate. Sum each judge's ratings across all items to get that judge's implied cut score, then average across judges. The mean (occasionally the median) becomes the exam's passing score.
- Iterate, usually. Most implementations run a second round: judges see the panel's rating distribution and, often, actual item difficulty data, discuss discrepant items, and re-rate. These additions are what make the procedure a "modified" Angoff, which is how the method is almost always run in practice.
A worked illustration
For a 4-item exam and three judges:
Item | Judge A | Judge B | Judge C |
|---|---|---|---|
1 | 0.90 | 0.85 | 0.80 |
2 | 0.60 | 0.55 | 0.65 |
3 | 0.75 | 0.70 | 0.70 |
4 | 0.40 | 0.50 | 0.45 |
Sum | 2.65 | 2.60 | 2.60 |
The panel mean is 2.62 out of 4, so the cut score is roughly 65%. On a real exam the same arithmetic runs over the full item set.
Strengths and cautions
Angoff standards are exam-specific — the cut score moves with the difficulty of the particular items — which is exactly what a fixed percentage pass mark cannot do. The method is transparent, well documented in the literature (Cizek & Bunch, 2007; Norcini, 2003), and straightforward to defend on appeal.
Its main weaknesses are cognitive: judges struggle to estimate probabilities for a hypothetical candidate, and ratings can drift toward what judges think the pass rate should be. Training, discussion rounds, and feeding back real item statistics mitigate this — one reason platforms that generate item analysis automatically, such as StudyDrome, make Angoff exercises considerably less laborious to run well. For OSCEs, where examiners observe real performances, examinee-centered alternatives are usually preferred; see the comparison of OSCE standard-setting methods.