What is the difference between the Angoff and modified Angoff methods?
Angoff's original proposal asked judges for yes/no judgments about whether a borderline candidate would answer each item correctly. Modified Angoff procedures — now the norm — use percentage or probability estimates instead, and typically add multiple rounds, discussion, and performance data between rounds. In practice, "modified Angoff" describes almost every Angoff study conducted today.
The short version: nearly every "Angoff" standard-setting exercise run today is technically a modified Angoff, and the label mostly signals which enhancements were bolted onto the original idea.
What Angoff originally proposed
Angoff (1971) described the method in a single paragraph — and, notably, in a footnote. The main text proposed that judges state, for each item, whether a "minimally acceptable person" would or would not answer it correctly: a dichotomous 0/1 judgment, with the count of "yes" items becoming the judge's cut score. The famous footnote suggested a variation: estimate the probability that the minimally acceptable person answers correctly, and sum the probabilities.
Strictly, then, the yes/no version is the "original" Angoff and the probability version is already a modification — one so universally adopted that most people now think of it as the base method.
What "modified Angoff" means in practice
Because the original description was so sparse, implementers filled in the operational details, and the accumulated additions travel under the "modified Angoff" name. Typical modifications include:
Modification | Purpose |
|---|---|
Probability/percentage ratings instead of yes/no | Finer-grained judgments, less information loss |
Structured borderline-candidate discussion before rating | Shared conceptualization across judges |
Multiple rating rounds with discussion of discrepant items | Convergence and error-checking |
Feedback of empirical item difficulty (p-values) between rounds | Reality-checks implausible ratings |
Feedback of the implied pass rate | Guards against standards no one would accept |
Dropping or restricting the rating scale (e.g., yes/no/maybe) | Reduces cognitive load |
Not every modification pulls in the same direction — performance-data feedback, in particular, nudges a purely test-centered method toward the empirical, and how much data to show judges is a genuine design decision. Cizek & Bunch (2007) and Hambleton & Plake (1995) discuss these variants and their trade-offs in detail.
Does the distinction matter?
For search engines and textbooks, yes; for defensibility, less than you might think. What matters when a standard is challenged is not which label you used but whether the procedure was documented, the judges qualified, the borderline candidate defined, and the resulting cut score evaluated for stability and plausibility (Norcini, 2003). A well-run "modified Angoff with two rounds and difficulty feedback" is easier to defend than a purist single-round exercise, precisely because the modifications exist to reduce judge error.
When writing up your process — for an assessment committee, an accreditor, or an appeal — name the specific modifications rather than just saying "Angoff." If you are choosing a method rather than a variant, start with what standard setting is and, for performance exams, the comparison of OSCE standard-setting methods.