Every published exam produces item statistics. You do not have to ask.
Statistic | How it is computed |
|---|---|
Difficulty index (p-value) | Correct responses ÷ total, to 4 decimal places |
Discrimination index | Top 27% minus bottom 27%, ÷ group size. Groups are |
Distractor analysis | Selection count and share per option; any wrong option chosen by under 5% is flagged non-functioning |
Distractor efficiency | Functioning distractors ÷ total distractors × 100, to 2 dp |
Classification | Good, Review or Poor, derived from the flags |
Likert items are excluded from every calculation.
Items with fewer than 10 responses are marked "Insufficient Data" and classified NotRated. Their flags are discarded rather than shown.
The reason is blunt. Below about ten, the top and bottom 27% groups are one or two people, and the numbers are noise. The gate applies twice: once for the report, once per item on its own count.
OSCE stations get their own analysis. Facility, discrimination on top and bottom 33% groups, and item-rest correlation per checklist item.