Exam Manager names every statistic it computes. Four of six competitors we track name no reliability coefficient at all.
Statistic | How it works |
|---|---|
KR-20 or Cronbach's α | Chosen automatically. KR-20 when every scorable item is MCQ or True/False and none allows partial credit. Otherwise α. The report names which one it used. |
Standard error of measurement | SEM = SD × √(1 − reliability). Computed for written exams, OSCE stations and the OSCE total. |
Difficulty index (p-value) | Correct responses divided by total, to 4 decimal places. |
Discrimination index | Top 27% minus bottom 27%, divided by group size. OSCE uses 33% groups. |
Distractor analysis | Any wrong option chosen by under 5% of candidates is flagged non-functioning. Rolls up to one efficiency percentage. |
Borderline regression | Ordinary least squares, checklist score on Global Rating. Returns slope, intercept and R². |
Cohen's d and Welch's t-test | For comparing forms, cohorts and semesters. |
Point-biserial | Computed on every item and available through the API. |
Reliability uses each candidate's latest attempt, so retakes cannot inflate it. Items with fewer than 10 responses are marked "Insufficient Data".
Read our item analysis guide for the method behind these.