Exam Manager for assessment leads and psychometricians

The evidence layer, computed on every exam

StudyDrome Exam Manager computes item statistics on every published exam without being asked. Difficulty, top-and-bottom-27% discrimination and distractor analysis per item. KR-20 or Cronbach's alpha selected automatically, with the standard error of measurement beside it. Items under 10 responses are marked Insufficient Data rather than rated.

What runs on every exam?

Every published exam produces item statistics. You do not have to ask.

Statistic

How it is computed

Difficulty index (p-value)

Correct responses ÷ total, to 4 decimal places

Discrimination index

Top 27% minus bottom 27%, ÷ group size. Groups are ceil(n × 0.27), minimum 1

Distractor analysis

Selection count and share per option; any wrong option chosen by under 5% is flagged non-functioning

Distractor efficiency

Functioning distractors ÷ total distractors × 100, to 2 dp

Classification

Good, Review or Poor, derived from the flags

Likert items are excluded from every calculation.

Items with fewer than 10 responses are marked "Insufficient Data" and classified NotRated. Their flags are discarded rather than shown.

The reason is blunt. Below about ten, the top and bottom 27% groups are one or two people, and the numbers are noise. The gate applies twice: once for the report, once per item on its own count.

OSCE stations get their own analysis. Facility, discrimination on top and bottom 33% groups, and item-rest correlation per checklist item.

How are problem items surfaced?

Flags carry a severity and a plain recommendation. A reviewer gets a next action, not a number to interpret.

Trigger

Severity

What it says

Difficulty ≥ 0.8

Warning

"80%+ of students answered correctly. Consider increasing difficulty or reviewing distractors."

Difficulty ≤ 0.3

Warning

"Less than 30% answered correctly. Verify question clarity and correctness of answer key."

Discrimination < 0

Critical

"Low-performing students answer correctly more than high-performing ones. Possible wrong answer key or ambiguous wording."

Discrimination < 0.2

Warning

"Question does not effectively differentiate between strong and weak students. Consider revision."

Distractor efficiency < 40

Warning

"Less than 40% of wrong answer options are attracting responses. Update non-functioning distractors."

Negative discrimination is the only trigger that marks an item Poor on its own.

Calibration compares the difficulty an author assigned against the difficulty candidates actually demonstrated. A gap of 3 or more raises a flag.

The calibration list uses a lower bar. It surfaces every item with a gap of 2 or more. You see the drift before it becomes a flag.

How is reliability reported?

The coefficient is chosen for you, not configured. KR-20 runs when every scorable item is multiple choice or true/false, and no item type allows partial credit. Alpha runs otherwise. The report always names the method.

Reliability uses each candidate's latest attempt only. A retake-heavy cohort cannot inflate internal consistency by duplicating correlated rows. The rule applies to the coefficient and its recommendations.

The coefficient becomes a five-tier verdict, verbatim from the API:

Coefficient

Verdict

0.90 and above

Excellent — suitable for high-stakes decisions

0.80 to 0.89

Good — suitable for most assessment purposes

0.70 to 0.79

Acceptable — adequate for classroom testing

0.60 to 0.69

Questionable — consider revision

Below 0.60

Poor — significant measurement error

Recommendations attach to the weak end. Increase item count below 0.7. A note fires when the paper carries fewer than 20 scorable items. And a count of negatively discriminating items to revise.

The standard error of measurement sits alongside it: SD × √(1 − reliability). It is computed for written exams, for OSCE stations and for the OSCE total. It is the number that matters at the pass mark, and the one an appeal turns on.

How healthy is the bank?

Bank health answers a different question. Not "how did this exam perform", but "can this bank still carry an exam".

A quality score reports the share of administered items classified Good, to one decimal place.

Alongside it sits the count never administered, and the difficulty spread across Easy (0–3), Moderate (4–6) and Hard (7–10) bands. Average difficulty, discrimination and distractor efficiency follow.

Items are keyed by their original identity. A question used in five exams is one row, not five.

The assembly profile turns that into a readiness check. It cross-tabs topic against difficulty band. It counts Good, Review and Poor per cell. And it raises named warnings:

  • "No questions available in {band}"
  • "Fewer than 5 good questions in {band}"
  • "Topic '{tag}' has only poor-quality questions"
  • "{n} questions ({pct}%) have never been administered" — above 20%

It reports what the bank can support. You build the paper.

How do you compare forms and cohorts?

Form comparison answers whether two versions of a paper behaved the same way.

Per form: candidate count, mean, pass rate, standard deviation, average item difficulty.

Per pair: the score gap, Cohen's d as an effect size, and Welch's t-test with a p-value.

A warning fires when a difference is both statistically significant and larger than 3 points: "{lower} mean score is {gap}% lower than {higher} (p={p}) — forms may not be equivalent".

This is an equivalence check on observed performance. It compares how two forms behaved. It is the evidence you take to a board when someone asks whether the January and June papers were comparable.

Cohort comparison does the same across course groups. It reports a pair when the gap exceeds 5 points or the p-value falls below 0.05.

Longitudinal views compare semesters and track a student across exams, using first-half against second-half means.

Around 35 report screens sit across the Analytics workspace and the per-course tab. That is 28 in the workspace and 7 inside a course.

What keeps the numbers honest?

Several guards sit inside the calculations, not in the reader's head.

  • Small-N suppression: Items under 10 responses are classified NotRated and their flags discarded. Individual statistics carry their own minimums. Polytomous discrimination needs 3 per group. Skewness needs 3. Cohen's d and Welch's t need 2 per group.
  • Latest attempt only: The reliability coefficient uses one attempt per candidate. Retakes cannot inflate it.
  • Likert excluded: Likert items are removed from the scorable set before any calculation runs.
  • True/False excluded from the distractor problem list: Zero distractor efficiency on a true/false item means it was easy, not that its distractors failed. The bank-health rollup leaves it out.
  • Clamped and rounded consistently: Reliability clamps to 0–1. Discrimination clamps to −1 to 1. Each statistic carries a stated precision.

Who is this for?

An assessment lead, psychometrician or quality office. Someone who has to defend a result, not just announce it.

The evidence chain is the point.

Publishing writes an immutable snapshot of 40 record types. The paper each candidate sat is kept exactly as they sat it.

Item statistics run on that snapshot. Reliability and the standard error of measurement quantify the decision at the pass mark. For OSCEs, standard setting derives the cut score with its own warnings.

Every grade, re-grade and un-grade writes a numbered row. Nothing is overwritten.

Read our item analysis guide for the underlying method. Read KR-20 or Cronbach's alpha for the choice between the two coefficients.

Frequently asked questions

Which reliability coefficient does StudyDrome report?

KR-20 or Cronbach's alpha, selected automatically. KR-20 applies when every scorable item is multiple choice or true/false, and no item type allows partial credit. Alpha applies otherwise. The report names the method it used. It shows the standard error of measurement alongside, plus a five-tier reading of the coefficient.

How does StudyDrome decide an item is a problem?

Five automatic triggers, each with a severity and a written recommendation. Difficulty at or above 0.8. Difficulty at or below 0.3. Discrimination below zero. Discrimination below 0.2. Distractor efficiency below 40. Negative discrimination is Critical, and the only trigger that marks an item Poor on its own.

What stops a small cohort producing misleading statistics?

A suppression gate. Items with fewer than 10 responses are classified NotRated and marked Insufficient Data. Their flags are discarded rather than displayed. Individual statistics carry their own minimums on top. And reliability uses one attempt per candidate, so retakes cannot inflate it.

Can we check whether two versions of a paper were equivalent?

Yes. Form comparison reports each form's mean, pass rate, standard deviation and average item difficulty. Per pair it reports the score gap, Cohen's d and a Welch's t-test p-value. A warning fires when a difference is both significant and larger than 3 points. It is an equivalence check on the forms you ran.

How do we know a bank can still carry an exam?

Bank health reports the share of administered items classified Good. It reports how many have never been administered, and the difficulty spread across three bands. The assembly profile then cross-tabs topic against band. It raises named warnings for empty bands, thin bands, all-poor topics and a large never-administered share.

What evidence survives if a result is challenged?

Publishing writes an immutable snapshot of 40 record types. The paper each candidate sat is preserved exactly. Item statistics, the reliability coefficient and the standard error of measurement all attach to that published version. Every grade, re-grade and un-grade writes a numbered history row, and nothing is overwritten.

Book a pilot