Item analysis in StudyDrome Exam Manager

Every question graded, the moment the exam is

StudyDrome Exam Manager scores every question once an exam closes. Each item gets a difficulty index, a top and bottom 27% discrimination index, and a distractor breakdown. Automatic flags then classify it Good, Review or Poor. Items with fewer than 10 responses stay unrated, so a small cohort never produces a false verdict.

What does the item analysis report show?

Open Analytics and pick the Question Quality Analysis tile, in the Item Analysis group. Its own description reads "Evaluate question quality, calibration, and integrity after exams". Three tabs follow: Question Quality, Calibration and Integrity.

The Question Quality tab is one row per question, five columns wide.

Column

What it measures

PI

Performance Index. The share who answered correctly, 0 to 1. Higher means easier.

DI

Discrimination Index. How well the question separates strong from weak, −1 to 1. Higher is better.

DE

Distractor Efficiency. The share of wrong options that drew responses, as a percentage.

Classification

Good, Review or Poor, read straight from the flags.

Flags

The named faults on that question, each with a severity.

Point-biserial correlation is computed on every item and available through the API. Whole-paper reliability has its own report: see exam reliability.

For the method behind the numbers, read the complete guide to item analysis. Or try it on your own data with the free item analysis calculator.

How is the discrimination index calculated?

  1. 1
    Candidates are ranked by the points they earned. One ranking serves the whole report.
  2. 2
    The top and bottom 27% become two groups. Each holds at least one candidate, so a small sitting still resolves.
  3. 3
    StudyDrome counts correct answers in each group, then divides the gap by the group size. That is the discrimination index, from −1 to 1.
  4. 4
    The difficulty index is simpler: correct answers over total answers. Both run to four decimal places.
  5. 5
    Likert questions stay outside the maths by design. They carry no correct answer to score.

How does StudyDrome flag a weak question?

Flag

Severity

Fires when

Too easy

Warning

80% or more of candidates answered correctly

Too hard

Warning

30% or fewer answered correctly

Negative discrimination

Critical

Weaker candidates answered correctly more often than stronger ones

Poor discrimination

Warning

The discrimination index falls below 0.2

Ineffective distractors

Warning

Under 40% of the wrong options attracted responses

Misclassified difficulty

Info

Assigned and observed difficulty differ by 3 points or more

Every flag carries a written recommendation. "Negative discrimination" reads: Low-performing students answer correctly more than high-performing ones. Possible wrong answer key or ambiguous wording.

Severity then sets the classification. A Critical flag makes the item Poor. Any other flag makes it Review. A clean item is Good. Negative discrimination is the only trigger that marks an item Poor on its own. A key error is the one fault worth stopping for.

How does distractor efficiency work?

Expand a multiple choice or true/false row and a sub-table opens: Label, Selection Count, Selection %, Is Correct, Functioning. Option labels run A, B, C in the order the author wrote them.

  • Selection share: every option gets the share of candidates who chose it, from the responses actually given.
  • The 5% rule: a wrong option chosen by fewer than 5% of candidates is marked non-functioning. It looks like a choice and behaves like decoration.
  • One number out: distractor efficiency is the share of wrong options still working. Four plausible wrong options score 100%. One dead option in four scores 75%.
  • Bank rollup: the bank-wide non-functioning list leaves true/false items out. A low score there means the item is easy, not that its options are broken.

What does the "Insufficient Data" badge mean?

A question with fewer than 10 responses wears an "Insufficient Data" badge. Its tooltip says so plainly: "Not enough responses to reliably analyze this question yet."

  • The reason is in the maths. Below ten responses the top and bottom groups are one or two people, so discrimination is noise.
  • The item is classified Not Rated, and its flags are discarded rather than shown. StudyDrome withholds a judgement the data cannot support.
  • Difficulty survives the gate. A simple proportion holds up at small numbers, so the p-value is still reported and averaged.
  • The gate applies twice: to the sitting, and to each question on its own count. A late question answered by a handful of candidates is suppressed inside a large exam.

How does calibration compare assigned and observed difficulty?

Authors assign a difficulty from 0 to 10 when they write a question. The cohort then reports its own. Observed difficulty is one minus the p-value, times ten, rounded. A p-value of 1.0 becomes 0. A p-value of 0.0 becomes 10.

The Calibration tab lists every question where the two disagree by 2 points or more. Its columns are Question Title, Current Difficulty, Observed PI, Suggested Difficulty and Reason. When the bank agrees with the cohort, the tab reads "No calibration suggestions found".

The Misclassified difficulty flag uses a wider gap. It fires at 3 points or more. The list is therefore more sensitive than the flag, on purpose. You can browse near-misses without each one touching a classification.

The author sets the new difficulty, and difficulty feeds difficulty-weighted scoring on written exams.

How do the flags fit a question review process?

  • Flags are advisory. Nothing leaves the bank, and no answer key changes, until a person decides it should.
  • Each flag names the fault and recommends an action, so the row hands a reviewer something to do.
  • The report covers one published test. Each sitting is judged on the responses it collected.
  • The Question Trends tab follows one source question across a course, keyed to survive the publish snapshot. Under two administrations it is labelled New.
  • OSCE checklist items get their own partial-credit report, with a facility index and an item-rest correlation. See [OSCE exams](/exam-manager/osce/).

Who uses item analysis software most?

  • Assessment leads defending a paper at a board. See [for assessment leads](/exam-manager/for/assessment-leads/).
  • Item writers checking whether a rewrite worked, one sitting later.
  • Exam offices sifting a large bank for the few items that need a human eye.
  • Course teams recalibrating difficulty so next year's paper lands where they aimed it.

Frequently asked questions

How does StudyDrome calculate the discrimination index?

Candidates are ranked by the points they earned. The top 27% and the bottom 27% form two groups, each holding at least one candidate. For a question, StudyDrome subtracts the bottom group's correct answers from the top group's. Dividing by the group size gives the index, from −1 to 1.

What does the "Insufficient Data" badge mean?

It means the question drew fewer than 10 responses. Its tooltip reads: "Not enough responses to reliably analyze this question yet." Below that line the top and bottom groups are one or two people, so discrimination is noise. StudyDrome classifies the item Not Rated and drops its flags. Difficulty is still reported.

How is distractor efficiency measured?

StudyDrome counts how many candidates picked each option, then turns that into a share of all responses. A wrong option chosen by fewer than 5% is marked non-functioning. Distractor efficiency is the share of wrong options still working, as a percentage. An item with no wrong options scores 100%.

Which questions get an item analysis?

Every question in the published test except Likert, which has no correct answer to score. Difficulty and discrimination are computed for the rest. The option-by-option distractor breakdown runs on multiple choice and true/false. OSCE checklist items get their own partial-credit report instead.

What happens to a question flagged as Poor?

It stays exactly where it is. Flags are advisory, so nothing leaves the bank and nothing blocks the next publication. The classification and its written recommendation go to the person who owns the item. They rewrite it, recalibrate its difficulty, or watch the next sitting.

Book a pilot

Bring one past paper and its responses. We will publish it, run the item analysis, and walk you through every flagged question.