Exam reliability in StudyDrome

The coefficient is chosen for you

StudyDrome Exam Manager computes reliability for every published exam. It picks KR-20 or Cronbach's alpha from the paper itself. The standard error of measurement sits beside it. Both coefficients report on a 0-to-1 scale. A five-tier verdict reads the number, and 0.70 or above is acceptable.

How does StudyDrome choose between KR-20 and Cronbach's alpha?

  1. 1
    The paper decides, not the user. Nobody picks a formula from a dropdown.
  2. 2
    Likert items drop out first. They are excluded from scoring by design, so they never affect the choice.
  3. 3
    The remaining scorable items are then read. KR-20 needs every one of them to be multiple choice or true/false.
  4. 4
    Partial credit is read too. Checkboxes, matching, sequence and fill in the blank all allow it.
  5. 5
    All multiple choice or true/false, and no partial credit anywhere? KR-20 runs on a right/wrong matrix. Anything else routes to Cronbach's alpha, on the earned-points matrix.
  6. 6
    The report names the method it used. You never read a coefficient without seeing whether it is KR-20 or Cronbach's Alpha.

Which papers get KR-20, and which get Cronbach's alpha?

Essay, short answer and file upload papers always route to Cronbach's alpha. Those items are marked on points, so a right/wrong matrix would throw information away.

What the paper contains

Coefficient used

Multiple choice and true/false only

KR-20

Any checkboxes, matching, sequence or fill in the blank

Cronbach's alpha

Any essay, short answer or file upload

Cronbach's alpha

KR-20 scores each response as right or wrong. Alpha uses the points the response earned. KR-20 is the right/wrong case of alpha, which is why one engine serves both. For the difference in full, read KR-20 versus Cronbach's alpha.

What does a reliability coefficient of 0.82 mean?

The report does not hand you a bare number. Every coefficient carries one of five verdicts, rendered word for word from the API:

Coefficient

Verdict

0.90 and above

Excellent — suitable for high-stakes decisions

0.80 to 0.89

Good — suitable for most assessment purposes

0.70 to 0.79

Acceptable — adequate for classroom testing

0.60 to 0.69

Questionable — consider revision

Below 0.60

Poor — significant measurement error

A paper at 0.82 therefore reads Good — suitable for most assessment purposes. The bands are fixed in code, so the same number reads the same way for every exam and every reviewer.

For how assessment teams use those bands, see what counts as a good KR-20 value, or the complete guide to item analysis.

What does the report tell you to do next?

A weak coefficient is only useful if it points somewhere. The report attaches actions, drawn from this paper rather than from general advice:

  • Under 0.70: it suggests increasing the item count. A longer paper measures the same trait more steadily.
  • Short papers: on fewer than 20 scorable questions it warns that the item count may be limiting reliability, and names the count you have.
  • Negative discrimination: when strong candidates get an item wrong more often than weak ones, the report asks you to remove or revise it, and gives the number found.
  • The recount: that number is worked out from the top and bottom 27% of the cohort, on each candidate's latest attempt.

What is the standard error of measurement?

Reliability describes the paper. The standard error of measurement describes one candidate's score.

StudyDrome computes it from the spread of the scores and the reliability coefficient: the standard deviation multiplied by the square root of one minus reliability. On written exams it runs on percentage scores and reports two decimal places. On OSCEs it runs twice — once per station on earned scores, and once across the exam on candidate totals.

Read it as the expected error band around a single score. The tooltip on screen says it plainly: Estimated error margin around each student's score. Lower is more precise.

This is the number an exam board argues about at the pass mark. A candidate two marks under the cut, on a paper with an SEM of three, is a conversation. The same candidate on a paper with an SEM of one is not. Pair the statistic with standard setting and with item analysis.

Why does reliability count each candidate once?

Retakes correlate with themselves. Ten students sitting a paper three times are not thirty independent candidates.

  • Before the item-score matrix is built, every student is reduced to their latest attempt.
  • Duplicate, correlated rows therefore cannot lift internal consistency. A retake-heavy cohort reports the reliability it earned.
  • The same latest-attempt rule applies when the report recounts negatively discriminating items for its recommendation.
  • Item analysis, exam difficulty and cohort comparison sample attempts in their own way. Reliability is the one that deduplicates.

Where do you find reliability on screen?

In Cohort Analytics, on the Exam & Group Performance tile. Its own description reads: analyze exam difficulty, reliability, and group performance.

  • The reliability tab carries the coefficient, the method name beside it, the standard error of measurement and the recommendations.
  • The tooltip states the rule in one line: internal consistency of the exam, on a 0 to 1 scale, with 0.7 and above acceptable.
  • OSCE papers report the same SEM statistic on the OSCE item analysis screen, per station and across the exam.

How should you read the coefficient itself?

  • Both coefficients are reported on a 0-to-1 scale, so one reading scale covers every paper.
  • The method name travels with the number, which keeps KR-20 and alpha from being read as the same statistic.
  • A coefficient describes one published test and one cohort. It belongs to the sitting, not to the item bank.
  • The maths needs at least two candidates and at least two scorable items, and it needs real spread in the totals.
  • Reliability is one line of evidence. Read it beside item statistics and the score distribution.

Who needs this most?

A coefficient is a footnote until someone questions a result. Then it is the case.

  • Assessment leads signing a paper off for high-stakes use.
  • Exam offices asked to show that a pass list is defensible.
  • Item writers who want to know which questions held the paper back.

Frequently asked questions

How does StudyDrome decide which reliability coefficient to use?

The paper decides, not the user. Likert items drop out first, because they carry no score. If every remaining scorable item is multiple choice or true/false, and none of them allows partial credit, KR-20 runs. Everything else routes to Cronbach's alpha: essay, short answer, file upload, checkboxes, matching, sequence and fill in the blank.

What reliability score should we aim for?

It depends on the stakes. The report reads 0.90 and above as excellent and suitable for high-stakes decisions. From 0.80 it reads good and suitable for most assessment purposes. From 0.70 it reads acceptable and adequate for classroom testing. Below 0.70 it asks for revision. A weekly progress quiz and a final paper are held to different bars.

Can retakes inflate our reliability figure?

The calculation counts each student once. Before the item-score matrix is built, every student is reduced to their latest attempt, so correlated duplicate rows cannot lift internal consistency. The same latest-attempt rule applies when the report recounts negatively discriminating items. The rule is written into the reliability calculation itself, which is why it holds on every cohort.

What does the report suggest when reliability is low?

It names actions rather than leaving you with a number. Under 0.70 it suggests increasing the item count. On a paper with fewer than 20 scorable questions it warns that the count may be limiting reliability. Where items discriminate negatively it asks you to remove or revise them, and says how many it found.

Does reliability cover OSCE exams as well as written papers?

Yes. Written papers report the coefficient and the standard error of measurement together in Cohort Analytics. OSCE papers report SEM at two levels: per station, on earned scores, and across the exam, on candidate totals. Both exam types draw from the same item bank and run through the same statistics engine, so the numbers mean the same thing.

Book a pilot

Bring one past paper and its results. We will publish it, run the reliability report, and read the coefficient, the SEM and the recommendations with you.