Skip to content

How the reliability figures are computed

The conventions, written down

StudyDrome Exam Manager computes reliability with one of two coefficients, chosen from the paper itself, and reports the standard error of measurement beside it. Both use the sample variance. Both need at least two candidates and two scorable items: the score distribution and the bell curve run on the same responses.

How is the coefficient chosen?

Nobody picks the coefficient. The paper does, and the rule runs in a fixed order.

  1. Partial credit is checked first. If any item type on the paper can award part marks, the exam takes Cronbach's alpha, and the check stops there.
  2. Otherwise, every scorable item must be multiple choice or true/false. If they all are, the exam takes KR-20. If any other type is present, it takes alpha.

Likert items are removed from the scorable set before the rule runs. They collect opinion, so they carry no right answer to be consistent with.

The report names the coefficient it used, next to the number. Exam reliability covers what the figure then means.

What are the formulas?

Both coefficients are the classical internal-consistency formulas, unmodified.

KR-20 = (k ÷ (k − 1)) × (1 − Σpq ÷ σ²)

Here k is the number of scorable items. p is the share of candidates answering an item correctly and q the share answering it wrong. σ² is the variance of the candidates' total scores.

Alpha = (k ÷ (k − 1)) × (1 − Σσᵢ² ÷ σₜ²)

Here σᵢ² is the variance of the scores on one item and σₜ² the variance of the totals. Where every item is right or wrong, alpha reduces to KR-20. That is why one engine serves both.

SEM = SD × √(1 − reliability)

The standard error of measurement takes the spread of the scores and the coefficient. The shared helper reports it to four decimal places. For written papers, it is computed as a percentage score and shown to two decimal places.

Which variance is used?

This is the line a reviewer needs, and almost nobody publishes.

Every variance in those formulas is the sample variance — the form that divides by one less than the count. It applies to the item variances, the variance of the totals, and the standard deviation from which the error band is built.

Recompute one of our figures using the population form, and you will end up slightly below it. The gap is the convention, not an error in either calculation. Ours is written down so you can match it.

Reliability itself is reported on a 0-1 scale.

What does a coefficient need?

A coefficient describes agreement between items across candidates, so it needs enough of both.

Requirement

Why it exists

At least two candidates

One candidate produces no variance to compare against

At least two scorable items

Internal consistency is agreement between items

Equal-length response rows

Every candidate is measured on the same set

Some variation in the totals

A cohort that all scored the same has no consistency to measure

Item-level statistics carry their own minimums, and what an assessment lead sees lists them.

Reliability also reduces each candidate to a single attempt before the matrix is built, so retakes cannot lift it. Exam reliability sets out exactly which figures that rule covers.

What does the distribution report compute?

The distribution report describes the cohort rather than any one candidate.

Figure

Detail

Mean, median, minimum, maximum

The center and the range of the scores

Standard deviation

Sample form, as above

Pass rate

Candidates who passed, as a percentage, to two decimals

Average item difficulty

The mean difficulty across the paper

Skewness

Fisher's, with the small-sample adjustment, to four decimals

Histogram

Ten bands across the percentage range, the last one taking a full 100%

Skewness is the figure worth reading. The report turns it into one sentence of its own: approximately symmetric, right-skewed with many low scores and a paper that may be too difficult, or left-skewed with many high scores and a paper that may be too easy.

Under half, treat the shape as symmetric. Between half and one, moderately skewed. Above one, strongly skewed. Skewness needs at least three candidates and some spread in the scores.

How is the bell curve drawn?

The bell curve overlays a normal curve on the cohort's histogram, on a shared count axis.

The curve is fitted to the mean and standard deviation of the scores actually earned. Each bar then gets an expected count from that fitted curve, so the observed shape and the theoretical one can be compared without a second axis.

Setting

Detail

Bands

Twelve by default, adjustable between four and fifty

Range drawn

The mean, plus and minus three standard deviations

Basis

Points as they stand, or percentages handled as a proportion

Failing region

Shaded at the pass mark

Pass and fail are always determined by the total score, regardless of the basis you use.

The curve is drawn over the scores. The scores are not changed by it. It is a picture of the cohort you had, offered next to the results rather than applied to them.

Notes for a reviewer

A statistic is only as useful as the conventions behind it. These are the ones that decide whether your number and ours agree.

  • The formulas are the classical ones, published here in full, so any figure can be rebuilt.

  • Every variance is the sample form, including the one the error band is built from.

  • The coefficient is chosen from the paper's own item types, not from a menu.

  • Reliability is reported on a 0 to 1 scale, with the method named beside it.

  • The distribution statistics describe a cohort. They never alter a mark.

  • Every figure is computed from the published exam, so a re-run reproduces it.

Frequently asked questions

Which variance convention do you use?

The sample form throughout, dividing by one less than the count. It applies to the item variances, to the variance of the total scores, and to the standard deviation behind the error band. If you rebuild one of our figures using the population form, you will end up slightly below it, and the difference is due to convention rather than a fault in either sum.

What is the smallest cohort that produces a reliability figure?

Two candidates and two scorable items, with response rows of the same length and some variation in the totals. Those are the conditions for internal consistency to exist at all. Skewness needs three candidates and some spread. Small cohorts satisfy the arithmetic long before they satisfy a board, so read a figure from twelve candidates as a hint.

Does StudyDrome curve or scale exam scores?

Scores stay as they were earned. The bell curve fits a normal distribution to the cohort's mean and standard deviation, then overlays it on the observed histogram so the two shapes can be compared. It is a picture offered next to the results. Pass and fail are decided on the total score, whichever basis you display.

Can I reproduce your reliability number in my own tool?

That is what this page is for. Take the item scores from the results export, apply the formula in the report names, and use the sample variance throughout. Reduce each candidate to their latest attempt first, since that is what the coefficient runs on. The figures should then agree to the decimal place we report.

Book a pilot

Send us a past paper and the scores it produced. We will run it in a sandbox, and you can reconcile every figure against your own calculation.