Cronbach's alpha answers one question: how consistently this exam ranked these candidates. It cannot say which part of the exam caused the inconsistency. Nor how many more stations would fix it. Nor whether this year's stations behaved like last year's. Three method families answer those, and none of them is needed for a routine sitting. This page explains what they do and when a school should reach for one.
What alpha leaves out
Alpha gives a single number for the entire exam. Underneath it sit several separate sources of error: the stations sampled, the examiners who marked, the simulated patients, the day, the circuit. A single coefficient blends them, so two exams with the same alpha can have completely different problems.
It also moves with length. Add stations of equal quality and alpha rises. That is why a long exam of ordinary stations can post a comfortable figure. The foundations page covers how to report it honestly.
The scale of the gap is visible in the published record. A systematic review of OSCE reliability pooled 39 studies and 188 coefficients. The mean alpha across stations was 0.66. Within a station, across its items, it was 0.78. And the mean generalizability coefficient, which accounts for multiple sources of error simultaneously, was 0.49.
Those three numbers describe the same exams. The most flattering one is measured inside a station. There, candidates are scored on the same task by the same person. The least flattering accounts for the ways an OSCE can vary. Neither is wrong; they answer different questions.
Generalizability theory: naming the sources of error
Generalizability theory extends classical reliability by separating these sources rather than pooling them. Bloch and Norman's AMEE guide calls it a generalization of classical reliability theory. In it, the various sources of error that make a measurement inaccurate are explored. They present it as a tool for judging the quality of an assessment method and improving its precision.
In practice, it answers questions a single coefficient cannot. How much of the variation comes from stations, and how much from examiners? Would another examiner per station help more than another station? Is the exam's weakness in what it samples or in who marks it?
For an OSCE, the usual finding is that stations dominate, because performance is case-specific. That is the statistical form of the advice to add stations before polishing any one of them.
The decision study: what would make it better
The second half is the more useful one for a school. Having estimated where the error comes from, a decision study projects what would happen to reliability under a different design: more stations, fewer stations, two examiners instead of one, a longer station.
It converts an argument into a number. "We should add stations" becomes something sharper. Four more stations of this type move the coefficient from 0.61 to 0.72. A second examiner at every station does less. That is a budget conversation the cost model can take.
Item response theory: comparing across years
The third family puts each item and each candidate on a common scale. Performances from different sittings can then be compared, even when the stations are not identical. Its OSCE use is mostly about equivalence over time.
A longitudinal study of a ten-station OSCE used it exactly that way. Six stations were reused one year apart, with 191 candidates in the first year and 236 in the second. Generalizability passed 0.70 in both years, and the test characteristics were equivalent. The authors concluded that, with careful security measures, high-stakes OSCEs may use the same stations. That is a finding with a direct operational consequence for the station bank.
When should a school call a psychometrician?
Not for a routine sitting. The metrics in the quality report answer the ordinary questions. A school that runs them every time is already ahead of most.
Five triggers are worth the call.
- The exam is about to become high stakes, or a regulator is about to inspect it.
- You are changing the design: adding or dropping stations, moving to two examiners, splitting across sites.
- Reliability has been poor for two cycles, and the station-level work has not moved it.
- You are making claims across years, such as comparing cohorts or reusing stations under an equivalence argument.
- A result is being challenged in a way that needs an expert to answer.
Ask for two things when you do. A written report the committee can read without the mathematics, and the model stated plainly enough that someone else could repeat it.
A note on software
These are analyses, not features. An exam platform computes the routine metrics. The methods on this page are normally run by a specialist, in dedicated statistical software, on the exported score file. Assume nothing here is a button unless a vendor shows you the output. And do not put these terms in a requirements list without knowing why you want them. The bridge page is explicit about which analyses are in the product.
Limits
Better statistics do not rescue a weak exam. A decision study on twelve poor stations tells you how many poor stations you would need, which is not the answer anyone wants.
These methods also require sufficient data volume and a stable design to be worth running. A school with 25 candidates, one circuit, and a changing-station set will get estimates that are too wide to act on. The ordinary metrics, read across several sittings, serve it better.
Frequently asked questions
What does Cronbach's alpha not tell you about an OSCE?
Which source of error caused the inconsistency, how many more stations would fix it, and whether this year's exam behaved like last year's. It blends stations, examiners, patients, and circuits into one number, and it rises as the exam gets longer.
What is generalizability theory used for in an OSCE?
Separating the sources of error rather than pooling them, so you can see how much variation comes from stations, examiners, and other factors. It is described as a generalization of classical reliability theory and as a way to judge and improve the precision of an assessment.
What is a decision study?
A projection of what reliability would be under a different design: more stations, fewer, a second examiner, a longer station. It turns a design argument into a number that a budget discussion can use.
Why do published OSCE reliability figures differ so much?
Because they measure different things. A review of 39 studies found a mean alpha of 0.66 across stations, 0.78 within a station across its items, and a mean generalizability coefficient of 0.49. The last accounts for more sources of error, so it is lower.
When does a school need a psychometrician?
When the exam becomes high-stakes or faces inspection, when the design is changing, when two cycles of station-level work have not improved reliability, when claims are being made across years, or when a result is challenged in a way that requires expert evidence.