A station metrics table has one row per station and four numbers that matter. What the station contributes to the exam's consistency. How well its checklist agrees with the examiner's global judgment. How much the score rises per grade step. And how many candidates failed. This page reads a worked table row by row and says what to do about each pattern. The thresholds come from AMEE Guide 49.
Which level are you reading?
Two different analyses are both called item analysis, and mixing them wastes a morning.
Station level, across the exam. Each station is one unit. You ask what that station adds to the whole exam and whether its scoring holds together. That is this page, and it is the table the exam board reads.
Item level, inside one station. Each checklist item is one unit. You ask whether an individual item separates strong candidates from weak ones. That is what the OSCE item-analysis calculator computes. It is the analysis you run after the station-level table flags a station.
The written-exam versions of these statistics behave differently. A question is scored right or wrong, and a checklist item is not. The complete guide to item analysis covers that family, and the answers on item analysis define each one.
The worked table
Twelve stations, one sitting, maximum checklist mark 30.
Station | Alpha if removed | R squared | Inter-grade discrimination | Failures |
|---|---|---|---|---|
1 Chest pain history | 0.71 | 0.70 | 3.1 | 4 |
2 Shoulder examination | 0.72 | 0.68 | 2.9 | 6 |
3 Breaking bad news | 0.79 | 0.40 | 1.2 | 2 |
4 Blood pressure | 0.73 | 0.62 | 3.4 | 1 |
5 Medication counseling | 0.71 | 0.66 | 2.8 | 19 |
6 Laboratory interpretation | 0.74 | 0.51 | 1.0 | 11 |
The exam's overall alpha is 0.74. Read the first column against that number and the second against 0.5. Read the third against a tenth of the maximum mark, which is 3.0 here.
Stations 1, 2 and 4 need no action. Three rows do.
A station that raises alpha when removed
Station 3 is the clearest signal in the table. The exam is more internally consistent without it. AMEE Guide 49 provides four explanations, each requiring a different response.
The station may measure a different construct. Breaking bad news is a plausible candidate because a communication station legitimately ranks candidates differently than a physical examination does. That is not a fault. It is a blueprint decision, and the answer is to record it rather than to fix the station.
Or the station may be poorly designed. Or the teaching may not have covered it. Or the examiners may not be working to a common standard. Those three are faults. The other columns usually say which one. Station 3 also has a low R-squared and weak discrimination, which points to the scoring, not the blueprint.
A low R squared
Station 3 sits at 0.40 and station 6 at 0.51. AMEE Guide 49 treats above 0.5 as a reasonable relationship. Its own worked examples contrast a good station at 0.697 with a poor one at 0.404.
A low value indicates that the checklist and the examiner's holistic judgment are diverging. Three causes are common. The checklist rewards behavior the examiner does not weigh heavily, which is a design problem. The global scale has no shared meaning, which is a training problem. Or the station is one where a checklist tracks performance poorly anyway. A systematic review compared checklists with global rating scales. The two were similar on agreement between raters. The global scale did better at consistency across items and stations, and at capturing expertise.
Station 6 is worth a second look for another reason. A laboratory interpretation station often has few observable steps, so a short checklist ends up close to all-or-nothing.
Discrimination that is too flat
Inter-grade discrimination is the slope: how far the checklist score climbs per step of the global scale. AMEE Guide 49 puts a healthy value at about a tenth of the maximum checklist mark. That is 3.0 on a 30-mark station.
Stations 3 and 6 return 1.2 and 1.0. The checklist barely separates a good pass from a borderline performance. Candidates whom the examiner thought were clearly different scored almost the same. That station cannot carry a defensible cut score, because the line falls where the data is flattest.
The usual fixes are on the marking scheme, not the exam. Fold several trivial items into one. Add a middle level between the extremes, so the scale has somewhere to land. Write anchors that say what each level looks like. The marking scheme page covers the wording.
A high failure count
Station 5 failed 19 candidates. The reflex is to call the station too hard, and AMEE Guide 49 warns against it directly. Review the teaching first.
A count like that has three plausible readings. The cohort was not taught the content, or not taught it in a form the station tests. The station is mis-marked; in that case, the other columns will usually be odd too, but here they are not. Or the station is correctly identifying something the program needs to fix.
Station 5 has healthy alpha, R squared, and discrimination. That is a station working as designed on material the cohort had not mastered. The finding belongs in the teaching report, not in a request to withdraw the station.
Plot the scatter every time
AMEE Guide 49 calls plotting a scatter graph of checklist score against global grade routine good practice, and it is the fastest check on the table. One glance shows range restriction, a grade nobody used, a cluster of candidates on one value, or a trend resting on two outliers.
Outliers matter less than people fear. The guide tested removing them: a near-zero checklist, a high checklist with a failing grade, a plainly wrong grade. The pass mark moved by less than 0.2 percent. The scatter is for understanding the station, not for cleaning the data.
Limits
Every number here describes one sitting of one cohort. A station that looks weak once is a question; a station that looks weak three times is a finding. Keep the tables so the third time is visible.
The table also cannot tell you whether the station was worth testing. A well-behaved station on a trivial task passes every check on this page. Content sits with the blueprint and the station review. No metric replaces it.
Frequently asked questions
What does it mean if alpha goes up when a station is removed?
That station ranks candidates differently from the rest of the exam. AMEE Guide 49 gives four reasons: it measures a different construct, it is poorly designed, the teaching did not cover it, or the examiners are not working to a common standard. The first is acceptable and is recorded; the other three are faults.
What is a good R-squared for an OSCE station?
Above 0.5 indicates a reasonable relationship between checklist scores and examiner global grades. AMEE Guide 49's worked examples contrast a good station at 0.697 with a poor one at 0.404. A low value points to the checklist design, examiner training, or a station where a checklist tracks performance poorly.
What is inter-grade discrimination in an OSCE?
The rise in checklist score per step of the global rating scale. A healthy value is about a tenth of the station's maximum checklist mark. A flat slope means the checklist does not separate performances that the examiner could tell apart, so the cut score sits where the data is weakest.
Does a high failure count mean the station is too hard?
Not on its own. AMEE Guide 49 says to review the teaching first. If the station's other metrics are healthy, it is usually working as designed on material the cohort had not mastered, which is a finding for the teaching report rather than a reason to withdraw it.
Is station-level analysis the same as checklist item analysis?
No. Station-level metrics treat each station as one unit and ask what it contributes to the exam. Item analysis treats each checklist item as a unit and asks whether that item separates candidates. Run the station table first, then the item analysis on whatever it flags.