OSCE Item-Analysis Calculator

Free · no signup · your data stays in your browser

OSCE Item-Analysis Calculator

Paste your station's examiner-scored checklist results and get the polytomous item analysis OSCE units normally only see in commercial exam software — facility, discrimination, item-rest correlation, Cronbach's α, SEM — with every flagged item explained and cited. Computed entirely in your browser: candidate data never uploads anywhere.

Load your station's scores

Paste from Excel / Sheets: item names in the first row, one row per candidate. Optional: a candidate-ID first column, a “Max” row with per-item maximums, a “Total” column.

Your data is parsed and analyzed locally — open your browser's network tab to verify nothing is uploaded.

Methodology

Every statistic mirrors the OSCE item-analysis engine inside StudyDrome — the implementation was ported from it and is regression-tested against the product's own test fixtures — so the numbers here are the numbers the exam software would report. Definitions and citations per metric:

Facility
Mean awarded score as a proportion of the item maximum (the polytomous difficulty index). Bands: ≥ 0.80 Easy, ≤ 0.30 Hard, otherwise Moderate. (Pell et al. 2010, How to measure the quality of the OSCE — AMEE Guide 49)
Discrimination
Mean item score of the top third of candidates minus the bottom third, divided by the item maximum — candidates ranked by station total, group size ⌊N × 0.33⌋ (shown as “—” when a group would have fewer than 3 candidates). Below 0.20 is flagged; below 0 is critical. (Pell et al. 2010, How to measure the quality of the OSCE — AMEE Guide 49)
Item-rest correlation
Pearson correlation between an item's score and the candidate's total over the other items (total minus this item, avoiding self-inflation on short checklists). Below 0.20 is flagged. (Pell et al. 2010, How to measure the quality of the OSCE — AMEE Guide 49)
Cronbach's α
Internal consistency of the station across checklist items, computed over candidates scored on every item. Pell suggests interpreting station-level α alongside R² from borderline regression rather than in isolation. (Pell et al. 2010, How to measure the quality of the OSCE — AMEE Guide 49)
SEM
Standard error of measurement = SD × √(1 − α), in station points — the error bar around any candidate's observed station score. (Pell et al. 2010, How to measure the quality of the OSCE — AMEE Guide 49)

These are structural screening statistics: they find items that behave strangely, not items that measure the wrong thing. Judging clinical relevance and content validity still needs your examiners and blueprint — and small cohorts make every statistic noisy; treat single-administration flags as prompts to look, not verdicts.

Every flag and band, explained

The exact rules the analysis applies — thresholds, severities and what to do about each. Every example below is executed against the engine by the test suite, so this reference cannot drift from what the tool actually does.

F1Too easy

Warning

Facility ≥ 0.80 — candidates averaged 80%+ of the item's maximum mark.

Example: On "Introduces self and confirms identity", 15 of 16 candidates score the full 2 points: facility 0.97.

What to do: 80%+ of the maximum mark was achieved on average. Item may not discriminate; consider reviewing.

Pell et al. 2010, AMEE Guide 49

F2Too hard

Warning

Facility ≤ 0.30 — the average score is below 30% of the item's maximum.

Example: Only 4 of 16 candidates score anything on a 2-point item: facility 0.13. Is the task unclear, or were examiners marking to different standards?

What to do: Average score is below 30% of the maximum. Verify the item wording, scoring scale and examiner calibration.

Pell et al. 2010, AMEE Guide 49

F3Negative discrimination

Critical

Discrimination below 0 — the bottom third of candidates outscored the top third on this item.

Example: The five strongest candidates average 0 on the item while the five weakest average 2: discrimination −1.0. Classic signs: a reversed scale, a miskeyed checklist row, or examiners interpreting the item differently.

What to do: Lower-performing candidates scored higher than top performers on this item. Possible scoring or wording problem.

Pell et al. 2010, AMEE Guide 49

F4Poor discrimination

Warning

Discrimination in [0, 0.20) — exactly 0.20 is not flagged. The item barely separates strong from weak candidates.

Example: Top-third and bottom-third candidates average almost the same score: discrimination 0.10.

What to do: Item does not separate strong from weak candidates effectively. Consider revision.

Pell et al. 2010, AMEE Guide 49

F5Low item-total correlation

Warning

Item-rest correlation below 0.20 — exactly 0.20 is not flagged. The item's scores barely track the rest of the station.

Example: Scores on the item are unrelated to how candidates did on everything else (r = 0.05) — it may be measuring a different construct, or examiner noise.

What to do: Item correlates weakly with the rest of the station score; it may be measuring something different.

Pell et al. 2010, AMEE Guide 49

B1Easy band

Info

Facility ≥ 0.80 (0.80 itself is Easy).

Example: Facility 0.80 lands in Easy; 0.79 is Moderate.

What to do: A few easy items are fine — they reassure candidates and catch absolute non-performance — but an easy-heavy station wastes testing time.

Khan et al. 2013, AMEE Guide 81 Part II

B2Moderate band

Info

Facility between 0.30 and 0.80, both exclusive.

Example: Facility 0.55 is Moderate.

What to do: The productive zone — moderate items carry most of the measurement information.

Khan et al. 2013, AMEE Guide 81 Part II

B3Hard band

Info

Facility ≤ 0.30 (0.30 itself is Hard).

Example: Facility 0.30 lands in Hard; 0.31 is Moderate.

What to do: Check hard items before blaming candidates: unclear task wording and uncalibrated examiners produce the same numbers as a genuinely difficult skill.

Khan et al. 2013, AMEE Guide 81 Part II

N1Discrimination needs ~10+ candidates

Info

Discrimination compares the top and bottom 33% and needs at least 3 candidates in each group: with fewer than 10 candidates, group size floor(N × 0.33) drops below 3 and the metric shows “—” instead of a misleading number.

Example: With 8 candidates, group size is floor(8 × 0.33) = 2, so discrimination is not computed.

What to do: Don't read discrimination on tiny cohorts — collect more administrations before acting on it.

Daniels & Pugh 2018, Twelve tips for developing an OSCE

N2Item-rest needs variation

Info

The item-rest correlation needs at least 3 scored candidates AND variation on both sides — if every candidate got the same item score (or the same rest score), it shows “—”.

Example: Every candidate scored 2 on the item: no variance, so no correlation exists.

What to do: A dash here on an item everyone scored identically is expected, not a data problem.

N3Reliability needs complete cases

Info

Cronbach's α uses only candidates scored on every item, and needs at least 2 such candidates and 2 items — otherwise the report says “Insufficient data”.

Example: One complete-case candidate is not enough to estimate internal consistency.

What to do: Blank cells (not-assessed) remove that candidate from the α calculation only; their scored items still count everywhere else.

Pell et al. 2010, AMEE Guide 49

References

Share this tool

Frequently asked questions

Does my candidate data get uploaded anywhere?

No. Parsing and every statistic run in your browser. The only network calls this page makes are anonymous aggregate counters (which tool, which action) and — only if you use it — the export form, which sends your email address and nothing else. You can verify this in your browser's network tab.

Are these the same numbers StudyDrome reports?

Yes — the engine was ported from StudyDrome's OSCE item-analysis calculator and is regression-tested against the product's own test fixtures, including its rounding behavior. Facility, discrimination, item-rest correlation, Cronbach's α, SEM, the flag thresholds and the Good/Review/Poor classification all match.

Why does discrimination show “—” for my station?

Discrimination compares the top and bottom 33% of candidates and needs at least 3 candidates in each group, so it requires roughly 10+ candidates. With fewer, showing a number would be misleading — the product suppresses it the same way. Item-rest correlation likewise shows “—” when fewer than 3 candidates were scored on the item or when scores don't vary.

What input format do I need?

A simple matrix: item names in the first row, one row per candidate, one score column per checklist item. Optional extras are auto-detected: a candidate-ID first column, a “Max” row under the header with per-item maximum scores, and a station-total column (used to rank candidates for discrimination). Paste from Excel/Sheets or upload CSV/XLSX. Blank cells mean not assessed; zeros are real scores.

Is this really free?

Yes. The analysis is free with no account; the XLSX/PDF export asks only for an email address. The tool exists so you can see the kind of OSCE analytics StudyDrome produces automatically on every exam.