Skip to content

OSCE analytics in StudyDrome: item analysis, examiner reports and the candidate scorecard

Last updated · 6 min read

The quality section asks for four things after a sitting. Statistics per station and per item. A comparison across examiners. A record of who marked what. And something to hand a candidate. StudyDrome Exam Manager produces all four from the scores that the examiners have already entered on the day marking closes. This page says what each report contains and what it deliberately does not claim.

The product overview is on which report answers your question. The wider OSCE model is on OSCE exams in StudyDrome.

Checklist item analysis

The OSCE item analysis is the partial-credit version of the written-exam report. A checklist item is scored on a scale rather than as right or wrong. It runs on one published exam and produces three numbers per item.

Facility is the mean score as a proportion of the item's maximum, from 0 to 1, with bands: 0.8 or above is Easy, 0.3 or below is Hard, anything between is Moderate. Higher means easier.

Discrimination compares the top and bottom third of candidates at that station. It divides the gap between their mean scores by the item maximum. The fraction is 33 percent for OSCE items. The written-exam analysis uses 27. It returns nothing when the groups would be smaller than three candidates.

Item-total correlation is the corrected version. The item is removed from the station total before correlation, so a short station cannot inflate its own items. It returns nothing for fewer than three candidates or when the scores have no variance.

The report then rolls up. Station reliability is Cronbach's alpha across that station's items. Exam reliability is alpha across the station totals. Both need at least two complete cases. Where there is insufficient data, the report states so rather than printing a number.

Two rules decide who counts. A blank score means not assessed, and that candidate is excluded from the item's denominator. A scored zero is a real score and stays in. Sessions still in progress are excluded entirely, so partial marking cannot skew facility.

The screen heads with candidates, mean, standard deviation, standard error of measurement, reliability, item count, difficulty, and quality. Below it sits a table of items with their band and classification. Exports go to Excel and PDF.

What that covers, and what the report reads as

Match this against the station metrics page. The mapping is direct at item level and partial at station level. Facility, discrimination, and the item-total correlation are the item-level analysis. That is the layer you reach for once a station is flagged. Alpha across station totals is the exam-level figure the quality report opens with.

Two metrics are intentionally missing from this list. R squared and the fitted line come out of the standard-setting calculation, so they sit on the standard-setting page.

The analysis is scoped to one published exam. Comparisons across sittings are made from the exports. That is where the methods on beyond alpha start too. Reliability here is always Cronbach's alpha. The written-exam coefficient selector does not apply to OSCEs. The discrimination figure is a top and bottom group statistic, not a slope from a measurement model. There is no decision study.

Examiner activity and workload

A date range in, a per-examiner picture out. It covers exams, stations, candidates marked, and whether each session was complete, partial, or never started. Two-hour figures come with it. On-duty hours are the headline. It spans each session from start to completion. Active hours sum only the per-candidate assessment times.

The report defends itself against its own worst input. Pause durations are not stored. So a session that was paused or reopened can read absurdly long. Any span over 16 hours is flagged, as is one more than six times the measured active time. The total then reverts to active hours, with a note indicating so. Where a session was never formally completed, the end time comes from the last candidate marked. That is noted too.

The product states the limit itself, and the wording matters: worked hours are a derived estimate, not a clocked timesheet. Between-candidate waiting counts as work, because the examiner is on station. The activity counts are exact, whatever the timing quality. Use it for workload planning and for the examiner record a review file wants. Do not use it for payroll without that note attached.

Examiner leniency and agreement

This is the report the examiner and site effects page leads to. It pivots completed sessions into an examiner-centered marksheet and adds two figures.

Leniency is an examiner's mean candidate score at a station, minus that station's overall mean across all examiners. It is shown in percentage points, to one decimal. A positive number means more lenient. It is computed only where two or more examiners staffed the station. A sole examiner has nothing to compare against.

Agreement applies where two or more examiners scored the same candidate. It is the mean absolute percentage difference between them.

Both ship with a caveat: the product prints on screen, in the PDF, and in the Excel. Quote it rather than soften it. The metric is unadjusted. Examiners are not randomly assigned candidates. A difference may reflect the ability of the candidates an examiner happened to see, rather than that examiner's stringency. It is a screening signal for review, not a verdict on an examiner.

That is a mean absolute percentage difference and nothing more. It is not a chance-corrected agreement statistic. No report here scores an examiner's calibration. A conversation with an examiner starts at the station and with the candidates, never with the number.

The candidate scorecard

One candidate, one sitting. It shows every station, every checklist item score, the global rating and comments, and the pass/fail status, with any critical failures marked. It exports to PDF or Excel for that person alone, in one of two layouts.

It is reached from the OSCE student results report rather than from a tile. It is a manager screen. Candidates receive it because someone sends it. That is the model the candidates section assumes. How results and feedback reach candidates is covered in the next section of this guide.

What stays with the committee

The reports produce the numbers. Reading them is still the psychometrician's job. Three judgments never leave the room. Whether a flagged station is a design fault or a blueprint decision. Whether an examiner difference is worth acting on. And whether anything found changes a result. The quality report is a document your committee signs. This is the arithmetic underneath it.

Was this page helpful?