---
title: "The post-OSCE quality report: the metrics, their thresholds and who signs it off"
description: "What belongs in an OSCE quality report: seven metrics with the thresholds AMEE Guide 49 sets, who prepares it, and who signs it off."
canonical: https://studydrome.com/docs/osce/quality/the-post-osce-quality-report/
updated: 2026-10-10
---

# The post-OSCE quality report: the metrics, their thresholds and who signs it off

What belongs in an OSCE quality report: seven metrics with the thresholds AMEE Guide 49 sets, who prepares it, and who signs it off.

An OSCE quality report is one document produced after every sitting. It carries the same metrics in the same order every time, each against a published threshold. The exam board can then see, on one page, whether the exam worked. AMEE Guide 49 sets out the metric family and the values against which to read them. This page turns that into a report with an owner, a deadline and a signature.

## Why a standing report and not an analysis

An analysis is done when someone is worried. A report is done every time. The difference matters for two reasons.

The first is comparison. A metric means little on its own and a lot next to the same metric from the last three sittings. A station whose failure count doubled is visible only if last year's count was written down in the same place.

The second is evidence. A quality review asks what the school does routinely, not what it did once. A report that exists for every sitting, with the same headings, answers that question by existing. The [file a reviewer asks for](/docs/osce/quality/evidence-for-accreditors/) is largely a run of these reports.

## What goes in it

| Metric | What it says | Threshold to read it against |
| --- | --- | --- |
| Cronbach's alpha across stations | How consistently the whole exam ranks candidates | 0.7 or above is usually regarded as acceptable; above 0.9 may mean the stations are redundant |
| Alpha with one station removed | What one station contributes | It should sit below the overall alpha; a station that raises it is measuring something else |
| R squared of checklist on global grade | Whether the checklist and the examiner's holistic judgment agree | Above 0.5 is a reasonable relationship |
| Inter-grade discrimination | How much the checklist score rises per step of the global scale | About a tenth of the station's maximum mark |
| Number of failures per station | Where candidates did not reach the mark | No fixed value; a high count is a question about teaching first, not proof of a hard station |
| Between-group variation | How much of the score variance comes from the circuit, site, or examiner group | Ideally under 30 percent; over 40 percent should give cause for concern |
| Simulated patient ratings | The patient's own view of the encounter | More than 10 percent adverse ratings flags a station |

Two additions make the page readable. Put last year's value in a second column. Give every flagged station a named owner and a date rather than a comment.

The report says what the numbers are. What each pattern means and what to do about it are on the [reading station metrics](/docs/osce/quality/reading-station-metrics/).

## Who prepares it and who signs it

The psychometrician prepares it, or the person playing that role. In a small school, that is often the assessment lead with a spreadsheet. That is fine, as long as the method is written down and repeated.

The assessment lead signs that the report is complete. The exam board receives it with the results, not after them, because two of its decisions depend on it: whether to withdraw a station, and whether the computed pass mark stands. The [ratification page](/docs/osce/scoring/results-ratification-and-release/) covers that meeting.

One rule keeps the report honest. It is produced before the pass list is circulated. Its flags are recorded regardless of whether the board acts on them. A metric that appears only when it supports a decision is not a metric.

## When is it produced?

Same day as the score file closes, or as near as the marking window allows. Examiner memory is the perishable input. A station flagged three months later gets a shrug. A station flagged while the examiner remembers the morning gets an explanation.

Three cadences work together. Per sitting, the report goes to the board. Per year, the reports go to the assessment committee as a trend, which is where [closing the loop](/docs/osce/quality/closing-the-loop/) starts. Per cycle, the entire run is added to the review file.

## Does this work for a small cohort?

Yes, with wider error bars and more caution. A published review of a chiropractic program ran these metrics on one sitting of 24 examinees. It reported what small cohorts often show. Cut scores above the traditional 50 percent pass, roughly 59 to 68 percent. Failure counts range from 0 to 16.7 percent, except for one laboratory-analysis station at 50 percent. R-squared values range from 0.67 to 0.97. And between-group variation from 67.3 to 95.9 percent, far above the level that should worry anyone.

The authors did not respond by dropping the metrics. They read the pattern as missed teaching, stations open to examiner interpretation, and inconsistent marking. Their recommendations were more examiner training, a global scale written for each station, and the same review next time. That is the right use of a report on a small cohort. It points at work, not at a verdict.

## Limits

The thresholds are conventions, not laws. Alpha rises with the number of stations, so a long exam of weak stations can post a respectable figure. That is why the report carries station-level metrics beside it. A value inside its threshold is not proof that a station is sound, and a value outside it is a question rather than an answer. The [foundations page](/docs/osce/foundations/validity-and-reliability/) covers what the coefficients can and cannot support.

The report also says nothing about content. A perfectly consistent exam can sample the wrong things. That check belongs to the blueprint, and the report should be read next to it.

## Frequently asked questions

### What should an OSCE quality report contain?

Cronbach's alpha across stations, the alpha with each station removed, R-squared of the checklist on global grade, inter-grade discrimination, failure counts, between-group variation, and simulated patient ratings. Each with its threshold, last year's value, and a named owner for anything flagged.

### What is an acceptable Cronbach's alpha for an OSCE?

AMEE Guide 49 treats 0.7 or above as usually acceptable, and warns that a value above 0.9 may indicate redundancy between stations. It also notes that 0.9 is often quoted but unobtainable at the local level. Read it alongside the station-level metrics, not alone.

### Who writes the OSCE quality report?

The psychometrician, or the assessment lead playing that role in a smaller school. The assessment lead signs that it is complete. The exam board receives it together with the results, because its decisions on withdrawing a station and accepting the pass mark depend on it.

### When should the OSCE quality report be produced?

As soon as the score file closes, while examiners still remember the day. It goes to the board with the results, then into an annual trend review, then into the file an accreditation visit reads.

### Can a small school run these metrics?

Yes. A published review ran the full metric set on one sitting of 24 examinees. The values are less stable, and between-group variation, in particular, can be very high, so the report is read as a list of questions to investigate rather than as a verdict on the exam.

## Sources

**Sources**

1. [Pell G, Fuller R, Homer M, Roberts T. How to measure the quality of the OSCE: a review of metrics. AMEE Guide No. 49](https://doi.org/10.3109/0142159X.2010.507716) (Medical Teacher, 2010)
2. [Cade AE, Meuller N. Measuring the quality of the OSCE in a chiropractic programme: a review of metrics and recommendations](https://doi.org/10.7899/JCE-22-29) (Journal of Chiropractic Education, 2024)

> [!TIP]
> **In practice**
> Most of this report is arithmetic over the score file, which is why software can produce it the same day. To see which of these metrics a platform computes and which stay with the committee, read /docs/osce/quality/in-studydrome-osce-analytics/.
