---
title: "Examiner and site effects: between-group variation and what you may do about it"
description: "How to measure OSCE examiner and site effects with between-group variation, the thresholds to read it against, and which adjustments are defensible."
canonical: https://studydrome.com/docs/osce/quality/examiner-and-site-effects/
updated: 2026-10-10
---

# Examiner and site effects: between-group variation and what you may do about it

How to measure OSCE examiner and site effects with between-group variation, the thresholds to read it against, and which adjustments are defensible.

Candidates at the same station do not all meet the same examiner, circuit, or site. Between-group variation measures how much of the score spread is attributable to that rather than to the candidates. AMEE Guide 49 puts it ideally under 30 percent, with anything over 40 percent a cause for concern. This page covers how to read the number and which corrections a board may apply afterward.

## What the number actually measures

Group your candidates at one station by whatever they did not choose: the examiner who marked them, the circuit they rotated through, the morning or afternoon session, the site. Then ask how much of the total variation in that station's scores lies between those groups rather than within them.

Suppose candidates were allocated at random and every group marked alike. The between-group share should then be small. A large share says one of two things. Either the groups were not equivalent, which is an allocation problem. Or they were not marked alike, which is an examiner problem.

The calculation is an analysis of variance at station level. It belongs in the [quality report](/docs/osce/quality/the-post-osce-quality-report/), next to the other station metrics. Report it per station, not for the exam as a whole. An exam-level average hides the one station where a single examiner marked a whole circuit.

## Reading the thresholds

| Between-group share | Reading | Action |
| --- | --- | --- |
| Under 30 percent | Within the range the guide treats as acceptable | Record it |
| 30 to 40 percent | Watch | Note the station and compare with last sitting |
| Over 40 percent | Cause for concern | Investigate before the board signs |
| Over 60 percent | The groups, not the candidates, are driving the score | Expect a finding and a named action |

Very high values are not rare in small cohorts. A published review of one 24-candidate sitting reported between-group variation from 67.3 to 95.9 percent across its stations. The authors read that as inconsistent marking rather than as a property of the students.

Two cautions come with the number. It rises when groups are small, because a single unusual examiner is a larger share of a smaller pool. And it says nothing about direction. One examiner may have been harsh. Another may have been generous. The two circuits may have seen different candidates. The number does not say which.

## Why groups differ in the first place

The size of examiner effects is well established. The [examiner variance page](/docs/osce/people/examiner-variance/) covers it, with the people-side response: training, calibration and how examiners are spread across circuits. Two findings are worth carrying here, because they frame what the analysis can and cannot fix.

Examiner stringency is a large and stable effect, not noise. It persists across sittings, which is why it shows up as a group effect rather than averaging out. And a candidate of a given ability can expect a materially different score depending on which examiner cohort they met.

That has one consequence for this page. A high between-group value is usually a real signal, not a statistical artifact. Treating it as noise because the cohort was small is the common mistake.

## Comparing sites when nobody shares candidates

Multi-site exams have a hard problem. Each site sees its own candidates. A difference in mean score could be due to the site or the students. The data cannot separate them.

The method that does separate them is to give examiners at different sites the same performances to mark. Yeates and colleagues did this by video. Examiners score a shared set of recorded encounters alongside their live candidates. That links the sites onto a common scale. It is work, and it is the only way to make a between-site claim that survives challenge.

Short of that, two weaker checks help. Compare each site against itself over time, which removes the cohort question. And compare the spread of global grades rather than the checklist means. Grades are less sensitive to a station's own scale.

## May you adjust the scores?

Sometimes, and the choice is a board decision recorded in the minutes, not an analyst's call. AMEE Guide 49 describes three post-hoc remedies with very different standing.

**Adjustment to a common mean across sites** is the routine one. Where a linking design shows a stable offset, scores are brought onto a common scale.

**Station-level adjustment** is rare. It corrects one station where the evidence is strong and specific. It needs a recorded justification.

**Removing a station** is rare and, done without evidence, indefensible. It is for a station that has been shown to have failed. It applies to the whole cohort. The [ratification page](/docs/osce/scoring/results-ratification-and-release/) covers that decision and who takes it.

Three rules keep any of them defensible. Decide the principle before you see who is affected. Apply it to everyone at that station. Write the reason and the evidence in the minute. An adjustment chosen after the pass list is visible is not a correction. It is a change of standard.

## What to do before the results go out

1. Produce the per-station between-group figures with the rest of the report, before ratification.
2. Investigate anything over 40 percent: how many examiners staffed the station, how many candidates each saw, and whether one group sat a disrupted session.
3. Check the [incident log](/docs/osce/delivery/incidents-and-contingencies/) before blaming an examiner. A replacement examiner, a paused station, or an equipment failure explains many of the group differences.
4. Take a recommendation to the board: accept, adjust by the recorded rule, or withdraw.
5. Feed the finding into examiner training and into next year's allocation, which is where it actually gets fixed.

## Limits

Between-group variation is a screening statistic. It tells you where to look and never who was wrong. Naming an examiner on one station's number is unfair and usually unsupportable. The comparison is unadjusted for the candidates each examiner happened to see.

The published effect sizes also come from large licensing exams with stable examiner pools. A school with twelve examiners and one circuit cannot reproduce those designs. Read your own numbers as a trend across sittings, not as a measurement of a person.

## Frequently asked questions

### What is between-group variation in an OSCE?

The share of a station's score variation that comes from the group a candidate was in, such as the examiner, circuit, session, or site, rather than from the candidates themselves. AMEE Guide 49 treats under 30 percent as acceptable and over 40 percent as a cause for concern.

### What should I do if a station shows high between-group variation?

Investigate before results are ratified. Check how many examiners staffed the station and how many candidates each saw. Read the incident log for a disrupted session. Then take a recommendation to the board: accept it, adjust it according to a rule set in advance, or withdraw the station for everyone.

### Can OSCE scores be adjusted for a harsh examiner?

Adjustment to a common mean across sites is routine where a linking design supports it. Station-level adjustment is rare and needs a recorded reason. Decide any adjustment by principle before anyone sees who is affected. Apply it to everyone at that station and minute the evidence.

### How do you fairly compare two OSCE sites?

Give examiners at both sites the same performances to mark, usually recorded encounters scored alongside their live candidates. That links the sites onto a common scale. Without it, a difference in mean score cannot be separated from a difference in the candidates.

### Does high between-group variation mean an examiner is biased?

No. It is a screening signal that points at a station and a grouping, not a verdict on a person. The comparison is unadjusted for the candidates each examiner happened to see. Use it to trigger review, training, and better allocation, never to label an examiner.

## Sources

**Sources**

1. [Pell G, Fuller R, Homer M, Roberts T. How to measure the quality of the OSCE: a review of metrics. AMEE Guide No. 49](https://doi.org/10.3109/0142159X.2010.507716) (Medical Teacher, 2010)
2. [Homer M. Pass/fail decisions and standards: the impact of differential examiner stringency on OSCE outcomes](https://doi.org/10.1007/s10459-022-10096-9) (Advances in Health Sciences Education, 2022)
3. [Yeates P, Moult A, Cope N, et al. Measuring the effect of examiner variability in a multiple-circuit objective structured clinical examination (OSCE)](https://doi.org/10.1097/ACM.0000000000004028) (Academic Medicine, 2021)
4. [Cade AE, Meuller N. Measuring the quality of the OSCE in a chiropractic programme: a review of metrics and recommendations](https://doi.org/10.7899/JCE-22-29) (Journal of Chiropractic Education, 2024)

> [!TIP]
> **In practice**
> Comparing examiners at a station needs two or more of them on it, and the comparison is unadjusted whichever tool computes it. To see how that screening signal is presented, read /docs/osce/quality/in-studydrome-osce-analytics/.
