---
title: "Hawks and doves: how much examiners move OSCE results, and what an organizer can do about it"
description: "How much examiners move OSCE results, from large studies of examiner stringency, and the five controls an organizer can apply."
canonical: https://studydrome.com/docs/osce/people/examiner-variance/
updated: 2026-10-09
---

# Hawks and doves: how much examiners move OSCE results, and what an organizer can do about it

How much examiners move OSCE results, from large studies of examiner stringency, and the five controls an organizer can apply.

Examiners move OSCE results by a measurable amount. In the largest study to date, examiner stringency explained about 16 percent of the variation in station scores. It explained 8 percent of the variation in global grades. Examiners fell into two clusters, hawks and doves. An organizer cannot remove that. An organizer can sample enough stations and decide at exam level. An organizer can monitor examiners after the day, rotate them across circuits, and adjust where the design allows.

The kinds of examiner effect are described in [how to train OSCE examiners](/answers/osce-assessment/how-to-train-osce-examiners/). This page carries the numbers and the controls.

## Why the numbers matter to an organizer

Every OSCE lead knows some examiners are harder than others. The board asks how much it matters. The answer decides the design. If examiner effects are small, a 10-station exam with one examiner per station is fine. If they are large, the exam needs more stations. It needs a rotation that spreads each examiner's effect across candidates. It needs a way to find the outliers afterward. The research says they are large.

## How much do examiners move results?

| Study | Data | What examiners accounted for |
| --- | --- | --- |
| Homer 2022 | 442 administrations of an 18-station OSCE for international medical graduates seeking UK registration | About 16 percent of the variance in domain scores and 8 percent in global grades; examiners consistent in their stringency across both (r = 0.76); two clusters, hawks and doves |
| Yeates et al. 2021 | A multiple-circuit undergraduate OSCE with video-linked examiner comparison | The same candidate's expected score ranged from 68.8 percent under the most hawkish examiner cohort to 75.9 percent under the most doveish |
| Yeates et al. 2023 | The same method applied across sites in a distributed OSCE | Examiner cohorts at different sites could be compared and their influence estimated |

Two details from the 2022 study bear on design. Exam-level reliability was 0.75 for domain scores and 0.69 for grades. Examiner variance was roughly twice as large for scores as for grades. And correcting for examiner stringency produced lower cut scores under borderline regression than the raw marks did. That would have raised pass rates. The author's conclusion: performance exams need sufficient domain sampling to make sound pass-and-fail decisions.

The 2021 study's figure is the one to show a board. A seven-point swing on the same performance separates a comfortable pass from a borderline one. It comes from which circuit a candidate happened to sit in.

## What does an organizer do about it?

Five controls, in the usual order.

1. **Sample more stations.** The defense on which the OSCE was built. Many short stations, each with a different examiner, so that no single examiner dominates a candidate's result. The [reliability page](/docs/osce/foundations/validity-and-reliability/) gives the station counts the evidence supports.
2. **Decide at exam level.** A hawk on one station fails a candidate at that station. The pass decision is based on the total across stations, where hawks and doves partly cancel out. Station-level pass rules, such as must-pass stations, bring single-examiner risk back. They need the justification the [critical-error policy](/docs/osce/design/critical-errors-and-must-pass/) asks for.
3. **Rotate examiners across circuits and sessions.** Suppose the same examiner marks the same station in the morning and the afternoon. Their stringency then falls equally on every candidate. When a hawk marks one circuit and a dove the other, it falls on half the cohort. Spread the effect rather than concentrate it.
4. **Monitor after the day.** The AMEE guide on OSCE quality metrics excludes examiner effects from the data. It uses between-group variation across parallel circuits, and station-level comparison of examiners marking the same station. Outliers get feedback and refresher training. The [quality section](/docs/osce/quality/) sets out the metrics.
5. **Adjust where the design allows.** The video method records a sample of performances. Examiners from several circuits score the same ones. That links the circuits, so the effect of each examiner cohort can be estimated. A 2024 simulation study tested the accuracy of the adjustment. The method is a research tool today. The point for an organizer is that examiner effects can be measured, not only suspected.

## What does training contribute?

Less than most programs hope. The [training page](/docs/osce/people/examiner-training-program/) carries the evidence. A 2026 scoping review found the effects of training and calibration on scoring inconsistent. Benefits were most evident for borderline performances. Training is still required. It teaches the instruments and makes the standard explicit. It lets the school say every examiner was prepared. It does not make a hawk into a dove. Design does the rest.

## What the monitoring report shows

A one-page report after each sitting, per station and per examiner:

- Each examiner's mean checklist score and mean global grade, against the station mean across circuits.
- The gap between circuits for the same station. That is the between-group variation the AMEE metrics guide uses.
- The examiners far from their station's mean, with the number of candidates they scored.
- What happened next: feedback was given, a refresher was scheduled, and the station was reviewed.

The report names people, so the [consent and data page](/docs/osce/planning/consent-recording-and-data-protection/) applies. Examiners are told the comparison exists and who sees it.

## Limits

The two large studies are British. One is a licensing exam for international graduates, with a stable pool of examiners. Effects in a small school with a new cohort of examiners may be larger. The percentages come from statistical models and depend on the model. The direction and the rough size are the robust findings. Adjustment methods are still being tested. They are not a substitute for design.

## Frequently asked questions

### How much do examiners affect OSCE scores?

A lot. One study covered 442 sittings of an 18-station licensing OSCE. Examiner stringency explained about 16 percent of the variance in station scores, and 8 percent in global grades. In a multi-circuit student OSCE, the same candidate's expected score ranged from 68.8 to 75.9 percent by examiner cohort.

### What are hawks and doves in OSCE marking?

Hawks mark harder than their peers. Doves mark softer. A 2022 study found examiners fall into the two groups. They stay consistent across checklist scores and global grades. The effect is stable, so design rather than training has to absorb it.

### Can examiner stringency be corrected statistically?

In research, yes. Examiners from different circuits score the same recorded performances. That links the circuits, and each cohort's effect can be estimated and adjusted. A 2024 simulation study tested the accuracy of the adjustment. Most schools use the monitoring, not the adjustment.

### Does examiner training fix hawk-and-dove effects?

No. Training reduces variation and makes the standard explicit. A 2026 scoping review found its effects on scoring mixed. The controls that work are design controls: enough stations, exam-level decisions, examiner rotation across circuits, and post-exam monitoring.

### Should a hawkish examiner be removed from the panel?

Not for being a hawk. A consistent hawk spread across circuits harms nobody in particular. Remove or retrain an examiner whose marks are erratic or who ignores the instrument. Do the same for an outlier on one station whose marks the review cannot explain.

## Sources

**Sources**

1. [Homer M. Pass/fail decisions and standards: the impact of differential examiner stringency on OSCE outcomes](https://doi.org/10.1007/s10459-022-10096-9) (Advances in Health Sciences Education, 2022)
2. [Yeates P, Moult A, Cope N, et al. Measuring the effect of examiner variability in a multiple-circuit objective structured clinical examination (OSCE)](https://doi.org/10.1097/ACM.0000000000004028) (Academic Medicine, 2021)
3. [Yeates P, et al. Using video-based examiner score comparison and adjustment (VESCA) to compare the influence of examiners at different sites in a distributed objective structured clinical exam (OSCE)](https://doi.org/10.1186/s12909-023-04774-4) (BMC Medical Education, 2023)
4. [Yeates P, et al. Investigating the accuracy of adjusting for examiner differences in multi-centre Objective Structured Clinical Exams (OSCEs): a simulation study of video-based Examiner Score Comparison and Adjustment (VESCA)](https://doi.org/10.1186/s12909-024-06462-3) (BMC Medical Education, 2024)
5. [Pell G, Fuller R, Homer M, Roberts TE. How to measure the quality of the OSCE: a review of metrics. AMEE Guide No. 49](https://doi.org/10.3109/0142159X.2010.507716) (Medical Teacher, 2010)
6. [Thampy H, Callanan N, Ahmed A, Taromsari S. Examiner training and calibration for simulated clinical examinations: a scoping review](https://doi.org/10.1111/medu.70253) (Medical Education, 2026)

> [!TIP]
> **In practice**
> Monitoring needs every score captured, with the examiner and the station attached to each one. An electronic scoring screen provides that and paper does not. To see how StudyDrome's examiner app records scores per examiner and station, read /docs/osce/people/in-studydrome-examiner-app/.
