Skip to content

Hawks and doves: how much examiners move OSCE results, and what an organizer can do about it

Last updated · 7 min read

Examiners move OSCE results by a measurable amount. In the largest study to date, examiner stringency explained about 16 percent of the variation in station scores. It explained 8 percent of the variation in global grades. Examiners fell into two clusters, hawks and doves. An organizer cannot remove that. An organizer can sample enough stations and decide at exam level. An organizer can monitor examiners after the day, rotate them across circuits, and adjust where the design allows.

The kinds of examiner effect are described in how to train OSCE examiners. This page carries the numbers and the controls.

Why the numbers matter to an organizer

Every OSCE lead knows some examiners are harder than others. The board asks how much it matters. The answer decides the design. If examiner effects are small, a 10-station exam with one examiner per station is fine. If they are large, the exam needs more stations. It needs a rotation that spreads each examiner's effect across candidates. It needs a way to find the outliers afterward. The research says they are large.

How much do examiners move results?

Study

Data

What examiners accounted for

Homer 2022

442 administrations of an 18-station OSCE for international medical graduates seeking UK registration

About 16 percent of the variance in domain scores and 8 percent in global grades; examiners consistent in their stringency across both (r = 0.76); two clusters, hawks and doves

Yeates et al. 2021

A multiple-circuit undergraduate OSCE with video-linked examiner comparison

The same candidate's expected score ranged from 68.8 percent under the most hawkish examiner cohort to 75.9 percent under the most doveish

Yeates et al. 2023

The same method applied across sites in a distributed OSCE

Examiner cohorts at different sites could be compared and their influence estimated

Two details from the 2022 study bear on design. Exam-level reliability was 0.75 for domain scores and 0.69 for grades. Examiner variance was roughly twice as large for scores as for grades. And correcting for examiner stringency produced lower cut scores under borderline regression than the raw marks did. That would have raised pass rates. The author's conclusion: performance exams need sufficient domain sampling to make sound pass-and-fail decisions.

The 2021 study's figure is the one to show a board. A seven-point swing on the same performance separates a comfortable pass from a borderline one. It comes from which circuit a candidate happened to sit in.

What does an organizer do about it?

Five controls, in the usual order.

  1. Sample more stations. The defense on which the OSCE was built. Many short stations, each with a different examiner, so that no single examiner dominates a candidate's result. The reliability page gives the station counts the evidence supports.
  2. Decide at exam level. A hawk on one station fails a candidate at that station. The pass decision is based on the total across stations, where hawks and doves partly cancel out. Station-level pass rules, such as must-pass stations, bring single-examiner risk back. They need the justification the critical-error policy asks for.
  3. Rotate examiners across circuits and sessions. Suppose the same examiner marks the same station in the morning and the afternoon. Their stringency then falls equally on every candidate. When a hawk marks one circuit and a dove the other, it falls on half the cohort. Spread the effect rather than concentrate it.
  4. Monitor after the day. The AMEE guide on OSCE quality metrics excludes examiner effects from the data. It uses between-group variation across parallel circuits, and station-level comparison of examiners marking the same station. Outliers get feedback and refresher training. The quality section sets out the metrics.
  5. Adjust where the design allows. The video method records a sample of performances. Examiners from several circuits score the same ones. That links the circuits, so the effect of each examiner cohort can be estimated. A 2024 simulation study tested the accuracy of the adjustment. The method is a research tool today. The point for an organizer is that examiner effects can be measured, not only suspected.

What does training contribute?

Less than most programs hope. The training page carries the evidence. A 2026 scoping review found the effects of training and calibration on scoring inconsistent. Benefits were most evident for borderline performances. Training is still required. It teaches the instruments and makes the standard explicit. It lets the school say every examiner was prepared. It does not make a hawk into a dove. Design does the rest.

What the monitoring report shows

A one-page report after each sitting, per station and per examiner:

  • Each examiner's mean checklist score and mean global grade, against the station mean across circuits.
  • The gap between circuits for the same station. That is the between-group variation the AMEE metrics guide uses.
  • The examiners far from their station's mean, with the number of candidates they scored.
  • What happened next: feedback was given, a refresher was scheduled, and the station was reviewed.

The report names people, so the consent and data page applies. Examiners are told the comparison exists and who sees it.

Limits

The two large studies are British. One is a licensing exam for international graduates, with a stable pool of examiners. Effects in a small school with a new cohort of examiners may be larger. The percentages come from statistical models and depend on the model. The direction and the rough size are the robust findings. Adjustment methods are still being tested. They are not a substitute for design.

Frequently asked questions

How much do examiners affect OSCE scores?

A lot. One study covered 442 sittings of an 18-station licensing OSCE. Examiner stringency explained about 16 percent of the variance in station scores, and 8 percent in global grades. In a multi-circuit student OSCE, the same candidate's expected score ranged from 68.8 to 75.9 percent by examiner cohort.

What are hawks and doves in OSCE marking?

Hawks mark harder than their peers. Doves mark softer. A 2022 study found examiners fall into the two groups. They stay consistent across checklist scores and global grades. The effect is stable, so design rather than training has to absorb it.

Can examiner stringency be corrected statistically?

In research, yes. Examiners from different circuits score the same recorded performances. That links the circuits, and each cohort's effect can be estimated and adjusted. A 2024 simulation study tested the accuracy of the adjustment. Most schools use the monitoring, not the adjustment.

Does examiner training fix hawk-and-dove effects?

No. Training reduces variation and makes the standard explicit. A 2026 scoping review found its effects on scoring mixed. The controls that work are design controls: enough stations, exam-level decisions, examiner rotation across circuits, and post-exam monitoring.

Should a hawkish examiner be removed from the panel?

Not for being a hawk. A consistent hawk spread across circuits harms nobody in particular. Remove or retrain an examiner whose marks are erratic or who ignores the instrument. Do the same for an outlier on one station whose marks the review cannot explain.

Sources

Was this page helpful?