Skip to content

Designing the OSCE marking scheme: items, anchors, weighting and global ratings

Last updated · 10 min read

An OSCE marking scheme has two parts. The first is a checklist of observable items for the station, scored as binary or on a short, anchored scale, and grouped under section headers. The second is an independent global rating on a 4- or 5-point scale. The checklist gives the station score and the feedback. The global rating feeds standard setting and quality checks. The two are never added together.

The debate over checklists versus global ratings has its own page: should OSCEs use checklists or global rating scales? This page assumes the answer most schools reach, which is both, and covers the design decisions that follow.

Why the scheme is a design document

The scheme decides what the examiner looks for. An item on the sheet gets observed; a behavior not on the sheet gets forgotten. The AMEE guide on OSCE organization identifies four factors that reduce OSCE reliability: poorly standardized patients, untrained examiners, poor questions, and an unsuitable scoring instrument. A weak scheme is not a scoring problem found after the exam. It is a design fault built in before it.

The scheme also decides what the results can show. Item-level scores make item analysis possible. A global rating enables borderline regression. A scheme without either leaves the assessment committee with a total and nothing to explain it.

What are the design decisions?

Decision

Options

What the evidence says

Recommendation

Item size

Fine-grained actions vs. broader tasks

Very fine items reward completeness over competence; broad items lose feedback value

Items an examiner can see in one glance, each one an action

Item scale

Binary (done / not done) vs. 3- to 7-point scaled items

The AMEE guide describes both binary checklists and 5- to 7-point rating-scale checklists; scaled items need anchors to stay consistent

Binary for procedural steps; 3-point (not done / partly / done) where quality matters

Anchor wording

Numbers only vs. labeled levels vs. labeled and described levels

Unlabeled levels drift between examiners

Every level labeled; describe the middle level

Section headers

None vs. sections by phase or domain

Sections give the examiner a map and give the report a domain score

Sections for any station above eight items

Weighting

Equal items vs. weighted sections vs. weighted items

Weights should follow the blueprint cell, not the writer's opinion

Weight sections, not single items; record the reason

Checklist length

10 vs. 20 vs. 40 items

Adding 20 items to a 20-item list did not change observer accuracy or consistency

As many items as the task has observable steps, and no more

Global rating

None vs. added to the score vs. independent

Global ratings show higher inter-station reliability than checklists and separate levels of training

Independent, never summed with the checklist

Binary or scaled items?

A binary item asks one question: did it happen? It is fast to score, easy to train, and simple to analyze. It suits procedural steps and safety actions, where the step is either done or not.

A scaled item asks how well it happened. It suits communication, explanation, and examination technique, where a step can be done badly. The cost of consistency is that two examiners must agree on what "partly done" means. That agreement comes from the anchor wording and from examiner training, not from the item itself.

A 2015 systematic review pooled 45 studies that used both instruments. Inter-rater reliability was much the same for checklists and global ratings. That finding applies to well-anchored items. It does not rescue a scale with numbers and no words. Keep scales short. A 3-point item is easier to anchor than a 7-point one, and the extra levels rarely earn their place.

How should anchors be written?

An anchor is the sentence under each level that tells the examiner what that level looks like. Three rules:

  1. Describe behavior, not judgment. "Asked about allergies before prescribing" is an anchor. "Safe prescribing" is a verdict.
  2. Describe the middle. Examiners agree on the ends and disagree in the middle. "Asked about allergies after writing the prescription" defines a middle level.
  3. Use the same words as the candidate instructions. If the task said "explain", the anchor scores explaining.

The scoring scale itself is a reusable object. Define a scale once, such as Not done / Partially done / Done, with a value for each level. Then attach it to every station in the bank. Per-item scales exist for stations where one item requires a different scale from the rest. Keep them rare, because every extra scale is another thing to train.

How long should the checklist be?

Shorter than most first drafts. A study at Dalhousie University provided 57 observers with scripted station videos using either 20- or 40-item checklists. Adding the second 20 items changed neither accuracy nor consistency. Long lists do not buy agreement. They buy examiners who stop watching the candidate to find the next box.

The better test is observability within the time. If the station lasts 8 minutes and the list has 30 items, the examiner is scoring 4 items per minute. Cut until each item is something the examiner will see, in order, with time to record it.

Where does the global rating sit?

The global rating is the examiner's overall judgment of the performance. It uses a short scale such as fail / borderline / pass / good / excellent. It is scored after the checklist and recorded separately. It is never added to the checklist score.

Two reasons. First, it carries information the checklist cannot. A 1998 study of surgical residents compared the two. Global ratings by expert examiners showed higher inter-station reliability and better construct validity than checklists. Adding the checklist did not improve the global rating. A 2003 study found that analytic global ratings separated third-year from fourth-year students. Their inter-station reliability was 0.70, against 0.54 for the checklist.

Second, borderline regression needs its independence. The method regresses checklist scores on global grades, station by station, to find the score that a borderline performance predicts. If the global grade were derived from the checklist, the regression would be circular. The scoring section covers the method. The disagreement between the two instruments, read as R squared, is the audit the quality section uses.

The borderline grade needs care. Examiners hesitate at the cut, and a "borderline pass" grade invites a coin toss. One medical school replaced it with a plain "borderline" grade. A formula was used to resolve each one after the exam, based on the station's difficulty and the candidate's performance in the same domain elsewhere. Of 14,634 observations, 4.69 percent were borderline. The exam failure rate rose from 0.7 to 4.1 percent, and examiners reported that the grade matched what they meant.

Critical items and must-pass stations

Some items carry more than their mark. An action that would harm the patient can be written to fail the station regardless of total. So can a failure of professional conduct. That is a policy decision with appeal consequences, so it has its own page: critical errors and must-pass stations. The marking scheme records which items are critical and at what level they fail. The policy says why.

What does the sheet look like?

  • Station title, version and date at the top.
  • Candidate identifier field, not a name.
  • Sections in the order the task unfolds.
  • Items numbered, one line each, with the scale levels as columns.
  • Critical items marked in a way the examiner cannot miss.
  • The global rating last, on its own, with its anchors printed.
  • A comments box, because written feedback is one of the examiner-training outcomes the AMEE guide lists.

On paper, the sheet is single-sided. On a screen, the same structure becomes a form. A 2022 study of a national licensing exam found 14.2 percent of paper checklists had at least one missing rating. Electronic checklists had none. The electronic form has both a fairness argument and a convenience argument. The delivery section covers the choice.

Limits

The evidence base skews toward medicine and toward procedural skills. The 2015 review's 45 studies were mostly surgery, endoscopy, resuscitation and anesthesia, and all involved physicians or physicians in training. Findings about global ratings assume expert examiners. Where non-clinician examiners score, checklists carry more weight, and the global rating needs more training.

Frequently asked questions

How many items should an OSCE checklist have?

As many as the task has observable steps within the time, and usually fewer than a first draft. A study that compared 20-item and 40-item checklists found no difference in observer accuracy or consistency. Ten to twenty well-chosen items are typical for an eight- to ten-minute station.

Should the global rating count toward the station score?

No. Record it separately and never add it to the checklist total. The global rating carries the examiner's overall judgment and feeds borderline regression standard setting, which needs it to be independent of the checklist. Adding the two together would make the pass mark circular.

What is the difference between a checklist and a rating scale in an OSCE?

A checklist scores each action as done or not done, or on a short scale. A global rating scale rates overall performance across a few anchored levels. Most schools use a checklist for the station score and feedback, and a global rating for standard setting and quality review.

How do you weigh items in an OSCE marking scheme?

Weight sections rather than single items, and take the weights from the blueprint cell the station fills. A history station might weigh the history section more than the closing summary. Record the reason for each weight, because an appeal will ask.

What is a borderline grade in an OSCE?

A global rating level that says the performance was neither a clear pass nor a clear fail. In borderline regression, it anchors the pass mark for the station. Some schools resolve borderline grades into pass or fail after the exam using a formula, rather than letting the examiner decide at the desk.

Sources

Was this page helpful?