An OSCE marking scheme has two parts. The first is a checklist of observable items for the station, scored as binary or on a short, anchored scale, and grouped under section headers. The second is an independent global rating on a 4- or 5-point scale. The checklist gives the station score and the feedback. The global rating feeds standard setting and quality checks. The two are never added together.
The debate over checklists versus global ratings has its own page: should OSCEs use checklists or global rating scales? This page assumes the answer most schools reach, which is both, and covers the design decisions that follow.
Why the scheme is a design document
The scheme decides what the examiner looks for. An item on the sheet gets observed; a behavior not on the sheet gets forgotten. The AMEE guide on OSCE organization identifies four factors that reduce OSCE reliability: poorly standardized patients, untrained examiners, poor questions, and an unsuitable scoring instrument. A weak scheme is not a scoring problem found after the exam. It is a design fault built in before it.
The scheme also decides what the results can show. Item-level scores make item analysis possible. A global rating enables borderline regression. A scheme without either leaves the assessment committee with a total and nothing to explain it.
What are the design decisions?
Decision | Options | What the evidence says | Recommendation |
|---|---|---|---|
Item size | Fine-grained actions vs. broader tasks | Very fine items reward completeness over competence; broad items lose feedback value | Items an examiner can see in one glance, each one an action |
Item scale | Binary (done / not done) vs. 3- to 7-point scaled items | The AMEE guide describes both binary checklists and 5- to 7-point rating-scale checklists; scaled items need anchors to stay consistent | Binary for procedural steps; 3-point (not done / partly / done) where quality matters |
Anchor wording | Numbers only vs. labeled levels vs. labeled and described levels | Unlabeled levels drift between examiners | Every level labeled; describe the middle level |
Section headers | None vs. sections by phase or domain | Sections give the examiner a map and give the report a domain score | Sections for any station above eight items |
Weighting | Equal items vs. weighted sections vs. weighted items | Weights should follow the blueprint cell, not the writer's opinion | Weight sections, not single items; record the reason |
Checklist length | 10 vs. 20 vs. 40 items | Adding 20 items to a 20-item list did not change observer accuracy or consistency | As many items as the task has observable steps, and no more |
Global rating | None vs. added to the score vs. independent | Global ratings show higher inter-station reliability than checklists and separate levels of training | Independent, never summed with the checklist |
Binary or scaled items?
A binary item asks one question: did it happen? It is fast to score, easy to train, and simple to analyze. It suits procedural steps and safety actions, where the step is either done or not.
A scaled item asks how well it happened. It suits communication, explanation, and examination technique, where a step can be done badly. The cost of consistency is that two examiners must agree on what "partly done" means. That agreement comes from the anchor wording and from examiner training, not from the item itself.
A 2015 systematic review pooled 45 studies that used both instruments. Inter-rater reliability was much the same for checklists and global ratings. That finding applies to well-anchored items. It does not rescue a scale with numbers and no words. Keep scales short. A 3-point item is easier to anchor than a 7-point one, and the extra levels rarely earn their place.
How should anchors be written?
An anchor is the sentence under each level that tells the examiner what that level looks like. Three rules:
- Describe behavior, not judgment. "Asked about allergies before prescribing" is an anchor. "Safe prescribing" is a verdict.
- Describe the middle. Examiners agree on the ends and disagree in the middle. "Asked about allergies after writing the prescription" defines a middle level.
- Use the same words as the candidate instructions. If the task said "explain", the anchor scores explaining.
The scoring scale itself is a reusable object. Define a scale once, such as Not done / Partially done / Done, with a value for each level. Then attach it to every station in the bank. Per-item scales exist for stations where one item requires a different scale from the rest. Keep them rare, because every extra scale is another thing to train.
How long should the checklist be?
Shorter than most first drafts. A study at Dalhousie University provided 57 observers with scripted station videos using either 20- or 40-item checklists. Adding the second 20 items changed neither accuracy nor consistency. Long lists do not buy agreement. They buy examiners who stop watching the candidate to find the next box.
The better test is observability within the time. If the station lasts 8 minutes and the list has 30 items, the examiner is scoring 4 items per minute. Cut until each item is something the examiner will see, in order, with time to record it.
Where does the global rating sit?
The global rating is the examiner's overall judgment of the performance. It uses a short scale such as fail / borderline / pass / good / excellent. It is scored after the checklist and recorded separately. It is never added to the checklist score.
Two reasons. First, it carries information the checklist cannot. A 1998 study of surgical residents compared the two. Global ratings by expert examiners showed higher inter-station reliability and better construct validity than checklists. Adding the checklist did not improve the global rating. A 2003 study found that analytic global ratings separated third-year from fourth-year students. Their inter-station reliability was 0.70, against 0.54 for the checklist.
Second, borderline regression needs its independence. The method regresses checklist scores on global grades, station by station, to find the score that a borderline performance predicts. If the global grade were derived from the checklist, the regression would be circular. The scoring section covers the method. The disagreement between the two instruments, read as R squared, is the audit the quality section uses.
The borderline grade needs care. Examiners hesitate at the cut, and a "borderline pass" grade invites a coin toss. One medical school replaced it with a plain "borderline" grade. A formula was used to resolve each one after the exam, based on the station's difficulty and the candidate's performance in the same domain elsewhere. Of 14,634 observations, 4.69 percent were borderline. The exam failure rate rose from 0.7 to 4.1 percent, and examiners reported that the grade matched what they meant.
Critical items and must-pass stations
Some items carry more than their mark. An action that would harm the patient can be written to fail the station regardless of total. So can a failure of professional conduct. That is a policy decision with appeal consequences, so it has its own page: critical errors and must-pass stations. The marking scheme records which items are critical and at what level they fail. The policy says why.
What does the sheet look like?
- Station title, version and date at the top.
- Candidate identifier field, not a name.
- Sections in the order the task unfolds.
- Items numbered, one line each, with the scale levels as columns.
- Critical items marked in a way the examiner cannot miss.
- The global rating last, on its own, with its anchors printed.
- A comments box, because written feedback is one of the examiner-training outcomes the AMEE guide lists.
On paper, the sheet is single-sided. On a screen, the same structure becomes a form. A 2022 study of a national licensing exam found 14.2 percent of paper checklists had at least one missing rating. Electronic checklists had none. The electronic form has both a fairness argument and a convenience argument. The delivery section covers the choice.
Limits
The evidence base skews toward medicine and toward procedural skills. The 2015 review's 45 studies were mostly surgery, endoscopy, resuscitation and anesthesia, and all involved physicians or physicians in training. Findings about global ratings assume expert examiners. Where non-clinician examiners score, checklists carry more weight, and the global rating needs more training.
Frequently asked questions
How many items should an OSCE checklist have?
As many as the task has observable steps within the time, and usually fewer than a first draft. A study that compared 20-item and 40-item checklists found no difference in observer accuracy or consistency. Ten to twenty well-chosen items are typical for an eight- to ten-minute station.
Should the global rating count toward the station score?
No. Record it separately and never add it to the checklist total. The global rating carries the examiner's overall judgment and feeds borderline regression standard setting, which needs it to be independent of the checklist. Adding the two together would make the pass mark circular.
What is the difference between a checklist and a rating scale in an OSCE?
A checklist scores each action as done or not done, or on a short scale. A global rating scale rates overall performance across a few anchored levels. Most schools use a checklist for the station score and feedback, and a global rating for standard setting and quality review.
How do you weigh items in an OSCE marking scheme?
Weight sections rather than single items, and take the weights from the blueprint cell the station fills. A history station might weigh the history section more than the closing summary. Record the reason for each weight, because an appeal will ask.
What is a borderline grade in an OSCE?
A global rating level that says the performance was neither a clear pass nor a clear fail. In borderline regression, it anchors the pass mark for the station. Some schools resolve borderline grades into pass or fail after the exam using a formula, rather than letting the examiner decide at the desk.