Skip to content

What Reports Does a Board of Examiners Need?

The evidence a board needs before an exam runs, after it runs, and across the year — and the line where reporting stops and judgement starts.

board of examiners

A board of examiners needs three kinds of evidence, not one. Before the exam: proof that the paper covers the syllabus, and a pass mark rule agreed in advance. After it: reliability, item flags, and how each attempt ended. Across the year: whether the question bank and the cohort are still healthy.

Most boards see the first kind of evidence and almost none of the second. The results spreadsheet arrives on time. The reliability figure does not arrive at all. That gap is where appeals begin. It is also where an external examiner asks the question nobody prepared for.

What is a board of examiners deciding?

A board signs off results. That is the visible job. Two more decisions sit underneath it.

The first decision is about each candidate. Did this person pass, fail, or need a referral?

The second is about the paper. Did the exam measure what it claimed to measure? A board that ratifies marks from a broken paper has ratified nothing.

The third is about the program. Is assessment getting better or worse across the year?

Each decision requires different evidence, gathered at different times. Treating all three as one meeting is why boards run long and decide little.

What does the board need before the exam runs?

This is the leg most boards skip. Once results are in, it is too late to fix coverage and too late to argue about how the pass mark was set.

Three things should be signed off before a paper is delivered.

Coverage. A blueprint maps the paper against two axes, usually topic and competency. Each cell shows how many questions cover it, and how those questions have performed before. The empty cells are the point. They tell the board what the exam does not test.

Bank readiness. Coverage assumes you have questions worth using. A readiness check crosses the bank by difficulty band and by topic, then counts the good, the weak, and the never-used in each cell. The useful output is a plain warning: fewer than five good questions in the hard band, or a topic carrying only poor-quality items. A bank can look large and still be unable to support one balanced paper.

The pass mark rule. Agree the standard-setting method before the exam, not after. This is the single easiest thing to get wrong and the hardest to defend later. A method chosen once results are visible is not a standard. It is a negotiation.

None of this needs psychometric training to read. Coverage gaps, thin difficulty bands and an agreed method are three questions any chair can ask.

What does the board need after the exam has run?

Three numbers and one record. Keep this part short, because the detail belongs elsewhere.

Reliability describes how consistently the paper separated candidates. It is usually reported with a plain-English band attached. A coefficient above 0.9 is typically considered suitable for high-stakes decisions. Below 0.6, the usual wording is poor and introduces significant measurement error. A high-stakes result should not be ratified based on a figure the software itself flags as questionable.

The standard error of measurement matters more at the pass mark than anywhere else. It is the number that tells a board how much a mark of 49 and a mark of 51 really differ. Bring it to any discussion about a borderline candidate.

Item flags name the questions that misbehaved. The first to look at is a question that strong candidates got wrong more often than weak ones did. That pattern usually means the key is wrong or the stem is ambiguous.

The conduct record shows how each attempt ended: finished by the candidate, timed out, ended by an administrator, or terminated. Boards need this to separate a technical failure from misconduct.

For the method behind these figures, read our guide to reading an item analysis report. For pass marks in clinical exams, cohort size and cut score stability covers the question boards ask most.

What does the board need across the year?

Annual monitoring is the leg with the least written about it and the most value. Three views carry it.

Bank health. Track the share of administered questions rated good, the spread across difficulty bands, and how many items have never been used. A bank where a third of the questions have never been delivered is not a bank of that size. It is a smaller bank with a backlog.

Trends across semesters. Treat these with care. Most trend figures in assessment software compare two halves of a series: the first-half mean against the second-half mean. A gap beyond a fixed threshold receives a label such as improving or declining. That is a rule of thumb, not a significance test. Minute it as one, or the label becomes a finding it was never entitled to be.

Form equivalence. Where two versions of a paper ran, compare them. Candidate counts, means, pass rates and an effect size, with a p-value on the difference. This is an equivalence check on how two papers behaved. It is not equating. Nothing is placed on a common scale, and no conversion table is produced. Boards should use the precise term, because the imprecise one implies a guarantee nobody made.

One caution applies to every comparison here. Candidates are not randomly assigned to groups, cohorts or teachers. A gap between two groups is a question worth asking. It is not a finding.

What does a full governance cycle look like?

Stage

The question being answered

Evidence to table

Who prepares it

Before the exam

Does this paper cover what we teach?

Blueprint matrix, with coverage gaps named

Assessment lead

Before the exam

Can our bank support this paper?

Readiness warnings by difficulty band and topic

Assessment lead

Before the exam

How will the pass mark be set?

Standard-setting method, agreed and minuted

Board

After the exam

Did the paper measure consistently?

Reliability coefficient, its band, and the SEM

Exam office

After the exam

Which questions misbehaved?

Item flags, with the suppression threshold stated

Assessment lead

After the exam

Did anything go wrong on the day?

Finish categories and the conduct record

Exam office

Across the year

Is the bank still usable?

Quality share, difficulty mix, never-used count

Assessment lead

Across the year

Were parallel papers comparable?

Equivalence check with effect size and p-value

Assessment lead

Across the year

Is the program improving?

Semester comparison, labeled as a rule of thumb

Board

Then, in minute four, things: the method, the numbers, the exceptions, and the reason for each exception. A minute that records only outcomes is not a record. It is a receipt.

Where does the evidence stop and the judgment start?

No system convenes a board. That sounds obvious, and in practice it is the line that gets blurred.

Assessment software can compute a cut score. It cannot approve one. There is no gate that refuses to release results because reliability came in low. Item flags are advisory. Standard-setting warnings are advisory. A conduct report classifies how an attempt ended; it does not decide whether anyone cheated. Adjudication is human, and it should stay that way.

So judge a system by whether it makes the board's job possible, not by whether it does the board's job. Three properties matter more than any feature list.

It should freeze the paper. Publishing an exam should write an immutable snapshot, so the paper each candidate sat is kept exactly as they sat it. Without that, every later statistic describes something you can no longer produce.

It should keep a history. Every grade, re-grade, and un-grade should be written in a numbered row so nothing overwrites. An appeal a year later is won or lost on that trail.

It should refuse to report what it cannot support. StudyDrome Exam Manager stops rating a question below ten responses and discards its flags rather than showing noise. A report that stays silent is more useful to a board than a confident number built on four answers.

That last property is the one buyers rarely ask about, and boards always need. Governance is not the absence of judgment. It is a judgment made in the open, on evidence gathered before anyone knew the answer.

Frequently asked questions

What is the difference between a board of examiners and an exam board?

In the UK, "exam board" usually refers to an awarding organization, such as AQA or OCR. Those bodies set and mark school qualifications. A board of examiners is a university committee. It ratifies program results, agrees on classifications, and considers mitigating circumstances. The two are confused constantly. Ask which one someone means before you send them a report.

Who sits on a board of examiners?

Membership varies by institution, but the pattern is stable. A chair, usually a senior academic. Internal examiners who taught and marked the modules. At least one external examiner from another institution. A secretary to record decisions. An exam office representative who brings the data. Some boards also invite an assessment lead when standard setting is in question.

What does an external examiner need to see?

An external examiner checks comparability and process. They are not marking every script. They need the paper, the marking scheme, the mark distribution, and a sample of scripts across the range. They also need to know how the pass mark was set and why. Give them the reliability figure too. If it is weak, say so before they ask.

Can we generate detailed reports and analytics?

Most assessment systems produce three families of reports: results by candidate, item-level statistics, and cohort comparisons. What varies is which of them leave the system as a file and which stay on screen. Ask for that list report by report, rather than a yes. Ask which statistics are named, too. "Analytics" with no named coefficient tells a board nothing.

Who changed this question, and when?

This is the question an appeal turns on. A board needs to know that the paper it approved is the paper that ran. Two records answer it: a version history on the question itself, and an immutable snapshot taken when the exam is published. The second matters more, because it fixes the whole paper rather than one item.

What should a board of examiners minute?

Record the method, the numbers, the exceptions, and the reason for each exception. Name the standard-setting method and when it was agreed. Record the reliability figure and the pass mark. List every candidate decision that departed from the rule, with its justification. A minute that records only outcomes cannot be defended a year later.

Where to go next

For the psychometrics behind the post-exam leg, start with reading an item analysis report and how reliability is computed.

For coverage evidence before delivery, see how a blueprint matrix works.

For the whole evidence chain in one place, see what an assessment lead sees.

Written by Dimitri · Jul 1, 2026

Put this into practice with Exam Manager

Run a real exam with your own questions and see the results analysis on your own data — guided setup, no commitment.

Book a pilot

Share this post

Get the next article by email

Assessment and edtech articles, straight to your inbox. Double opt-in, unsubscribe anytime.