Interview calibration: how to run a session that lines up your interviewers
On this page
- What calibration is, and what else goes by the name
- What to prepare
- The 60-minute session agenda
- Reading the scores: look at the bar, not the average
- Facilitating the discussion
- Shadowing and reverse shadowing new interviewers
- Reviewing drift in real scorecards
- Using recordings for calibration: consent and privacy
- Questions people ask
Interview calibration is a short, regular session where interviewers score the same answer on their own, compare the scores, and fix the anchor wording that let them disagree. The aim is that a 3 means the same thing from every interviewer on the loop. A good session takes an hour, uses sample answers rather than live candidates, and ends with a revised scorecard, not just a discussion.
This page covers the session itself (agenda, materials, facilitation), how to read the scores, how to bring new interviewers in with shadowing and reverse shadowing, and how to spot drift in real scorecards afterwards. The scale being calibrated is the one in the interview scorecard template, and the write-up method is evidence-based interview feedback.
What calibration is, and what else goes by the name
Three different meetings get called calibration. This page is about the first.
| Meeting | Who | Question it answers | Where it is covered |
|---|---|---|---|
| Interviewer calibration | Everyone who scores candidates for a role or role family | Do we read the rating scale the same way? | This page |
| Profile calibration | Recruiter and hiring manager, early in a search | Is this the kind of candidate you want? | Intake meeting questions |
| Debrief | The interview panel for one candidate | Do we hire this person? | Interview debrief template |
The research case for the first is solid. A meta-analysis by Conway, Jako and Goodman (1995, Journal of Applied Psychology) of 111 interrater reliability coefficients found that interviewer training and standardization of questions, of how answers are evaluated, and of how ratings are combined all moderated how closely interviewers agreed. Multiple ratings helped when combined mechanically; there was no evidence they helped when combined subjectively.
The method itself is close to what psychologists call frame-of-reference training: raters learn what each level of performance looks like, practice rating examples, and get feedback. An updated meta-analysis, Roch, Woehr, Mishra and Kieszczynska (2012), concluded it is an effective way to improve rating accuracy. Most of those studies are about performance appraisal rather than interviews, so treat it as strong support for the method, not a promise of a number. OPM's structured interview guide puts the same idea in its sample training plan: critiqued practice using a recorded interview.
What to prepare
- The scorecard with anchors, printed or shared, one competency per page. Calibrate one or two competencies per session, not the whole kit.
- Three sample answers per competency: one clearly strong, one clearly weak, and one on the line between "below the bar" and "meets the bar". The borderline answer does most of the work; the other two check that the scale has ends.
- An answer key, written by the kit owner before the session: the intended score for each sample and the anchor line that decides it. Keep it hidden until scores are in.
- A scoring sheet for each interviewer (below), or a form that hides everyone's answers until all are submitted.
- A facilitator who owns the kit and can edit anchors on the spot. It does not have to be the most senior person, and it is better if it is not.
Where the sample answers come from matters:
- Written answers are quick to make and easy to tune to the borderline. Write them in a candidate's voice, with the hesitations and "we" statements real answers have.
- Mock interviews recorded with colleagues who agreed to it show delivery as well as content, which is closer to the real task.
- Real candidate recordings or transcripts are the most realistic and the hardest to justify. See the section on recordings below before using one.
Copy this scoring sheet, one per interviewer per answer:
CALIBRATION SCORING SHEET
Interviewer: [name] Session: [date] Kit version: [v]
Sample answer: [#] Competency: [name]
Score (1-4, or NA = not assessed): [ ]
Confidence: sure / torn between [ ] and [ ]
Evidence, in the candidate's words:
1. "[ ]"
2. "[ ]"
Anchor line that decided the score: "[copy the words]"
Not higher because: [ ]
Not lower because: [ ]
Probe I would have asked next: [ ]
The 60-minute session agenda
| Minutes | Step | Output |
|---|---|---|
| 0–5 | Purpose and rules: we are calibrating the scale, not grading each other; scores stay private until everyone has submitted | Agreement on rules |
| 5–10 | Everyone reads the competency definition and anchors in silence | No discussion yet |
| 10–20 | Sample answers 1 and 2 (strong and weak): read or play, score alone, submit | Two sets of sheets |
| 20–25 | Reveal all scores at once; confirm the ends of the scale work | Any anchor that failed at the ends |
| 25–45 | Sample answer 3 (borderline): score alone, reveal, discuss the widest spread first | Anchor lines rewritten live |
| 45–55 | A fresh borderline answer, scored alone against the rewritten anchors | Evidence the rewrite helped, or not |
| 55–60 | Record changes: new kit version, anchor edits, who needs shadowing | Kit v[n+1] and an action list |
The second borderline answer is the easiest step to drop when time runs short. Without it you do not know whether the discussion changed how people score or only how they talk about scoring.
Reading the scores: look at the bar, not the average
An invented session: six interviewers score three sample answers on "prioritization" using a 1–4 scale, where 3 is "meets the bar". The answer key says A is a 4, B is a 2 and C is a 3.
| Answer | Key | Scores from six interviewers | Average | Spread | Same side of the bar as the key |
|---|---|---|---|---|---|
| A (strong) | 4 | 4, 4, 3, 4, 4, 4 | 3.83 | 1 | 6 of 6 |
| B (weak) | 2 | 2, 2, 1, 2, 3, 2 | 2.00 | 2 | 5 of 6 |
| C (borderline) | 3 | 3, 2, 4, 3, 2, 3 | 2.83 | 2 | 4 of 6 |
The arithmetic: A is 23 ÷ 6 = 3.83; B is 12 ÷ 6 = 2.00; C is 17 ÷ 6 = 2.83.
What it tells you:
- The averages look fine and hide the problem. Answer C averages close to the key, but two of six interviewers would have scored a candidate who meets the bar as below it. On a real loop, that is the difference between an offer and a rejection.
- Answer B has one generous reader. The interviewer who gave a 3 is worth hearing first: what did they see that the anchor allowed?
- A spread of two or more on any answer points to the anchor, not the people. Find the words that let both readings stand and rewrite them.
- If everyone agrees with each other but not with the key, the key may be wrong. Say so and change it.
Track one number across sessions: the share of scores on the same side of the bar as the key. Here it is 15 of 18, or 83%. Whether that is good enough is your call; the useful thing is whether it rises after you rewrite anchors.
Facilitating the discussion
The discussion should be about words on the page: the candidate's words and the anchor's words. The facilitator keeps it there. A script that works:
OPENING
"We're checking the scale, not each other. Nobody's score is wrong
until we find the anchor line that decides it."
REVEAL
"Everyone's scores are up. Let's start with answer C, where we're
furthest apart. [Name], you gave a 2. Read us the evidence you wrote
down, in the candidate's words."
"[Name], you gave a 4. Same question."
FIND THE ANCHOR
"Which line on the anchor did that evidence meet?"
"Is there anything in the anchor that rules out the other reading?"
FIX IT
"If both readings are reasonable, the anchor is ambiguous. What
would we add so the next person can't read it both ways?"
[Edit the anchor on screen. Read the new line aloud.]
WHEN SOMEONE ARGUES FROM EXPERIENCE
"That may be right for the role. If it is, it goes into the anchor
for everyone. What would the candidate have to say to show it?"
CLOSE
"Changes to the kit: [list]. New version is v[n]. Next real
scorecards use it from [date]."
Rules that keep it productive:
- Nobody changes their submitted score. The sheet records how the scale was read before the discussion, which is the data you need.
- Quote, do not characterize. "She said 'I dropped the vendor report'" beats "she prioritized well".
- The most senior person speaks last, for the same reason as in a debrief: once they have spoken, everyone else's reading moves.
- Every disagreement ends in an edit or a decision not to edit, written down.
Shadowing and reverse shadowing new interviewers
A calibration session teaches the scale. Shadowing teaches the interview: asking the kit's questions as written, probing without leading, and keeping time. Use two stages.
Stage 1: shadow
The new interviewer joins an experienced interviewer's real interviews as an observer, usually two. Both take notes and both submit a scorecard independently before comparing. Tell the candidate in the invite that a colleague in training will observe, and introduce them at the start.
Stage 2: reverse shadow
The new interviewer leads; the experienced interviewer observes, scores independently, and gives feedback afterwards on both the scores and the conduct of the interview. Only one person's scorecard counts toward the decision, decided in advance, so the candidate is not scored twice by the same panel seat.
When they are ready
There is no research threshold for this; set your own rule and write it down. A reasonable starting point: at least two reverse shadows where the new interviewer asked the kit's questions as written and landed on the same side of the bar as the experienced interviewer on every competency.
SHADOW / REVERSE SHADOW RECORD
New interviewer: [ ] Experienced interviewer: [ ]
Role / kit version: [ ] Type: shadow / reverse shadow #[ ]
Candidate reference: [ID, not name] Date: [ ]
Scorecard that counts for the decision: [name]
Competency New Experienced Same side of bar? Note
[C1] [ ] [ ] Y / N [ ]
[C2] [ ] [ ] Y / N [ ]
[C3] [ ] [ ] Y / N [ ]
Both scorecards submitted before comparing: Y / N (time: [ ])
Kit questions asked as written: Y / N Missed: [ ]
Probes: from the kit / improvised Leading questions: [ ]
Time kept: Y / N
Off-limits topics avoided: Y / N Detail: [ ]
Candidate experience (interruptions, tone, close): [ ]
Outcome: ready to interview alone / another reverse shadow
Signed: [experienced interviewer]
On a video panel, make sure shadowing does not become copying. Interview Signal, for example, syncs only which questions have been covered between panelists, not transcripts or notes, so each person's scorecard stays their own.
Reviewing drift in real scorecards
Interviewers who calibrated well in March can drift by June: a new favorite question, a tougher reading of one anchor, or fatigue in a busy quarter. After every batch of scorecards, or once a quarter, put each interviewer's ratings side by side. An invented example for one interview stage:
| Interviewer | Candidates scored | Ratings given | Ratings of 4 | Recommended to advance |
|---|---|---|---|---|
| Alex | 18 | 54 | 24 (44%) | 15 (83%) |
| Bea | 20 | 60 | 9 (15%) | 10 (50%) |
| Chen | 16 | 48 | 9 (19%) | 7 (44%) |
The arithmetic: 24 ÷ 54 = 44%, 9 ÷ 60 = 15%, 9 ÷ 48 = 19%; 15 ÷ 18 = 83%, 10 ÷ 20 = 50%, 7 ÷ 16 = 44%. Alex gives roughly three times as many 4s as Bea and recommends far more candidates.
Before concluding anything, check:
- Same stage, same pool? If Alex only sees candidates who already passed two interviews, a higher rate is expected. Compare interviewers on the same stage.
- Enough scorecards? With a handful each, differences are noise. Wait for more.
- Which competency? Drift is usually one anchor, not the whole scale. Break the 4s down by competency.
- What does the evidence say? Pull three of Alex's 4s and read the evidence against the anchor. Sometimes the evidence supports them and the others are strict.
Share the pattern privately, with examples from the interviewer's own scorecards, and bring the anchor that moved to the next calibration session. Other signals worth watching: a competency that is marked "not assessed" often (the question does not fit the time, or people skip it), and a question whose answers start to sound rehearsed (it may have been shared online and needs replacing).
Using recordings for calibration: consent and privacy
This is general information, not legal advice. Whether you can reuse a candidate's recording depends on what they agreed to and the law where you and they are. Check with your legal or privacy team.
A real interview recording is the most realistic calibration material you can find, which is why teams reach for it. The problem is purpose. A typical consent line, like the one in our interview recording consent script, tells the candidate the recording is used for this hiring process. Playing it to a room of interviewers months later is a different use. Under UK and EU data protection law, a new purpose for personal data generally has to be compatible with the original one or have its own lawful basis (see recording interviews under UK GDPR), and elsewhere it is still a matter of trust.
Options, from simplest to most involved:
- Write the sample answers. No consent question, and you control where the borderline sits.
- Record mock interviews with colleagues playing candidates, with their agreement. Rotate who plays the candidate so answers do not all sound alike.
- Use a short, de-identified transcript excerpt with names, employers and other identifying details removed, if your privacy team agrees.
- Ask the candidate for training use specifically, separately from the interview consent, after the process ends, and make clear that saying no changes nothing.
Whatever you use, keep calibration material in one access-controlled folder, list what is in it, and delete real-candidate material when the session is done. The consent guide has notice wording by jurisdiction if you decide to ask.
Questions people ask
How often should we run interview calibration?
Run a full session when a new interview kit is launched, when several new interviewers join, and when a drift review shows a problem. Between those, a short check at the start of a busy hiring period, scoring one sample answer together, is usually enough.
Can we use recordings of real candidates in calibration?
Only if the candidate agreed to that use. Many consent lines, including ours, say the recording is for this hiring process, which does not obviously cover training. Mock interviews recorded with colleagues, or written sample answers, avoid the question entirely.
What if the hiring manager is the outlier?
Treat it the same way as anyone else: compare the evidence and the anchor wording, not seniority. If the manager's reading of an anchor is what the role actually needs, rewrite the anchor for everyone. If not, the manager scores against the agreed anchor like every other interviewer.
Is calibration the same as a debrief?
No. A debrief decides about a real candidate. Calibration uses sample answers to check that interviewers read the scale the same way, before or between real decisions. Mixing them turns a debrief into a training session and a real candidate into practice material.