How to

Evidence-based interview feedback: the method, step by step

On this page
  1. Observation, interpretation, rating: three layers
  2. What the research says about structure and evidence
  3. The method, step by step
  4. What counts as evidence, and what does not
  5. A worked example: one answer, three pieces of feedback
  6. Calibrating interviewers
  7. Rating errors, and the evidence habit that counters each
  8. Questions people ask

Evidence-based interview feedback means every rating you give is backed by something the candidate actually said or did, written down separately from what you concluded about it, and scored against a scale the team agreed before the interview. The test is simple: a colleague who was not in the room should be able to read your feedback and see why you chose that rating, even if they would have chosen differently.

This page is the method: how to move from observation to evidence to rating, what counts as evidence and what does not, a worked example, and how to calibrate a group of interviewers so a 3 means the same thing from all of them.

Observation, interpretation, rating: three layers

Most weak feedback collapses three different things into one adjective. "Strong ownership" is an interpretation with no observation under it and no rating attached. Keep the layers apart and write them in order.

  1. Observation. What the candidate said or did, as close to their words as you have. Nobody who was there would dispute it.
  2. Interpretation. What you think the observation shows about a competency. Reasonable people can disagree here, which is why it has to sit next to the observation.
  3. Rating. Where the interpretation lands on the agreed scale, with a reason tied to the anchor.
ObservationInterpretationEffect on rating
"I rewrote the on-call runbook after the outage I caused." Raised it before being asked about failures.Takes ownership of their own mistakesSupports a 3 or 4 on ownership
Said "we decided" four times describing the migration; when asked "what was your part?", named the rollback plan.Personal role narrower than the story first suggestedHolds the rating at 3 on technical leadership until another example shows more
Asked what the success metric was before proposing a plan in the case question.Structures an ambiguous problem before solving itSupports problem solving
Could not say what their team's quota was after two follow-ups.Either did not track it or chose not to share itCounter-evidence on pipeline management; note both readings

What the research says about structure and evidence

The case for structured, evidence-based ratings comes from decades of selection research. In a widely cited meta-analysis, Schmidt and Hunter (1998, Psychological Bulletin) reported mean validity estimates, meaning correlations with later job performance, of .51 for structured interviews and .38 for unstructured ones.

In 2022, Sackett, Zhang, Berry and Lievens (Journal of Applied Psychology) argued that earlier corrections for range restriction had overstated many of those figures. Their revised estimates were .42 for structured interviews and .19 for unstructured interviews, and structured interviews came out as the top-ranked selection procedure in their comparison.

The U.S. Office of Personnel Management's Structured Interviews: A Practical Guide makes the practical version of the same point. It says structured interviews have shown high reliability, validity and legal defensibility, that unstructured interviews are more open to legal challenge, and that ratings should be supported by notes containing actual behavioral examples.

Two cautions. These are averages across many studies, not a forecast for your process. And "structured" in this research means the same questions and the same scoring for every candidate, not just a tidier form. Evidence-based feedback is the interviewer's half of that bargain.

The method, step by step

1. Agree what good looks like before the first interview

You cannot rate against an anchor that does not exist. Before the search starts, the hiring team agrees four to six competencies, the questions for each, and a description of what each score looks like. The artifact for this is an interview scorecard with an anchored rating scale. Without it, interviewers rate against their own private picture of the job, and calibration has nothing to hold on to.

2. Record observations, not conclusions

During the interview, write down what you would need later to justify a rating: short verbatim fragments, numbers, names of systems and people's roles, what the candidate did first, and what they could not answer. Write "said 'I should have flagged it in week four'" rather than "self-aware". The practical side of doing this while still listening is in how to take interview notes.

If you use a transcript, the quotes are easy; the discipline is still choosing which ones matter. Interview Signal, for example, checks each evidence quote against the transcript. Whatever tool you use, the interviewer still owns the judgment about what each quote shows.

3. Sort the evidence by competency, straight after the interview

Go through your notes and tag each observation to the one competency it best shows. Count each piece of evidence once: a great story about a migration should not earn points under technical depth, ownership and communication at the same time unless it genuinely shows each. Write down the counter-evidence too. A rating built only on supporting evidence is an argument, not an assessment.

4. Rate against the anchor, then write the reason

Read the anchor for each level and pick the one the evidence matches. Do not rate against the previous candidate or against yourself. Then write one sentence that a colleague could check:

[Competency]: [rating]
Evidence: [observation 1, quoted if exact]; [observation 2]
Why this rating: [how the evidence matches the anchor for this level]
Not higher because: [the gap, or what was missing]
Not lower because: [the evidence that rules out the level below]

The "not higher" and "not lower" lines do most of the work. They force you to look at the levels either side of your choice, which is where most disagreements in a debrief come from.

5. Write the recommendation last

The hire or no-hire recommendation comes after the ratings and points to them. If your recommendation disagrees with your own ratings, say why in a sentence. Sometimes that sentence reveals a must-have the scorecard missed; more often it reveals a feeling looking for a reason.

What counts as evidence, and what does not

TypeCounts?How to use it
A specific past example with their own actions and a resultYes, the strongest kindRecord the situation, what they did, and the outcome
A number they gave (team size, quota, volume, budget)Yes, as statedWrite "stated" next to it; verify later if the decision depends on it
Something they did in the interview (asked a clarifying question, corrected their own error in an exercise)YesDescribe the action, not the trait
A hypothetical "I would…"PartlyFine for situational questions with situational anchors; for behavioral questions, probe for a real instance
"We" throughoutWeak until probedAsk "what was your part?" and record that answer
A restatement of the resumeNoThe resume says they held the job, not how they did it
How confident or polished they soundedNot on its ownOnly where the competency is about communicating, and then describe what they did
A gut feeling, or "reminds me of…"NoFind the observation behind it or leave it out
What another interviewer said in the hallwayNoTheir evidence belongs on their scorecard

Some observations do not belong in feedback at all, however accurate they are, such as anything touching a protected characteristic. That list is in what not to write in interview notes.

A worked example: one answer, three pieces of feedback

An invented candidate, Sam Whitfield, answering "Tell me about a time a project you owned was going to miss its deadline":

"Our data migration was ten weeks and by week six we were two weeks behind. I pulled the three riskiest tables forward, told the product lead on a Tuesday we'd launch without the archive import, and we went live on the date. Archive came two weeks later. Honestly, I should have flagged it in week four. I was hoping we'd catch up."

Feedback A, all interpretation:

Great communicator, very honest, handles pressure well. Strong hire.

Nothing here can be checked, and nothing is tied to a competency or a rating. It is also the easiest kind to write at the end of a long day of interviews.

Feedback B, better but still conclusions:

Showed good judgment in prioritizing scope and was transparent with stakeholders. Planning: 4.

It names a competency and a rating, but a reader cannot tell what "good judgment" was, and the late flag, which argues against a 4, has disappeared.

Feedback C, evidence-based:

Planning and prioritization: 3. In week six of a ten-week migration, two weeks behind, moved the three riskiest tables forward and cut the archive import from launch; told the product lead directly and launched on the date. Not a 4: by their own account they raised the slip about two weeks later than they should have ("I should have flagged it in week four"). Not a 2: made an explicit scope trade-off and communicated it before the deadline.

Feedback C takes a minute longer to write and is far more useful. Another interviewer can disagree with the 3, but they will be disagreeing about the same facts. More pairs like this, across other competencies, are collected in internal interview feedback examples.

Calibrating interviewers

Anchors written on paper still get read differently. Calibration is how a team finds out where, and it works best when it is about evidence rather than about whose judgment is better.

Before the search

  • Give every interviewer the same sample answer, written out or from a mock interview recorded with the participant's agreement.
  • Each person rates it alone, with evidence and a reason, then the group compares.
  • Where ratings differ by two points, find the anchor wording that allowed both readings and rewrite it.

During the search

  • In the debrief, a spread of two or more points on one competency is a calibration signal, not only a disagreement about the candidate. The facilitation for that is in the interview debrief template.
  • New interviewers shadow an experienced one and rate independently, then compare scorecards before they interview alone.

After every batch of scorecards

  • Look at each interviewer's ratings across candidates. Someone who has never given a 1 or a 4, or who gives almost everyone a 4, is reading the scale differently from the rest.
  • Share the pattern privately and with examples from their own scorecards. The fix is almost always an anchor or a question, not the person.

Rating errors, and the evidence habit that counters each

OPM's guide lists common rating errors and interviewing mistakes. Each one has a matching habit in the method above.

ErrorWhat it looks like in feedbackHabit that counters it
Halo effectOne excellent answer lifts every ratingRate each competency only from evidence tagged to it
Similar to me"Reminds me of myself early on"Ask what observation supports it; drop it if none
Central tendencyA row of 3sWrite "not higher because" and "not lower because" for each
Leniency or strictnessEveryone gets a 4, or nobody doesReview your own ratings across candidates
First impressionsThe rating was settled in the first five minutesRecord evidence all the way through; rate only afterwards
Contrast effect"Much better than this morning's candidate"Rate against the anchor, never against another candidate
Negative emphasisOne weak moment outweighs three strong onesRecord supporting and counter-evidence side by side

None of these errors go away with awareness alone. They get smaller when the form makes you write down what you saw before it lets you write down what you think.

Questions people ask

Is evidence-based feedback the same as a structured interview?

They go together but are not the same. A structured interview standardizes the questions and the rating scale; evidence-based feedback is how each interviewer fills that scale in, with observations that another person could check. You can write evidence-based feedback in a loosely structured interview, but it is much harder to compare across candidates.

Can a hypothetical answer count as evidence?

Yes, if the question was designed as a situational question and you rate it against anchors written for that scenario. If you asked for a real past example and got 'I would…', treat it as weak evidence and probe for an actual instance before rating.

How much evidence does one rating need?

At least one specific observation, and preferably two, for each competency you rate. If you cannot write down a single thing the candidate said or did that supports the score, mark the competency not assessed instead.

What should I do with a strong gut feeling I cannot back up?

Look for the observation behind it. Often there is one, such as an evasive answer to a direct question, and you can write that down. If you cannot find one, leave the feeling out of the scorecard and, if it still bothers you, suggest a specific follow-up question for the next interviewer.