Interview scorecard template with an anchored rating scale
On this page
An interview scorecard is the form each interviewer fills in after an interview: the competencies the role needs, a rating for each against a written scale, the evidence behind every rating, and a hire recommendation. A good one is short enough to finish within the hour, anchored so a 3 means the same thing from every interviewer, and built so each score points to something the candidate actually said.
Below is a scorecard you can copy, a 1–4 rating scale with anchors, a filled example, and how to weight competencies without letting the arithmetic make the decision for you.
What goes on an interview scorecard
Every part of the form has a job. If a field does not change how someone rates or decides, cut it.
- Header. Candidate, role and level, interview stage, interviewer, date. The stage matters: a phone screen and a final round should not be scoring the same things.
- Competencies, four to six across the loop. Taken from the job, each with a one-line definition. The U.S. Office of Personnel Management's structured interview guide says a structured interview typically assesses between four and six competencies unless the job is unique or high-level.
- The questions for each competency. Written in advance and asked of every candidate, so ratings compare like with like.
- A rating scale with anchors. A description of what each score looks like, not "1 = poor, 4 = excellent".
- An evidence field. The two or three things the candidate said or did that justify the rating, in their words where you have them.
- An overall recommendation. Kept separate from the ratings, with a reason.
- A submitted time. Scorecards go in before the debrief, and the timestamp shows it.
The 1–4 anchored rating scale
Use one scale for every competency. Four points force a lean: there is no middle box to park a candidate you are unsure about, so an interviewer torn between 2 and 3 has to reread the evidence. OPM's guide asks for at least three levels and suggests aiming for five to seven. If your organization already uses five, keep them. The number of points matters less than whether each one is described.
| Score | Label | Anchor | What the evidence usually sounds like |
|---|---|---|---|
| 4 | Strong evidence | Specific, first-hand examples that meet the bar in situations at least as hard as the role; explained trade-offs and results without prompting | Named decisions they made, numbers, what they would do differently |
| 3 | Meets the bar | At least one specific, first-hand example that matches what the role needs; small gaps closed by follow-up questions | A real situation, their own actions, a result; needed a probe or two |
| 2 | Below the bar | Examples were general, hypothetical, about the team rather than them, or from a much simpler context; follow-ups did not close the gap | "We usually…", "I would…", no result, no clear personal role |
| 1 | No evidence, or the opposite | Could not give an example, or the example showed the opposite of the competency | Could not name a time; described skipping the step the role depends on |
| – | Not assessed | The question was not asked or time ran out | Leave it blank |
The "not assessed" row matters more than it looks. A competency nobody asked about that gets a polite 2 or a generous 3 turns a guess into a number, and that number ends up in the average.
Write question-specific anchors for each competency
The general anchors apply everywhere. What makes them usable is a short description of how each level answers your question. OPM's guide recommends writing example behaviors for each level of each competency with people who know the job well. Here is one for prioritization in an operations manager role, for the question "Tell me about a week when you had more work than time. What did you drop?"
| Score | What a candidate at this level describes |
|---|---|
| 4 | Names what they dropped and who they told before anything slipped; applied a stated rule (customer impact, safety, revenue); checked afterwards whether the call was right |
| 3 | Made an explicit choice and communicated it; the reasoning is sound but mostly reactive |
| 2 | Worked late to get everything done with no trade-off made; or describes a decision the team made with no personal role |
| 1 | Missed a commitment without warning anyone, or cannot name a time they had to choose |
Write these once at intake, with the hiring manager and someone who has done the job, and reuse them for every candidate in the search.
The interview scorecard template
Copy this into your ATS, a doc or a form. Delete the weights if you do not use them.
INTERVIEW SCORECARD
Candidate: [name] Role: [title, level]
Stage: [phone screen / onsite round 2 / final]
Interviewer: [name] Date: [date]
Competencies I own in this loop: [list]
RATING SCALE (same for every competency)
4 Strong evidence: specific, first-hand, above the bar
3 Meets the bar: specific, first-hand, matches the role
2 Below the bar: general, hypothetical, "we", or a simpler context
1 No evidence, or evidence of the opposite
– Not assessed (question not asked)
----------------------------------------------------------------
1. [Competency] Weight: [x]% [MUST-HAVE / nice-to-have]
Definition: [one line, in the words agreed at intake]
Question(s) asked: [Q1] [Q2]
Evidence (what they said or did; quotation marks only if exact):
- "[quote]" [timestamp, if you have one]
- [observable behavior, e.g. asked for the numbers before answering]
Gaps or counter-evidence:
- [what was missing; what the follow-up did not resolve]
Rating: [1 / 2 / 3 / 4 / –]
Why this rating: [one sentence linking the evidence to the anchor]
[repeat for each competency you own]
----------------------------------------------------------------
OVERALL
Recommendation: [Strong yes / Yes / No / Strong no]
Main reason: [one or two sentences that point to the ratings above]
What the next interviewer should probe: [a specific gap]
Not assessed: [competency, and why]
Submitted: [date, time]
(Submit before the debrief and before reading anyone else's scorecard.)
There is no "maybe" in the recommendation on purpose. An interviewer who cannot choose is telling you the evidence is thin, and the useful thing to write is what is missing, in "what the next interviewer should probe".
A filled example
An invented candidate for a mid-market customer success manager role, written the way a strong scorecard reads:
Candidate: Marcus Bell · Role: Customer Success Manager, mid-market · Stage: onsite, round 2 · Interviewer: Jordan Reyes
| Competency (weight) | Evidence | Rating |
|---|---|---|
| Renewal ownership, must-have (30%) | "I had 38 accounts and I owned the renewal number, not just the relationship." Walked through a renewal where the champion left: found a new sponsor through the finance contact and ran an executive review 90 days out. Gap: could not give the renewal rate for the whole book. | 3 |
| Handling escalations, must-have (25%) | Outage call with an angry VP: "I told him what we knew, what we didn't, and when he'd hear from me next, and then I called at four like I said." Sent a written timeline the same day. | 4 |
| Product feedback loop (20%) | "We'd send feedback to product." Could not name one request they had pushed, who received it, or what happened to it, after two follow-ups. | 2 |
| Using account data (15%) | Flags accounts when weekly logins drop; built a usage-by-seat sheet to pick expansion targets and described one expansion it led to. | 3 |
| Working with sales (10%) | Not asked; ran out of time. | – |
Recommendation: Yes. Both must-haves at 3 or above with first-hand examples; the escalation answer was specific down to the follow-up call.
Probe next: product feedback. Ask for one request, who they took it to, and the outcome.
Not assessed: working with sales. Round 3 owns that question.
Each rating can be checked by someone who was not in the room. The 2 is not a judgment of the person; it names a specific gap and hands it to the next interviewer. For more worked examples of the evidence lines, see internal interview feedback examples.
How to weight competencies
Weights are decided at intake, before anyone has met a candidate. Set afterwards, they tend to bend toward the person the team already likes.
- Treat must-haves as gates. A 1 on a must-have is a no, whatever the total. A 2 on a must-have needs a named plan, such as another conversation or a work sample, before a yes.
- Keep weights coarse. Multiples of 5 or 10 percent. A 23% weight suggests a precision that interview ratings do not have.
- Do not count what was not assessed. Divide by the weight you actually covered, or wait for the interviewer who owns that competency.
Using the example above, with "working with sales" not assessed:
Renewal ownership 0.30 x 3 = 0.90
Handling escalations 0.25 x 4 = 1.00
Product feedback 0.20 x 2 = 0.40
Using account data 0.15 x 3 = 0.45
----------------
Weight covered: 0.90 Total: 2.75
Weighted score: 2.75 / 0.90 = 3.06
The weighted score starts the conversation; it does not end it. Two candidates at 3.0 can be very different: one with steady 3s, another with two 4s and a 2 on a must-have. When you have several finalists, the side-by-side comparison lives in how to compare candidates after interviews.
Scorecard mistakes and the fix
| Mistake | Why it fails | Fix |
|---|---|---|
| Ten competencies on one card | Nobody can ask about ten things in an hour, so the last few are rated from impression | Four to six across the loop, two or three per interviewer |
| "1 = poor, 5 = excellent" | One interviewer's 4 is another's 3 | Describe what each score looks like for each competency |
| No evidence field | A rating with no reason cannot be challenged or defended | At least one quote or observation per rating |
| "Culture fit" as a competency | It cannot be anchored and it rewards similarity | Name the behavior, such as "gives direct feedback to peers" |
| Rating what was not asked | A guess becomes a number in the average | Use "not assessed" |
| Filling it in after the debrief | Ratings drift toward the most confident voice in the room | Lock submissions before the meeting starts |
| One card for every stage | The phone screen ends up judging what only an onsite can test | A shorter card per stage |
What interviewers write in the evidence field is also a record your company keeps. The phrases to keep out of it are covered in what not to write in interview notes.
Adapting the scorecard to your interview loop
- Recruiter phone screen: two or three must-haves plus logistics, and a recommendation of advance or do not advance. Keep it to one screen.
- Onsite or panel loop: give each interviewer two or three competencies so the whole list is covered once, not five times. The assignment belongs in a panel interview plan.
- Work sample or technical exercise: rate the work, not the conversation about it. Anchors describe the exercise: what a 3 solution includes, what a 4 adds.
- The hiring manager's own round: score the competencies you own, usually role judgment and motivation, and resist re-scoring what other interviewers covered.
- No ATS: one shared document per candidate, with each interviewer's section added before the debrief. The rest of that setup is in running a structured interview without an ATS.
Fill it in within the hour
OPM's guide tells interviewers to review their notes and rate the candidate immediately after the interview, and says ratings should be supported by actual behavioral examples. In practice that means booking fifteen to twenty minutes after every interview, before the next meeting takes over.
- Copy the evidence from your notes into each competency, before you think about a score.
- Rate each competency against its anchors, not against the last candidate you saw.
- Write the recommendation last, and make its reason point to the ratings.
- Submit, then read anyone else's scorecard.
How to turn what a candidate said into evidence that holds up is its own skill, covered in evidence-based interview feedback. How the submitted scorecards get used in the room is in the interview debrief template. If you run interviews with Interview Signal, the scorecard is drafted from the call and every score carries evidence quotes checked against the transcript, so your review starts from the evidence.
Questions people ask
Should an interview scorecard use a 1–4 or a 1–5 scale?
Either works if every point has a written anchor. A 1–4 scale removes the middle box, so an interviewer who is unsure has to go back to the evidence. If your organization already uses five levels, keep them and write anchors for each rather than switching.
How many competencies should one scorecard have?
Four to six for the whole loop, and two or three per interviewer. Beyond that, interviewers run out of time to ask about each one and start rating the last few from general impression.
Should interviewers see each other's scorecards before the debrief?
No. Each interviewer should submit before reading anyone else's ratings. Once you have seen a colleague's 2, your own 3 starts to look generous, and the debrief loses the independent views it exists to compare.
What if a candidate scores well overall but low on one must-have?
Treat must-haves as gates, not as part of the average. A 1 on a must-have is a no regardless of the total. A 2 means the team needs a specific plan to close the gap, such as another conversation or a work sample, before anyone says yes.