Templates

Interview rating scale: how many points, how to anchor them, how to calibrate

On this page
  1. The parts of an interview rating scale
  2. 1–4, 1–5, 1–7 or 1–10: choosing the number of points
  3. Behaviorally anchored rating scales
  4. Anchors written for three competencies
  5. Rules for writing anchors interviewers can use
  6. Decision rules that belong with the scale
  7. Calibrating the scale before and after launch
  8. Rating scale mistakes and the fix
  9. Questions people ask

An interview rating scale is the set of numbers interviewers use to score a competency, together with a written description of what each number means. The number of points matters less than the descriptions: a 1–4 or 1–5 scale where every level is anchored to observable evidence gives more consistent ratings than a 1–10 scale labeled "poor" to "excellent". Choose a range, write behavioral anchors for each level, add a "not assessed" option and a few decision rules, then calibrate interviewers on real answers before the scale goes live.

This page is about designing the scale itself. The form it sits on is the interview scorecard template, and ready-made question-level anchors are in behavioral interview questions with a scoring rubric.

The parts of an interview rating scale

  1. A range, the same for every competency. Mixing a 1–4 for one competency with a 1–5 for another makes every comparison harder.
  2. Labels for each level, short enough to remember: "strong evidence", "meets the bar".
  3. General anchors: what each level means for any competency.
  4. Competency or question anchors: what each level sounds like for a specific competency or question.
  5. A "not assessed" option, so a question nobody asked does not become a guess in the average.
  6. Decision rules: how must-haves work, how interviewers' ratings combine, and whether half points are allowed.

The U.S. Office of Personnel Management's structured interview guide describes the same core: one proficiency range for all competencies, at least three levels (it suggests aiming for five to seven), labels for at least three of them, and example behaviors for each level of each competency.

1–4, 1–5, 1–7 or 1–10: choosing the number of points

RangeWhat it does wellWhat goes wrongSuits
1–3Very quick; easy to anchor ("not shown, partly shown, shown")Too coarse to separate several strong finalistsPhone screens and knockout checks
1–4No middle box, so an unsure interviewer has to return to the evidence; four anchors are realistic to writeLoses a distinction between "clearly above the bar" and "exceptional"Most structured interview loops
1–5Familiar; matches OPM's example scale; room for an exceptional levelThe 3 becomes a parking spot unless its anchor is as specific as the othersTeams already on five points; roles where "exceptional" matters
1–7More room to separate candidatesFew teams can write seven levels an interviewer can tell apart in the momentTrained panels that interview often for one role family
1–10Feels preciseRarely anchored; interviewers use 6 to 8 for nearly everyone; hard to calibrateNot recommended for interview ratings

What does research say? Most studies on the number of scale points come from surveys, not interviews. In one widely cited study, Preston and Colman (2000, Acta Psychologica) had 149 people rate a store or restaurant on scales with 2 to 11 points. The 2-, 3- and 4-point versions did relatively poorly on several reliability and validity indices, the indices rose with more categories up to about 7, and respondents preferred 10 points. Customers rating service are not trained interviewers rating evidence against written anchors, so read it as a caution rather than a rule: short scales throw away distinctions, and long scales only help if the extra levels are described.

A practical recommendation: use 1–4 if you are starting from nothing, 1–5 if your organization already uses five points, and put your effort into anchors either way.

Behaviorally anchored rating scales

A behaviorally anchored rating scale (BARS) replaces adjectives with examples of behavior at each level. The method is usually traced to Smith and Kendall (1963, Journal of Applied Psychology), who built rating scales anchored by examples of expected behavior. OPM's guide gives a version suited to interviews:

  1. Bring together people who know the job well, such as high performers and their supervisors.
  2. Each person writes, alone, how employees at each proficiency level would answer the question.
  3. The group discusses the examples and keeps the ones they agree best reflect each level.
  4. Interviewers use the examples as a guide rather than a script, because candidates' experiences differ.

For situational questions, OPM adds a check worth borrowing for any anchor: a separate group reads each example and says which competency it shows. Examples that are not clearly linked to one competency are dropped. An anchor that two people would file under different competencies will produce two different ratings.

The rating scale template

Copy this, pick the range, and fill in the brackets for each competency.

INTERVIEW RATING SCALE
Range: 1-4 (same for every competency)     Half points: not allowed

GENERAL ANCHORS
4  Strong evidence   Specific, first-hand examples in situations at least
                     as demanding as the role; reasoning and results
                     explained without prompting
3  Meets the bar     At least one specific, first-hand example that matches
                     the role; small gaps closed by probes
2  Below the bar     General, hypothetical, team-level ("we"), or from a much
                     simpler context; probes did not close the gap
1  No evidence       No example, or evidence of the opposite
-  Not assessed      Not asked, or ran out of time (excluded from averages)

COMPETENCY: [name]
Definition: [one sentence, agreed at intake]
4  [what a candidate at this level describes, for this role]
3  [...]
2  [...]
1  [...]

DECISION RULES
- Must-haves: a 1 is a no; a 2 needs a named follow-up before a yes
- Each interviewer submits before seeing other ratings
- Differences of 2+ points on one competency are discussed at the debrief
- Final rating per competency: [consensus after discussion / median]

If you use five points, add a level 5 above "strong evidence" for examples that exceed the role's scope, such as leading the work others in the role only take part in, and keep 3 as "meets the bar".

Anchors written for three competencies

These are competency-level anchors on a 1–4 scale, written for invented roles. Rewrite the context for your own job; keep the pattern of observable behavior at each level.

Handling customer escalations (support team lead)

ScoreWhat the candidate describesWhat the evidence usually sounds like
4Takes ownership of a serious escalation, sets expectations on timing, keeps the customer updated at agreed points, fixes the root cause with the product or engineering team, and changes a process so it recurs less"I told her she'd hear from me at two and four. The fix shipped Thursday, and we added a check to the release list."
3Resolves the escalation with clear updates to the customer and a sensible handoff to the team that fixes itA specific case, regular updates, a resolution; little on preventing a repeat
2Passes the escalation on quickly and loses track of it, or calms the customer without resolving the issue"I escalated it to tier 3 and they handled it."
1Cannot describe a serious escalation, or describes arguing with the customer, missing promised updates, or blaming another team to the customer"Honestly, some customers just can't be helped."

Coaching direct reports (first-time people manager)

ScoreWhat the candidate describesWhat the evidence usually sounds like
4Diagnoses the specific gap, agrees a goal with the person, gives regular specific feedback, adjusts the approach when it is not working, and can show the person's improvementA named skill, a plan with checkpoints, and a before-and-after: "her first-contact resolution went from the bottom of the team to the middle in a quarter"
3Gives specific feedback and support over several conversations, with a visible improvementA real person, a real skill, several conversations, a result
2Gives general encouragement, or coaches only by doing the work alongside the person; improvement not tracked"I have an open-door policy," "I showed him how I do it"
1No example of developing anyone, or describes avoiding a performance conversation until it became a formal issueCannot name someone they helped improve

Accuracy with financial data (accounts payable specialist)

ScoreWhat the candidate describesWhat the evidence usually sounds like
4Uses a consistent checking method, catches an error before it causes loss, traces the cause, and changes a control so it cannot recur"Two invoices had the same number from different entities. I added a vendor-plus-amount match to our weekly check."
3Has a checking routine and a specific example of catching and correcting an errorA named reconciliation or review step and one real catch
2Describes being careful in general terms; errors were caught by reviewers or auditors"I always double-check my work."
1No routine, or describes errors reaching payment or reporting without a change afterwardsCannot describe how they check a batch

Rules for writing anchors interviewers can use

Weak anchorProblemBetter anchor
"Very good communication"An adjective; every interviewer supplies their own meaning"Explained a technical decision in the listener's terms and checked they understood"
"Always meets deadlines"Frequency cannot be observed in a one-hour interview"Describes raising a slipping deadline before it was missed, with a new plan"
"Confident and proactive"Personality traits, not behavior; rewards interview polish"Started work on a problem outside their role and told the owner"
"Better than most candidates"Compares candidates instead of evidence to the jobDescribe the behavior; never reference other candidates
"Strong leadership and strategic thinking"Two competencies in one anchorOne competency per anchor; split them
Level 3 and level 4 differ only by "significantly"Interviewers cannot tell where the line isAdd a concrete difference: scope, independence, prevention of recurrence

A quick test for any anchor: read it to someone who has never met the candidate and ask them to describe an answer that would earn it. If their description matches yours, the anchor works.

Decision rules that belong with the scale

  • Must-haves are gates, not averages. A 4 on a nice-to-have should never cancel a 1 on a must-have.
  • No half points. Ask for a choice and a sentence on what was missing for the next level.
  • "Not assessed" is excluded, not scored as a 2. Divide by what was covered, or wait for the owner of that competency.
  • Independent first, then discuss. OPM's guide has panel members rate independently, then compare notes and explore the basis for discrepancies before reaching a consensus rating. Decide in advance whether your final number is a consensus or a median, and write it down.
  • Evidence is required. A rating without an observation behind it is returned to the interviewer.

Calibrating the scale before and after launch

Before the first candidate

  1. Pilot the questions. OPM recommends a trial run of new questions with colleagues to check wording and whether they draw a range of answers.
  2. Rate the same answers independently. Use two or three written or recorded answers, recorded with the participant's agreement. Each interviewer rates alone, with a reason.
  3. Find the anchor behind each disagreement. Where ratings differ by two points, identify the wording that allowed both readings and rewrite it.

After the first batch of interviews

Look at each interviewer's distribution of scores. An invented example after eight candidates, four competencies each (32 ratings per interviewer):

Interviewer1s2s3s4sMeanReading
A291562.78Uses the full scale
B0210203.56Possible leniency; check anchors for 3 vs 4
C042712.91Central tendency; ask for "not higher because" lines

The mean for A is (2×1 + 9×2 + 15×3 + 6×4) ÷ 32 = 89 ÷ 32 = 2.78. Interviewers may simply have met different candidates, so compare them on candidates they both saw before drawing conclusions. OPM's list of rating errors includes leniency, strictness, halo and central tendency; a distribution table is the quickest way to spot the first and last. A full session plan is in interview calibration.

Rating scale mistakes and the fix

MistakeWhat happensFix
Labels only, no anchorsOne interviewer's 4 is another's 3Behavioral anchors for each level of each competency
Different ranges per stage or interviewerScores cannot be compared or combinedOne range for the whole loop
Changing the scale mid-searchEarly and late candidates are rated on different rulesCollect changes and apply them to the next search
Anchors written after meeting candidatesThe top anchor describes the favoriteWrite and agree anchors at intake
Averaging across competencies with gates ignoredA must-have failure disappears into a good totalGates first, averages second
No calibrationThe scale looks structured but is used privatelyIndependent ratings of shared answers before launch

A scale is only as consistent as the evidence each rating is based on. How to record that evidence is in evidence-based interview feedback. Interview Signal's scorecards attach quotes from the transcript to each score, which makes calibration conversations about what the candidate said rather than what each interviewer remembers.

Questions people ask

Is a 1–5 or a 1–10 interview rating scale better?

For most teams, 1–4 or 1–5. Every point needs a written description that interviewers can tell apart, and very few teams can write ten distinct levels of evidence for a competency. A 1–10 scale usually ends up used as a 1–5 scale with extra decimals of false precision.

Should interviewers be allowed to give half points?

No. A 3.5 means the interviewer could not decide which anchor the evidence matches, and that is useful information you lose. Ask them to choose, and write what would have moved the score up or down.

Do we need different anchors for every question?

One general scale for every competency, plus short question-specific anchors for the questions you ask most often. The general scale keeps the meaning of each number constant; the specific anchors make it easy to apply in the moment.

How do we know if interviewers are using the scale the same way?

Have them rate the same recorded or written answer independently and compare, before the search and again every few months. Between sessions, look at each interviewer's spread of scores. Someone who never uses the top or bottom of the scale is reading it differently.