Interview rating scale: how many points, how to anchor them, how to calibrate
On this page
- The parts of an interview rating scale
- 1–4, 1–5, 1–7 or 1–10: choosing the number of points
- Behaviorally anchored rating scales
- Anchors written for three competencies
- Rules for writing anchors interviewers can use
- Decision rules that belong with the scale
- Calibrating the scale before and after launch
- Rating scale mistakes and the fix
- Questions people ask
An interview rating scale is the set of numbers interviewers use to score a competency, together with a written description of what each number means. The number of points matters less than the descriptions: a 1–4 or 1–5 scale where every level is anchored to observable evidence gives more consistent ratings than a 1–10 scale labeled "poor" to "excellent". Choose a range, write behavioral anchors for each level, add a "not assessed" option and a few decision rules, then calibrate interviewers on real answers before the scale goes live.
This page is about designing the scale itself. The form it sits on is the interview scorecard template, and ready-made question-level anchors are in behavioral interview questions with a scoring rubric.
The parts of an interview rating scale
- A range, the same for every competency. Mixing a 1–4 for one competency with a 1–5 for another makes every comparison harder.
- Labels for each level, short enough to remember: "strong evidence", "meets the bar".
- General anchors: what each level means for any competency.
- Competency or question anchors: what each level sounds like for a specific competency or question.
- A "not assessed" option, so a question nobody asked does not become a guess in the average.
- Decision rules: how must-haves work, how interviewers' ratings combine, and whether half points are allowed.
The U.S. Office of Personnel Management's structured interview guide describes the same core: one proficiency range for all competencies, at least three levels (it suggests aiming for five to seven), labels for at least three of them, and example behaviors for each level of each competency.
1–4, 1–5, 1–7 or 1–10: choosing the number of points
| Range | What it does well | What goes wrong | Suits |
|---|---|---|---|
| 1–3 | Very quick; easy to anchor ("not shown, partly shown, shown") | Too coarse to separate several strong finalists | Phone screens and knockout checks |
| 1–4 | No middle box, so an unsure interviewer has to return to the evidence; four anchors are realistic to write | Loses a distinction between "clearly above the bar" and "exceptional" | Most structured interview loops |
| 1–5 | Familiar; matches OPM's example scale; room for an exceptional level | The 3 becomes a parking spot unless its anchor is as specific as the others | Teams already on five points; roles where "exceptional" matters |
| 1–7 | More room to separate candidates | Few teams can write seven levels an interviewer can tell apart in the moment | Trained panels that interview often for one role family |
| 1–10 | Feels precise | Rarely anchored; interviewers use 6 to 8 for nearly everyone; hard to calibrate | Not recommended for interview ratings |
What does research say? Most studies on the number of scale points come from surveys, not interviews. In one widely cited study, Preston and Colman (2000, Acta Psychologica) had 149 people rate a store or restaurant on scales with 2 to 11 points. The 2-, 3- and 4-point versions did relatively poorly on several reliability and validity indices, the indices rose with more categories up to about 7, and respondents preferred 10 points. Customers rating service are not trained interviewers rating evidence against written anchors, so read it as a caution rather than a rule: short scales throw away distinctions, and long scales only help if the extra levels are described.
A practical recommendation: use 1–4 if you are starting from nothing, 1–5 if your organization already uses five points, and put your effort into anchors either way.
Behaviorally anchored rating scales
A behaviorally anchored rating scale (BARS) replaces adjectives with examples of behavior at each level. The method is usually traced to Smith and Kendall (1963, Journal of Applied Psychology), who built rating scales anchored by examples of expected behavior. OPM's guide gives a version suited to interviews:
- Bring together people who know the job well, such as high performers and their supervisors.
- Each person writes, alone, how employees at each proficiency level would answer the question.
- The group discusses the examples and keeps the ones they agree best reflect each level.
- Interviewers use the examples as a guide rather than a script, because candidates' experiences differ.
For situational questions, OPM adds a check worth borrowing for any anchor: a separate group reads each example and says which competency it shows. Examples that are not clearly linked to one competency are dropped. An anchor that two people would file under different competencies will produce two different ratings.
The rating scale template
Copy this, pick the range, and fill in the brackets for each competency.
INTERVIEW RATING SCALE
Range: 1-4 (same for every competency) Half points: not allowed
GENERAL ANCHORS
4 Strong evidence Specific, first-hand examples in situations at least
as demanding as the role; reasoning and results
explained without prompting
3 Meets the bar At least one specific, first-hand example that matches
the role; small gaps closed by probes
2 Below the bar General, hypothetical, team-level ("we"), or from a much
simpler context; probes did not close the gap
1 No evidence No example, or evidence of the opposite
- Not assessed Not asked, or ran out of time (excluded from averages)
COMPETENCY: [name]
Definition: [one sentence, agreed at intake]
4 [what a candidate at this level describes, for this role]
3 [...]
2 [...]
1 [...]
DECISION RULES
- Must-haves: a 1 is a no; a 2 needs a named follow-up before a yes
- Each interviewer submits before seeing other ratings
- Differences of 2+ points on one competency are discussed at the debrief
- Final rating per competency: [consensus after discussion / median]
If you use five points, add a level 5 above "strong evidence" for examples that exceed the role's scope, such as leading the work others in the role only take part in, and keep 3 as "meets the bar".
Anchors written for three competencies
These are competency-level anchors on a 1–4 scale, written for invented roles. Rewrite the context for your own job; keep the pattern of observable behavior at each level.
Handling customer escalations (support team lead)
| Score | What the candidate describes | What the evidence usually sounds like |
|---|---|---|
| 4 | Takes ownership of a serious escalation, sets expectations on timing, keeps the customer updated at agreed points, fixes the root cause with the product or engineering team, and changes a process so it recurs less | "I told her she'd hear from me at two and four. The fix shipped Thursday, and we added a check to the release list." |
| 3 | Resolves the escalation with clear updates to the customer and a sensible handoff to the team that fixes it | A specific case, regular updates, a resolution; little on preventing a repeat |
| 2 | Passes the escalation on quickly and loses track of it, or calms the customer without resolving the issue | "I escalated it to tier 3 and they handled it." |
| 1 | Cannot describe a serious escalation, or describes arguing with the customer, missing promised updates, or blaming another team to the customer | "Honestly, some customers just can't be helped." |
Coaching direct reports (first-time people manager)
| Score | What the candidate describes | What the evidence usually sounds like |
|---|---|---|
| 4 | Diagnoses the specific gap, agrees a goal with the person, gives regular specific feedback, adjusts the approach when it is not working, and can show the person's improvement | A named skill, a plan with checkpoints, and a before-and-after: "her first-contact resolution went from the bottom of the team to the middle in a quarter" |
| 3 | Gives specific feedback and support over several conversations, with a visible improvement | A real person, a real skill, several conversations, a result |
| 2 | Gives general encouragement, or coaches only by doing the work alongside the person; improvement not tracked | "I have an open-door policy," "I showed him how I do it" |
| 1 | No example of developing anyone, or describes avoiding a performance conversation until it became a formal issue | Cannot name someone they helped improve |
Accuracy with financial data (accounts payable specialist)
| Score | What the candidate describes | What the evidence usually sounds like |
|---|---|---|
| 4 | Uses a consistent checking method, catches an error before it causes loss, traces the cause, and changes a control so it cannot recur | "Two invoices had the same number from different entities. I added a vendor-plus-amount match to our weekly check." |
| 3 | Has a checking routine and a specific example of catching and correcting an error | A named reconciliation or review step and one real catch |
| 2 | Describes being careful in general terms; errors were caught by reviewers or auditors | "I always double-check my work." |
| 1 | No routine, or describes errors reaching payment or reporting without a change afterwards | Cannot describe how they check a batch |
Rules for writing anchors interviewers can use
| Weak anchor | Problem | Better anchor |
|---|---|---|
| "Very good communication" | An adjective; every interviewer supplies their own meaning | "Explained a technical decision in the listener's terms and checked they understood" |
| "Always meets deadlines" | Frequency cannot be observed in a one-hour interview | "Describes raising a slipping deadline before it was missed, with a new plan" |
| "Confident and proactive" | Personality traits, not behavior; rewards interview polish | "Started work on a problem outside their role and told the owner" |
| "Better than most candidates" | Compares candidates instead of evidence to the job | Describe the behavior; never reference other candidates |
| "Strong leadership and strategic thinking" | Two competencies in one anchor | One competency per anchor; split them |
| Level 3 and level 4 differ only by "significantly" | Interviewers cannot tell where the line is | Add a concrete difference: scope, independence, prevention of recurrence |
A quick test for any anchor: read it to someone who has never met the candidate and ask them to describe an answer that would earn it. If their description matches yours, the anchor works.
Decision rules that belong with the scale
- Must-haves are gates, not averages. A 4 on a nice-to-have should never cancel a 1 on a must-have.
- No half points. Ask for a choice and a sentence on what was missing for the next level.
- "Not assessed" is excluded, not scored as a 2. Divide by what was covered, or wait for the owner of that competency.
- Independent first, then discuss. OPM's guide has panel members rate independently, then compare notes and explore the basis for discrepancies before reaching a consensus rating. Decide in advance whether your final number is a consensus or a median, and write it down.
- Evidence is required. A rating without an observation behind it is returned to the interviewer.
Calibrating the scale before and after launch
Before the first candidate
- Pilot the questions. OPM recommends a trial run of new questions with colleagues to check wording and whether they draw a range of answers.
- Rate the same answers independently. Use two or three written or recorded answers, recorded with the participant's agreement. Each interviewer rates alone, with a reason.
- Find the anchor behind each disagreement. Where ratings differ by two points, identify the wording that allowed both readings and rewrite it.
After the first batch of interviews
Look at each interviewer's distribution of scores. An invented example after eight candidates, four competencies each (32 ratings per interviewer):
| Interviewer | 1s | 2s | 3s | 4s | Mean | Reading |
|---|---|---|---|---|---|---|
| A | 2 | 9 | 15 | 6 | 2.78 | Uses the full scale |
| B | 0 | 2 | 10 | 20 | 3.56 | Possible leniency; check anchors for 3 vs 4 |
| C | 0 | 4 | 27 | 1 | 2.91 | Central tendency; ask for "not higher because" lines |
The mean for A is (2×1 + 9×2 + 15×3 + 6×4) ÷ 32 = 89 ÷ 32 = 2.78. Interviewers may simply have met different candidates, so compare them on candidates they both saw before drawing conclusions. OPM's list of rating errors includes leniency, strictness, halo and central tendency; a distribution table is the quickest way to spot the first and last. A full session plan is in interview calibration.
Rating scale mistakes and the fix
| Mistake | What happens | Fix |
|---|---|---|
| Labels only, no anchors | One interviewer's 4 is another's 3 | Behavioral anchors for each level of each competency |
| Different ranges per stage or interviewer | Scores cannot be compared or combined | One range for the whole loop |
| Changing the scale mid-search | Early and late candidates are rated on different rules | Collect changes and apply them to the next search |
| Anchors written after meeting candidates | The top anchor describes the favorite | Write and agree anchors at intake |
| Averaging across competencies with gates ignored | A must-have failure disappears into a good total | Gates first, averages second |
| No calibration | The scale looks structured but is used privately | Independent ratings of shared answers before launch |
A scale is only as consistent as the evidence each rating is based on. How to record that evidence is in evidence-based interview feedback. Interview Signal's scorecards attach quotes from the transcript to each score, which makes calibration conversations about what the candidate said rather than what each interviewer remembers.
Questions people ask
Is a 1–5 or a 1–10 interview rating scale better?
For most teams, 1–4 or 1–5. Every point needs a written description that interviewers can tell apart, and very few teams can write ten distinct levels of evidence for a competency. A 1–10 scale usually ends up used as a 1–5 scale with extra decimals of false precision.
Should interviewers be allowed to give half points?
No. A 3.5 means the interviewer could not decide which anchor the evidence matches, and that is useful information you lose. Ask them to choose, and write what would have moved the score up or down.
Do we need different anchors for every question?
One general scale for every competency, plus short question-specific anchors for the questions you ask most often. The general scale keeps the meaning of each number constant; the specific anchors make it easy to apply in the moment.
How do we know if interviewers are using the scale the same way?
Have them rate the same recorded or written answer independently and compare, before the search and again every few months. Between sessions, look at each interviewer's spread of scores. Someone who never uses the top or bottom of the scale is reading it differently.