How to compare candidates after interviews with a decision matrix
On this page
- Check that your evidence is comparable first
- The decision matrix template
- A worked example: three finalists and a tie
- Weighting must-haves without fooling yourself
- Biases that show up at comparison time
- When to add a work sample or another conversation
- When nobody clears the bar
- Write down the decision
- Questions people ask
To compare candidates after interviews, compare evidence rather than impressions. Put each finalist's scores for the same competencies side by side, set aside anyone who missed a must-have, weight what is left, and read the evidence behind any close call before you decide. The totals do not make the decision. They show you where to look.
Below is a decision matrix template you can build in any spreadsheet, a worked example with three invented finalists and a tie, the biases that show up at comparison time, and how to write down the decision so it still makes sense in a year.
Check that your evidence is comparable first
A matrix built on uneven evidence produces a confident wrong answer. Before you compare anyone, check four things:
- Same competencies for everyone. Every finalist was assessed on the same list, agreed at intake. If one candidate had a "culture" conversation and another did not, that session cannot count.
- Scorecards written independently, before any discussion. Scores written after the debrief tend to repeat what the loudest person said. If you have not held the debrief yet, run it with the interview debrief template before comparing.
- No guessed scores. A competency nobody asked about should be blank, not a polite 3. List the blanks; they are what you may need to fill before deciding.
- Written notes, not memory. If the interviews were two weeks apart, reread the scorecards. The candidate you met on Friday will otherwise feel stronger simply because you remember them.
The scores themselves should come from an anchored scale. The interview scorecard template uses 1 to 4 with a description for each point, and this page uses the same scale.
The decision matrix template
One row per competency, one column per candidate. Weights are set at intake, before anyone is interviewed, and must add up to 100%. Copy the layout into Excel or Google Sheets:
DECISION MATRIX: [role] Date: [date] Decision owner: [name]
Scale: 4 strong / 3 meets the bar / 2 below / 1 no evidence / blank = not assessed
A B C D E F
1 Competency Must-have? Weight [Cand. 1] [Cand. 2] [Cand. 3]
2 [C1] Yes 30% [score] [score] [score]
3 [C2] Yes 25%
4 [C3] No 20%
5 [C4] No 15%
6 [C5] No 10%
7 Weights total 100%
8 Must-have check [formula] ...
9 Weighted score (1-4) [formula] ...
10 Open questions [text] ...
11 Evidence to reread [text] ...
Row 8, must-have check (put in D8, copy right):
=IF(MIN(D2:D3)=1,"OUT",IF(MIN(D2:D3)=2,"NEEDS A PLAN","CLEARS"))
Row 9, weighted score over the weight actually assessed (D9, copy right):
=SUMPRODUCT($C$2:$C$6,D2:D6)/SUMPRODUCT($C$2:$C$6,--(D2:D6<>""))
DECISION
Selected: [name] Reason in two sentences: [...]
Second choice, if any: [name]
Not selected: [name] - [job-related reason per candidate]
The second formula matters. A plain weighted sum treats a blank as zero, so a candidate whose last interviewer ran out of time looks worse than one who was fully assessed. Dividing by the weight that was actually scored avoids that; the blank still shows up in "open questions".
A worked example: three finalists and a tie
Invented candidates for an invented role: Accounts Payable Supervisor, leading a team of six at a distribution company. The intake agreed two must-haves (team leadership, process control) and three nice-to-haves.
| Competency | Must-have | Weight | Morgan Adeyemi | Riley Novak | Casey Brennan |
|---|---|---|---|---|---|
| Team leadership | Yes | 30% | 2 | 3 | 4 |
| Process control and accuracy | Yes | 25% | 4 | 3 | 3 |
| Systems and automation | No | 20% | 4 | 3 | 2 |
| Vendor communication | No | 15% | 4 | 3 | 3 |
| Reporting to finance leadership | No | 10% | 3 | 4 | 3 |
| Must-have check | Needs a plan | Clears | Clears | ||
| Weighted score | 100% | 3.30 | 3.10 | 3.10 |
The arithmetic, so you can check it:
Morgan 0.30x2 + 0.25x4 + 0.20x4 + 0.15x4 + 0.10x3 = 0.60+1.00+0.80+0.60+0.30 = 3.30
Riley 0.30x3 + 0.25x3 + 0.20x3 + 0.15x3 + 0.10x4 = 0.90+0.75+0.60+0.45+0.40 = 3.10
Casey 0.30x4 + 0.25x3 + 0.20x2 + 0.15x3 + 0.10x3 = 1.20+0.75+0.40+0.45+0.30 = 3.10
Step 1: gates before totals
Morgan has the highest total and a 2 on team leadership, a must-have. The evidence on the scorecard: has supervised two contractors for six months, and could not give an example of dealing with a performance problem after two follow-ups. Morgan is not out, but is not a yes without more evidence on leadership. Sorting by the total would have put Morgan first and hidden exactly the gap the intake said mattered most.
Step 2: break the tie by reading the rows where they differ
Riley and Casey tie at 3.10 with different profiles. They differ on three rows: leadership (Casey 4, Riley 3), systems (Riley 3, Casey 2) and reporting (Riley 4, Casey 3). The team goes back to the intake notes: two new starters, an open performance issue on the team, and an ERP replacement next year that the software vendor will implement. Leadership is the problem this year; the systems gap is one that training and the vendor's rollout can close.
Step 3: test how fragile the answer is
Swap two weights and recalculate: systems at 30% and leadership at 20%. Riley becomes 3.10 and Casey 2.90. A decision that flips when two weights trade ten points is a judgment call, and it should rest on the reasoning in step 2, written down, not on the total.
The decision
Selected: Casey Brennan. Clears both must-haves, strongest leadership evidence (described coaching a clerk through a formal improvement plan to a successful outcome), and leadership is the team's most pressing need this year.
Plan for the gap: systems and automation scored 2; include ERP training in onboarding and the vendor implementation plan.
Second choice: Riley Novak. Recruiter to keep Riley informed until Casey accepts.
Not selected: Morgan Adeyemi. Below the bar on team leadership, a must-have, with no first-hand example of managing performance.
Weighting must-haves without fooling yourself
- Gates first, weights second. A must-have is not simply a heavy weight. A candidate can score high enough elsewhere to outweigh a 1 on a must-have, which is exactly the result the word "must" rules out.
- Set weights at intake, in writing, before anyone is interviewed. Weights adjusted after the interviews tend to drift toward the candidate the team already prefers.
- Keep them coarse: multiples of 5 or 10 percent. Interview ratings are not precise enough to justify more.
- Do not average away a disagreement. If one interviewer gave a 4 and another a 1 on the same competency, the average of 2.5 describes nobody's view. Resolve it in the debrief, from the evidence.
- If you change a must-have, change it for everyone. Reconsider anyone already turned down on it, and note why it changed.
Biases that show up at comparison time
Some rating errors happen in the interview. Others happen days later, in the room where candidates are compared. The US Office of Personnel Management's structured interview guide (Appendix F) names several of them; the checks are ours.
| Bias | What it sounds like | Check |
|---|---|---|
| Contrast effect (OPM) | "After Tuesday's candidate, Riley was a breath of fresh air." | Each score is against the anchors, not the previous candidate. Do not rescore after seeing the others. |
| Halo effect (OPM) | "Morgan's systems answer was so good, I'm sure the leadership side will come." | Compare one row at a time across candidates, not one candidate at a time down the rows. |
| Similar to me (OPM) | "Casey reminds me of myself at that stage." | Ask "what did they say that shows it?" and point to the scorecard. |
| Negative emphasis (OPM) | One awkward answer outweighs five solid ones. | Read the full row of scores before discussing any single moment. |
| Pressure to hire (OPM) | "We've been searching for three months, let's just pick." | Keep "none of them" as an explicit option in the matrix. |
| Recency | The last interview feels most vivid, so it feels strongest. | Work from written scorecards; review candidates in alphabetical order, not interview order. |
| Moving the goalposts | "Leadership matters less than we thought" right after meeting a favorite who is weak on it. | Weights are locked at intake. A change needs a written reason and applies to every candidate. |
The habit that counters most of these is the one in the scorecards themselves: every rating tied to something the candidate said or did. Evidence-based interview feedback covers how interviewers write that.
When to add a work sample or another conversation
Add a step when it will produce evidence you do not have. Do not add one to postpone a decision.
Add one when
- The decision hinges on a competency that is blank or has split scores.
- Finalists tie on the total with different gaps, and the gap that matters most has thin evidence.
- A must-have scored 2 and you want to know whether it is really a 2.
- A claim central to the role could not be checked in conversation, such as a technical skill or a writing standard.
Do not add one when
- You already have the evidence and are hoping for a feeling.
- You would give it to only one candidate. Extra steps go to every finalist still in contention, so the evidence stays comparable.
- The exercise tests something that is not on the scorecard.
How to run it
- Target the gap. For Morgan in the example, a 30-minute conversation built around two leadership scenarios, scored on the same leadership anchors, offered to any other finalist still being considered.
- Write the rubric before anyone takes it, using the same 1–4 scale.
- Keep it short and respect the candidate's time. Anything that looks like real work for your company should be small or paid.
- References can fill a gap too, if you ask about the specific competency. The reference check template has questions.
When nobody clears the bar
A matrix where every finalist has a 1 or an unresolved 2 on a must-have is a result, not a failure of the process. Your options, in the order worth considering:
- Reopen the search with what you learned: which must-have was hardest to find, and where the strongest near-misses came from.
- Revisit the must-haves with your recruiter, openly. If the requirement was unrealistic for the pay or location, fix it for the next round and reconsider the near-misses against the new version.
- Reshape the role: split it, change the level, or move one responsibility elsewhere.
Hiring the best of a slate that does not meet the bar, because it is the slate you have, is the pressure-to-hire error in its purest form.
Write down the decision
Save the matrix with a short note in the candidate folders. It should let someone who was not in the room understand the choice. Keep every reason job-related and tied to the competencies; what not to put in writing is covered in what not to write in interview notes.
- Role, date, who took part in the decision.
- Candidates considered, with the matrix attached.
- Who was selected and why, in two or three sentences that point to scores and evidence.
- For each candidate not selected, the job-related reason.
- Anything that changed since intake (a weight, a must-have) and why.
- Known gaps in the selected candidate and the onboarding plan for them.
Keep it. The EEOC's recordkeeping page explains that covered employers must keep personnel and employment records for one year, and that when a charge is filed, relevant records must be kept until the matter is finally resolved. The underlying rule is 29 CFR 1602.14, which runs the year from the date of the record or the personnel action, whichever is later. That is the federal rule as of September 2026; state law and your company's policy may require longer.
Questions people ask
Should I rank candidates or score them?
Score first, rank second. Score each candidate against the written anchors for each competency, without looking at the others, then put the scores side by side. Ranking straight away turns the decision into who seemed better than whom, which is where contrast effects creep in.
What if two finalists end up with the same weighted score?
A tie on the total usually hides different profiles. Compare them competency by competency, go back to the intake to decide which difference matters most for the first year, and if the evidence on that competency is thin, get more evidence on it from both candidates.
Can I hire someone who missed a must-have?
Only if you conclude the must-have was wrong, and then it changes for everyone: reconsider any candidate you turned down for the same reason, and write down why the requirement changed.
How many candidates should be in the final comparison?
Two to four is workable. With one, you are deciding whether to hire rather than whom; with more than four, the interviews are usually spread over weeks and memory starts doing the comparing.
How long should we keep the decision matrix and scorecards?
EEOC regulations require covered employers to keep hiring records for one year from the record or the decision, whichever is later, and longer if a charge is filed. State law or your own policy may require more; check both.