How to build an interview question bank your whole team can use
On this page
- What a question bank is for
- Structure: competencies first, role families second
- The fields every question needs
- Writing and approving new questions
- Versioning without breaking comparisons
- Retiring leaked and worn-out questions
- Calibration data: which questions work
- Where to keep it and who owns it
- Questions people ask
To build an interview question bank, organize approved questions by competency first and role family second, give every question the same fields (lead question, probes, a 1–4 anchored rubric, levels, status and version), approve new questions through a short review, and track how each question performs in real interviews so you can fix or retire it. The result is a shared library that turns building an interview guide into choosing, and makes ratings comparable across interviewers and searches.
This page is for the recruiter or people-operations lead who owns that library. If you need questions to start it with, behavioral interview questions with a scoring rubric has 24 with anchors.
What a question bank is for
Most teams already have a question bank: a folder of old interview guides and each hiring manager's favorite questions. That is a collection, not a bank. A bank does three jobs a folder cannot:
- Consistency. The U.S. Office of Personnel Management's structured interview guide describes structured interviews as asking all candidates the same questions in the same order and rating them on a common scale. A bank is how that holds across a team, not just within one search.
- Speed. An interview guide for a new search takes minutes when the competencies, questions and anchors already exist.
- Evidence over time. Every use of a question produces ratings. Collected together, they show which questions separate candidates and which just use up time.
Interviews also count as selection procedures under the federal Uniform Guidelines on Employee Selection Procedures, whose definition includes even informal interviews. A bank with documented, job-related questions is much easier to explain than a set of questions nobody can trace.
Structure: competencies first, role families second
Organize by competency, then tag each question with the role families and levels it suits. If you organize by role first, "handling a disagreement" gets written six times in six slightly different ways, with six sets of anchors.
| Layer | What it holds | Example |
|---|---|---|
| Core competencies | Assessed for almost every role | Ownership, collaboration, communication, learning |
| Role-family competencies | Shared by a group of similar roles | Customer-facing: handling escalations. Engineering: debugging. Finance: accuracy with financial data. |
| Level modifiers | How the anchors change with seniority | Individual contributor, first-line manager, manager of managers |
| Role-specific additions | A few questions only one role needs | A licensed trade's safety procedure, a regulated role's compliance scenario |
Give each question a stable ID made of a competency code and a number, such as COL-03 for
the third collaboration question. The ID never changes; the version does. Scorecards record the ID and
version, so a rating can always be traced to the exact wording and anchors used.
The fields every question needs
A question without these fields cannot be rated consistently. Copy this record format into a spreadsheet, a database or your ATS's question library.
ID: [competency code]-[number] e.g. COL-03
Version: [1.0] Status: [draft / pilot / approved / retired]
Competency: [name] - [one-line definition]
Role families: [e.g. customer-facing, operations]
Levels: [IC / first-line manager / manager of managers]
Type: [behavioral (past) / situational (hypothetical)]
Lead question: [exact wording, asked as written]
Probes: [3-5 allowed follow-ups]
Anchors (1-4):
4 [what a strong answer describes]
3 [...]
2 [...]
1 [...]
Does not assess: [what this question is not evidence for]
Typical time: [minutes with probing]
Equivalent to: [IDs of interchangeable questions]
Owner: [name] Approved by: [names, date]
Last reviewed: [date] Next review: [date]
Change log: [version, date, what changed, why]
An invented record, filled in:
COL-03, version 1.2, approved. Collaboration: works productively with people who think and work differently, and handles disagreement directly. Role families: all. Levels: IC and first-line manager. Type: behavioral.
Lead question: "Tell me about a disagreement with a colleague about how to do the work."
Probes: What was their view? What did you do first? How did it end? How is the working relationship now?
Anchors: 4, states the other view fairly, resolved it directly on the merits, relationship intact. 3, discussed directly, workable outcome. 2, escalated before talking to them, or describes them only as wrong. 1, went around them, or still unresolved and blamed on them.
Does not assess: conflict with a direct report (see LEAD-02). Typical time: 7 minutes. Equivalent to: COL-01, COL-05.
Change log: 1.1, probe "how is the relationship now" added after raters disagreed on 4 vs 3. 1.2, anchor 2 reworded to include escalation.
Writing and approving new questions
OPM's guide describes writing questions with a small group of people who know the job well, and it recommends writing more questions than you need so some can be discarded after review and a trial run. A lighter version works for most teams:
- Propose. Anyone can propose a question with a competency, the lead wording and draft anchors. No anchors, no proposal.
- Review against a checklist. The owner and one reviewer from the role family check it (below).
- Pilot. OPM recommends trying new questions on colleagues first to check wording and whether answers show a range. Mark the question "pilot" and use it in real interviews for a few searches, rated alongside an approved question for the same competency.
- Approve or revise. Look at the pilot ratings (see the calibration section). Approve, rewrite the anchors, or drop it.
- Record the decision. OPM's guide recommends documenting how questions and rating scales were developed, including the job analysis and the pilot. A dated change log is the practical version.
Review checklist for a proposed question
- Linked to one competency that comes from the job, not a general trait.
- Open-ended, clear and free of jargon, which OPM lists among its requirements for interview questions.
- Behavioral questions ask for a single past instance; OPM suggests words like "most", "last" or "worst" to focus the candidate on a specific incident.
- Answerable by candidates from different backgrounds, not only those who worked at a similar company.
- Does not invite answers about health, family, religion, age, national origin or other protected characteristics. The EEOC says pre-employment questions should be limited to what is needed to judge qualification.
- Has four anchors that describe behavior, not adjectives. Anchor-writing rules are in interview rating scale.
Versioning without breaking comparisons
| Change | Treat as | Why |
|---|---|---|
| Typo, punctuation | Same version | Nothing a candidate or rater would notice |
| Added or reworded probe; clarified an anchor | New minor version (1.1 to 1.2) | Ratings stay broadly comparable; the change log explains any shift |
| Changed the lead question's meaning, or rewrote anchors | New question ID | Old and new ratings measure different things and should not be pooled |
| Moved to a different competency | New question ID; retire the old | Scorecards for past candidates must still make sense |
Two rules prevent most versioning problems. First, never change a question in the middle of a search: candidates on the same slate should be rated on the same wording. Collect changes and apply them to the next search. Second, never delete a retired question. Keep it, with its last version, for as long as you keep the scorecards that used it; retention periods are covered in how long to keep interview notes.
Retiring leaked and worn-out questions
Candidates share questions. Glassdoor, for example, publishes interview questions and reviews posted by candidates for many employers. You cannot stop that, and for behavioral questions you do not need to panic: a candidate who knows the question still has to describe a real situation that survives three probes. Situational questions and work samples with a single right answer are more exposed.
Signals that a question needs rotating or retiring:
- Score drift. Ratings for the question rise over several months while other questions for the same competency stay flat.
- Rehearsed answers. Interviewers report very similar stories, structures or phrases from unrelated candidates.
- Found online. The exact wording appears in a public interview review.
- No longer discriminates. Almost everyone gets the same score.
- The job changed. The competency is no longer a must-have for the role family.
Rotation is simpler than retirement when each competency has two or three questions marked as equivalent. Swap in the next one at the start of a search, leave the anchors pattern the same, and watch whether scores settle back.
Calibration data: which questions work
Track a few numbers per question from submitted scorecards. An invented quarter of data:
| Question | Asked | Not assessed | 1s / 2s / 3s / 4s | Double-scored pairs within 1 point | Reading |
|---|---|---|---|---|---|
| PRI-01 | 40 | 2 (5%) | 3 / 10 / 17 / 8 | 12 of 12 | Healthy: uses the full scale, raters agree |
| PRI-02 | 36 | 9 (25%) | 0 / 3 / 22 / 2 | 9 of 11 | Often skipped and 22 of 27 ratings (81%) are 3s: too long, anchors do not separate |
| COL-04 | 30 | 1 (3%) | 0 / 1 / 6 / 22 | 10 of 10 | 22 of 29 ratings (76%) are 4s: too easy or leaked; check for rehearsed answers |
- Not-assessed rate. High means the question runs long or sits too late in guides. Shorten it or move it earlier.
- Score distribution. A question where nearly every rating is the same number is not telling you anything about the candidates. Rewrite the anchors around the level that is overused.
- Rater agreement. When two interviewers score the same interview, from a joint panel or a recording made with consent, count how often they land within one point. Low agreement points to an ambiguous anchor.
- Time. Ask interviewers to note roughly how long a question took with probing, so guides can be timed honestly.
Do not try to monitor fairness question by question on your own. If your organization analyzes outcomes by demographic group, that belongs with HR and counsel, using voluntarily self-identified data. For reference, the Uniform Guidelines describe a four-fifths rule of thumb: a selection rate for a group below 80 percent of the highest group's rate is generally regarded by federal agencies as evidence of adverse impact. How to run the interviewer side of calibration is in interview calibration.
Where to keep it and who owns it
- Start in a spreadsheet. One row per question version, with the fields above as columns. Move to your ATS's question library once the structure has settled.
- One owner, named reviewers. The owner controls status and versions. A reviewer from each role family approves questions for it.
- Limit access to interviewers. Anchors shared outside the hiring team become answer keys.
- Review quarterly. Pull the calibration numbers, rotate anything drifting, clear the proposal queue.
- Connect it to training. New interviewers learn the bank's scale and probes before their first interview, not during it.
A bank also makes tools more useful. Interview Signal builds a live question guide from the job description and ticks off must-ask questions as they are covered, which works best when your team has already agreed which questions are must-asks for each competency.
Questions people ask
How many questions should an interview question bank have?
Enough that every competency you assess has at least two or three interchangeable questions per level, so you can rotate them and replace one that leaks. For a team hiring into a handful of role families, that is often a few dozen approved questions, not hundreds. A small bank that people trust gets used; a large one gets skimmed.
Who should own the question bank?
One named owner, usually in talent acquisition or people operations, with a reviewer from each role family who knows the work. The owner controls status and versions; hiring managers propose questions and changes but do not edit approved questions directly.
Should interviewers be allowed to ask questions that are not in the bank?
Follow-up probes, yes, always. New lead questions, only through the proposal process, so they get anchors and a review before they count toward a score. An unapproved question can still be asked in the unscored part of an interview.
Is it a problem if candidates post our questions online?
Some leakage is unavoidable, and for behavioral questions it matters less than it seems, because a prepared candidate still has to describe a real situation that holds up under probing. Watch the data: if a question's scores drift upward and answers start to sound alike, rotate in an equivalent question.