Phase Three of the Scientific Hiring Method: Discovery
An interview rubric defines what the evidence must show before the candidate speaks.
Asking every candidate the same questions does not make an interview structured when the answers are still judged by instinct.
One interviewer may reward confidence. Another may value detail. A third may give the highest score to anyone who reminds them of a successful employee. Without a shared standard, the number written beside the answer only gives personal preference the appearance of precision.
What is an interview rubric?
An interview rubric is a set of rating anchors that describes what different levels of evidence look like for a specific competency or question.
It does not merely label the scale from poor to excellent. It explains the behavior, judgment, ownership, and results that justify each rating.
OPM identifies common rating scales as a key part of structured interviewing and advises that customized scales use behavioral or situational examples for the proficiency levels. See OPM’s customized rating-scale guidance.
Start with one competency
Suppose an Operations Manager needs strong process diagnosis.
The interview question is:
“Tell me about a time you inherited a process that produced unreliable results. How did you determine what was actually causing the problem?”
The rubric should describe how evidence of process diagnosis becomes stronger or weaker. It should not score unrelated qualities such as how smoothly the candidate tells the story.
A five-point interview rubric
| Rating | Evidence anchor for process diagnosis |
|---|---|
| 1 — Weak | Provides no relevant example or accepts the first explanation without gathering evidence. Cannot distinguish personal contribution from the team’s work and gives no credible result. |
| 2 — Limited | Describes a relevant problem but relies mostly on another person’s diagnosis, uses a narrow information source, or cannot explain why the chosen action addressed the real cause. |
| 3 — Acceptable | Gathers relevant information, separates symptoms from at least one likely cause, explains personal responsibility, chooses a reasonable action, and provides a credible result. |
| 4 — Strong | Tests competing explanations, uses several relevant sources, changes direction when evidence requires it, involves the right stakeholders, and connects the action to a measured result. |
| 5 — Exceptional | Shows unusually strong judgment across a complex or high-risk situation, anticipates second-order effects, creates a repeatable diagnostic method, and demonstrates the pattern across more than one comparable example. |
The anchors make the rating discussable. An interviewer who gives a 4 should be able to identify evidence of competing explanations, multiple sources, adaptation, stakeholder involvement, and measurement.
Do not write anchors that simply repeat the number
Weak rubric:
1 — Poor
3 — Average
5 — Excellent
These labels do not tell interviewers what to look for. They guarantee that every interviewer will supply a private definition.
Better rubric: Describe observable evidence, missing evidence, and the level of complexity or independence shown.
Use three levels when five creates false precision
A five-point scale is not automatically better. Some small teams can distinguish weak, acceptable, and strong evidence more consistently than five narrow levels.
OPM’s Structured Interview Guide notes that rating scales commonly use several proficiency levels and that the scale should be appropriate for the assessment. The important part is meaningful differentiation and clear benchmarks, not the appearance of mathematical sophistication. See OPM’s Structured Interview Guide.
Use the smallest scale that produces useful distinctions.
Build anchors from the work
Begin with the capability definition and ask:
- What would weak performance look like in this job?
- What is the minimum acceptable evidence from someone entering the role?
- What would strong performance look like under realistic conditions?
- Which details distinguish strong evidence from a polished but shallow answer?
- What should be tested through another assessment stage?
Use examples from actual work, past strong and weak performance, critical incidents, and knowledgeable people who understand the role.
Score one competency at a time
A candidate may provide strong evidence of process diagnosis and weak evidence of influence in the same answer.
Do not collapse several competencies into one overall score unless the question and rubric were deliberately designed that way. Separate ratings make it easier to understand what the candidate proved and where risk remains.
A worked disagreement
Two interviewers hear the same answer.
Interviewer A gives a 5: “The candidate was extremely confident and explained the result clearly.”
Interviewer B gives a 3: “The candidate found the cause and produced a result, but the manager supplied the data and chose the solution.”
The rubric shows that confidence is not part of process diagnosis. It also shows that a 5 requires unusually strong judgment and repeated evidence, which the answer did not contain.
After reviewing the candidate’s actual responsibility, both interviewers may settle on a 3. The purpose of the rubric is not to force agreement. It is to make disagreement about evidence rather than personality.
Require evidence notes beside every rating
A number without a reason cannot be audited or improved.
Weak note: “Good answer. Strong operator.”
Useful note: “Reviewed four weeks of failure data, interviewed sales and service managers, discovered conflicting definitions of confirmed, tested one shared status, and reduced delayed starts from fifteen to five per month. Manager selected the pilot team.”
The note supports the rating and preserves the detail needed for later Candidate Evaluation.
Score before the group discussion
Each interviewer should record the evidence and provisional rating independently.
OPM’s structured-interview materials emphasize standardized scoring and common rating scales. Independent scoring before discussion helps prevent the first or most senior opinion from changing everyone else’s initial judgment. See OPM’s Structured Interviews guidance.
After ratings are recorded, compare the evidence. Interviewers may revise a score when another person points to something they missed, but the reason should be visible.
Calibrate before the first candidate
Give interviewers a sample answer and ask them to score it using the rubric.
Compare the ratings. Where do interpretations differ? Which anchor is vague? Is one interviewer scoring communication style instead of the intended competency? Does “acceptable” mean the same thing to everyone?
Revise the rubric before live interviews begin. Calibration is easier when no candidate has yet become someone’s favorite.
Calibrate again after the search
Review whether interviewers used the anchors consistently and whether the ratings connected to later evidence.
A candidate rated highly for practical judgment may perform differently in the Candidate Test Drive. References may confirm or challenge a leadership claim. Post-hire results may reveal that the rubric rewarded something that did not matter.
Use the evidence to improve the rubric for the next search.
Do not average away a serious weakness
An overall average can hide a low score on a capability that is essential to the role.
Decide before interviewing whether any competency has a minimum acceptable rating. A candidate with strong communication and planning may still be too risky when the role requires technical judgment they did not demonstrate.
Minimums should be tied to the job and applied consistently, not invented after seeing which candidate the team prefers.
Do not confuse the interview rubric with the job scorecard
The Job Scorecard defines what successful performance in the role looks like.
The interview rubric defines what different levels of evidence look like when a candidate answers a particular question or demonstrates a competency.
The Interview Scorecard is the record used to capture the evidence and ratings for a specific candidate.
These tools connect, but they do different work.
Keep the scale job-related and fair
The EEOC advises that selection procedures should be job-related and consistent with business necessity, particularly when they disproportionately exclude protected groups. Rating criteria should measure the capability the employer claims to be assessing rather than personality similarity, protected information, or irrelevant presentation style. See the EEOC’s Employment Tests and Selection Procedures.
Reasonable accommodations may be required so the assessment measures the intended capability rather than a barrier created by the interview format.
This page is educational and is not legal advice.
Review the rubric before using it
- Each rating scale measures one defined competency or outcome.
- Anchors describe observable evidence rather than adjectives.
- The acceptable level reflects what the person must demonstrate upon entry.
- Strong ratings require more than confidence or polished delivery.
- Interviewers understand which details belong in evidence notes.
- Minimum ratings are set before candidates are evaluated.
- Interviewers practice scoring at least one sample answer.
- Ratings are recorded before the group discussion.
- The rubric connects to the job scorecard and later candidate evaluation.
- The scale will be reviewed after the hiring cycle.
What comes next
Use the Interview Scorecard to record ratings and evidence for each candidate. The following resource will provide the Interview Scorecard Template.
Use Structured Interviews for the complete method and Structured Interview Questions to build the question set.
Return to Discover What a Candidate Can Really Do or the main Scientific Hiring Method.
Sources and further reading
- U.S. Office of Personnel Management: Structured Interviews
- U.S. Office of Personnel Management: Structured Interview Guide
- U.S. Office of Personnel Management: Developing Customized Rating Scales
- U.S. Office of Personnel Management: Generic Rating Scales
- U.S. Equal Employment Opportunity Commission: Employment Tests and Selection Procedures
The five-point process-diagnosis rubric, calibration example, evidence-note standard, and Scientific Hiring implementation guidance are original to Scientific Hiring.