Phase Four of the Scientific Hiring Method: Decision

Candidate evaluation is the point where separate pieces of evidence become one fair comparison against the job.

A candidate should not win because one interviewer was impressed, another interviewer forgot to submit notes, or the team spent more time with that person than anyone else.

A fair evaluation gives each finalist the same standard, identifies what each person actually proved, and makes missing or conflicting evidence visible before the decision.

Build an evidence inventory

Before comparing candidates, gather every completed assessment record. Do not rely on memory.

Evidence source What it can show What to record
Résumé and application Career pattern, scope, progression, and claims that need testing Relevant evidence, assumptions, and questions raised
Phone screen Practical requirements, relevant examples, motivation, and basic alignment Advance, hold, or reject reason and unresolved questions
Structured interview Past behavior, proposed judgment, ownership, decisions, and results Criterion ratings, evidence notes, and confidence
Candidate Test Drive How the candidate handles a realistic piece of the work Observed decisions, output, score, and conditions
Reference checking Verification, context, patterns, support needs, and conflicting evidence Claims confirmed, context added, and risks identified

OPM’s assessment-strategy guidance explains that using multiple assessment methods can reduce error because people may perform differently across methods. The goal is not to collect more information for its own sake. It is to gather different forms of job-related evidence that complement one another. See OPM’s Designing an Assessment Strategy.

Return to the job scorecard

Create one evaluation row for every important outcome and capability on the Job Scorecard.

Do not begin with the candidates’ strengths. Begin with what the role requires.

For an Operations Manager, the evaluation might include:

  1. Establish a reliable weekly operating rhythm within 90 days
  2. Diagnose broken processes using evidence
  3. Gain cooperation across teams without relying on authority
  4. Choose priorities when several problems compete for attention
  5. Own mistakes and change behavior after feedback

Score the evidence, not the personality

Each rating should answer: how strongly does the evidence show that this candidate can meet this requirement?

The rating should not answer whether the interviewer enjoyed the conversation, admired the background, or could imagine having lunch with the person.

Use the behavioral anchors from the Interview Scorecard and the scoring standard created for the Candidate Test Drive. Keep communication style separate unless a defined form of communication is genuinely required by the work.

Add confidence beside every rating

A score of 4 based on one vague story is not the same as a 4 supported by a structured interview, realistic work sample, and confirming reference.

Use a simple confidence label:

High confidence: Several relevant sources point to the same conclusion.

Moderate confidence: The evidence is relevant but limited to one or two sources.

Low confidence: The conclusion depends on an assumption, a weak example, or unresolved conflict.

Do not multiply the score by the confidence label. Keep them separate so the team can see both the apparent capability and the strength of the evidence supporting it.

A worked comparison

Two finalists are being evaluated for the Operations Manager role.

Candidate A produced strong interview and Candidate Test Drive evidence of creating operating discipline. References confirmed the result. The weaker area is influence: one important rollout depended on executive intervention.

Candidate B showed excellent process diagnosis and stronger evidence of influencing peers. The weaker area is building a repeatable operating rhythm. The interview examples were smaller, and the Candidate Test Drive left several ownership questions unresolved.

Role requirement Candidate A Candidate B
90-day operating rhythm 4 — High confidence. Interview, test drive, and reference align. 3 — Moderate confidence. Relevant examples but smaller scope.
Process diagnosis 4 — High confidence. Several evidence sources align. 5 — High confidence. Strongest repeated pattern.
Influence without authority 3 — Moderate confidence. One rollout required executive support. 4 — High confidence. References confirmed repeated success.
Prioritization 4 — Moderate confidence. Good examples; limited reference detail. 3 — Low confidence. Test drive choices were not fully explained.
Ownership and learning 4 — High confidence. Clear mistake, consequence, and later change. 3 — Moderate confidence. Credible but less specific evidence.

This comparison does not automatically choose Candidate A. It shows why Candidate A may pose less risk when the first 90-day outcome is the most important requirement. It also identifies the influence risk that must be supported and monitored.

Do not rank candidates against one another too early

Evaluate Candidate A against the job. Then evaluate Candidate B against the job.

Only after both records are complete should the team compare them.

When the comparison begins too early, the standard starts moving. Candidate B may look strong merely because Candidate A struggled, even when neither person meets the role minimum.

Resolve gaps before the final meeting

Every open question should lead to an action.

Missing evidence: Add a focused interview question, short work sample, or reference question.

Conflicting evidence: Identify the exact conflict and ask the candidate or another qualified source for clarification.

Negative evidence: Decide whether the weakness violates a critical minimum or creates a manageable risk.

Irrelevant information: Remove it from the decision.

Do not keep interviewing simply because the team is uncomfortable making a choice. Additional assessment should answer a defined question.

Use minimums for critical requirements

A weighted average can hide a failure that the role cannot tolerate.

For each critical requirement, decide whether a minimum score is necessary. A candidate who falls below the minimum does not become qualified because several lower-priority strengths raise the overall average.

Minimums must be job-related, established before the final comparison, and applied consistently.

Record the final evaluation

For each finalist, record:

  • The rating for every role requirement
  • The evidence supporting each rating
  • The confidence in the conclusion
  • Any critical minimum not met
  • Missing, negative, or conflicting evidence
  • The candidate’s most important remaining risk
  • Whether that risk can be managed during the first 90 days

Then carry the completed evaluation into the Hiring Decision Matrix.

Keep candidate evaluation job-related

The EEOC explains that interviews, performance tests, application forms, experience requirements, and other procedures used to make employment decisions are selection procedures. When a procedure disproportionately excludes a protected group, it must be job-related and consistent with business necessity. See the EEOC’s Uniform Guidelines Questions and Answers and Employment Tests and Selection Procedures.

Evaluate the same job requirements for every candidate. Avoid adding personal characteristics, similarity, likability, family circumstances, protected information, or undefined “fit” to the final score.

This page is educational and is not legal advice.

Related resources

Use How to Make a Better Hiring Decision for the full Phase Four process and the Hiring Decision Matrix for the final comparison.

The evidence feeding this page comes from the Interview Scorecard, Candidate Test Drive, and Reference Checking.

Return to the main Scientific Hiring Method.