Have you ever reviewed a stack of applications, found several plausible candidates, and realized your team has no consistent way to decide who should progress? One reviewer prioritizes experience, another is persuaded by a polished answer, and a third is looking for evidence nobody else assessed.
That lack of alignment is widespread. In a Leadership IQ survey of 2,770 US leaders and employees, 68% of participating HR executives reported inconsistency in how their hiring managers evaluated candidates.
I’ve observed that the problem is rarely that reviewers are careless. Instead, broad requirements such as “strong communication” or “commercial mindset” do not specify what evidence counts, how much is sufficient, or how incomplete answers should be treated.
And when you’re dealing with high application volumes and AI-polished responses, each reviewer must fill those gaps with an unstated standard. The shortlist then reflects several interpretations of the role instead of one shared definition of what good looks like.
This guide shows you how to turn job requirements into observable criteria, define the evidence each one requires, and score every candidate against the same standard.
You’ll learn:
- How to turn broad job requirements into observable criteria and evidence rules.
- How to set clear scoring standards and resolve borderline decisions.
- Where AI can support the review while your hiring team retains responsibility.
- How to maintain the same standard across reviewers and high-volume batches.
- How to use the accompanying workbook to build the process for your next role.
How to shortlist candidates in seven steps
- Turn each job requirement into an observable criterion
- Define what evidence would prove each criterion
- Choose a method that can collect the required evidence
- Write questions or tasks that produce comparable responses
- Build a candidate shortlisting scorecard before reviews begin
- Review every candidate against the same scoring standards
- Resolve borderline cases and record who should progress
Start with the work the person must perform. Job descriptions often mix genuine requirements with broad traits, inherited preferences and credentials that are easy to list but hard to connect to performance.
Talent leaders like Justin Krulicki, Global Talent Acquisition Manager at Miovision (formerly MasterClass), have also observed that “bias starts to creep in” when multi-stage interview processes assess the same loosely defined requirements instead of clear benchmarks.
To avoid this, take each requirement and ask what a reviewer would need to observe before they could say the candidate meets it. “Commercial mindset” is too open to interpretation. “Uses customer and pipeline evidence to set account priorities” gives reviewers something they can look for.
Meet with relevant stakeholders (hiring managers, supervisors, team leads) to build a comprehensive list of the necessary skills and traits. Then classify each criterion according to how it affects the decision:
- Essential (mandatory): The minimum requirements a candidate must meet. If they do not meet an essential criterion, they do not progress.
- Weighted (important): Skills or traits that help distinguish candidates who clear the essentials. Strong evidence in one weighted area can offset weaker evidence in another, according to the priorities agreed in advance.
- Contextual (preferred): Useful experience or background that adds context but should not disqualify a candidate or receive the same weight as evidence tied directly to performance.
The table below turns four requirements for an example B2B sales manager role into observable criteria.
Three rules make these criteria easier to use:
- Make each category observable rather than a broad label. For every requirement, write down what a reviewer would need to see or hear before deciding that a candidate meets it. Those details will guide the review later, so define them now.
- Keep the list short enough to use. When every line in the job description becomes essential, the scorecard stops distinguishing between candidates and starts excluding people for requirements that may not matter.
- Leave protected characteristics such as age, race, sex and disability status out of the criteria. Review the framework against the employment and accessibility requirements that apply in each hiring jurisdiction.
For more on catching bias at the criteria-setting stage, including how to audit job descriptions for subtle exclusion signals, see this guide to inclusive hiring practices.
An observable criterion still leaves room for disagreement until you define what would count as proof. The right evidence depends on the criterion you are measuring because every source has strengths and limitations. “Can build a sales strategy” tells the reviewer what matters, but it does not yet tell them what a satisfactory answer must contain.
Write an evidence rule for every scored criterion. This is a short description of what the candidate's answer must show. It should identify:
- The information the candidate must provide.
- The minimum detail required for the evidence to count.
- Any result, example or demonstration the reviewer needs to see.
- What should happen when the evidence is missing.
Continue the sales manager example and the difference becomes clearer.
Treat missing evidence and weak evidence differently. Missing evidence means the process did not give the candidate a fair opportunity to demonstrate the criterion, or the reviewer cannot find the required information. Weak evidence means the candidate had the opportunity but their response did not meet the agreed rule.
That distinction matters later. If the process did not capture the required information, you may need clarification or a second review. Weak evidence is already a result.
With the criteria and evidence rules defined, choose how you will collect the evidence. The right method depends on what this stage must establish: eligibility, role-specific skill, or richer evidence about reasoning and judgment. Starting with a preferred format can produce questions that are convenient to administer but poorly matched to the criterion.
Use the simplest method that can produce credible evidence. A binary application question can confirm work authorization. It cannot show how someone handles an underperforming sales region. A structured work sample can demonstrate applied skill, but it creates more work for the candidate and the reviewing team.
You may use more than one method. A practical sequence confirms eligibility first, then tests skills or collects richer evidence from candidates who remain eligible. This narrows the pool by stage without asking every applicant to complete the most demanding activity.
If a candidate needs a different response format as an accommodation, keep the criterion and scoring rule the same. Consistency comes from applying the same standard, rather than forcing every person to respond in one format.
If you are comparing software for this stage, start with the types of evidence your process needs. Willo’s guide to candidate assessment tools explains how the main assessment formats differ.
Every question or task should collect evidence for a criterion already on the scorecard. If you cannot identify the criterion it supports, remove it.
Start with the evidence you decided you need, then turn it into a question or task the chosen method can collect. Each prompt should connect directly to one criterion and the evidence required. For the sales strategy criterion, a structured question could be:
Describe a sales strategy you developed for a team or territory. What evidence shaped your priorities, what did you change and what happened as a result?
The wording asks for the elements in the evidence rule. It gives every candidate the same opportunity to explain the context, evidence, action and result.
A work sample could test the same criterion differently. You might give candidates a small, standardized set of pipeline and customer information, then ask them to identify priorities and explain their reasoning. The work sample provides direct evidence of analysis, while the structured question provides evidence from past experience.
Check each prompt against four questions before using it:
- Which criterion does this assess?
- What evidence should the response contain?
- Can the candidate reasonably provide that evidence from the prompt?
- Can two reviewers apply the same scoring rule to the result?
The fourth question catches many weak prompts. “Tell us about yourself” may produce an interesting answer, but it gives reviewers little control over what they are comparing.
Clear prompts and evidence rules remove that guesswork. Candidates understand what they need to demonstrate, while reviewers receive answers that are easier to compare.
For more prompt ideas, see these candidate screening questions. Use them as starting points and map each one to your own criteria before adding it to the process.
A scorecard turns the criteria and evidence rules into a repeatable evaluation. Build it before candidate responses arrive, with input from the recruiters, hiring managers and other reviewers who will use it. That gives the team time to agree on the standard without being influenced by a particular applicant.
Generic labels such as “yes,” “maybe” and “strong yes” leave the most important judgment unstated. Instead, create scoring anchors that describe weak, sufficient and strong evidence for each criterion.
Consider two reviewers assessing the same sales-strategy answer: one may call it strong because the plan sounds ambitious, while another marks it weak because the candidate gives no measurable result. Agreeing in advance that sufficient evidence requires both a specific strategy and an outcome resolves that difference before scoring begins.
Mark essential criteria clearly. A candidate who does not meet an essential criterion should not compensate by collecting points from several preferred ones. Weighted criteria can help distinguish candidates who clear the essentials, but the weighting should be agreed before scoring starts.

Include a notes field beside each score. The reviewer should record the evidence that produced the rating, especially when it is missing, unclear or close to the agreed bar. Those notes give the team something concrete to discuss during a second review or when aligning reviewers on the scoring standard.
That combination of role clarity and one shared scorecard keeps different reviewers evaluating the same job, rather than applying their own version of what a good candidate looks like.
Score candidates against the criteria before comparing them with one another. Apply the scoring anchors from Step 5: weak evidence does not meet the agreed bar, sufficient evidence meets it and strong evidence clearly exceeds it. Record information that was not provided as missing or not evidenced instead of forcing it into a weak score. This prevents the strongest or weakest application in the current batch from quietly changing the bar.
Reviewers should work independently first. Each reviewer records the relevant evidence, applies the scoring anchor for that criterion and identifies anything missing. The panel can compare results after those initial judgments are complete.
The table below illustrates how three candidates might be assessed against the sales manager scorecard.
Candidate C performs strongly on several weighted criteria, but does not meet the essential requirement. Candidate B meets the essential criterion but has two areas that need another look, which makes a second review more appropriate than an immediate decision.
This is why the scorecard needs evidence notes. The panel can discuss the reason behind a score instead of debating labels in the abstract.
A borderline candidate usually meets every essential criterion but has incomplete, conflicting or closely matched evidence elsewhere. Candidate B in the example above meets the full-cycle sales requirement but has incomplete outcome evidence for sales strategy and no evidence for coaching. The team should apply the priorities it agreed earlier and request a second review, rather than treating every gap as equally important or inventing a new standard.
Use this sequence:
- Identify what is unclear. Is the evidence missing, open to interpretation or below the agreed bar?
- Request an independent second review. The second reviewer uses the same criteria and scoring anchors before seeing the first reviewer’s conclusion.
- Ask for more information only when the process left a real gap. Use the same clarification rule for any candidate in the same position.
- Compare the evidence, not just the scores. Reviewers explain which response detail and scoring anchor produced their judgment. Avoid averaging scores that came from different interpretations of the evidence.
- Record the final decision. Capture the outcome, the evidence considered, the reason and the next step.
Do not rewrite a criterion after seeing a candidate you particularly want to include or exclude. If the team discovers that a criterion is wrong, document the change and apply the revised rule consistently to every affected candidate.
Once the shortlist is complete, communicate the outcome to everyone who applied. Willo’s guide to candidate experience covers how to handle that next stage. Together, these seven steps form one connected system: define the criteria, specify the evidence, choose the method, write the prompts, anchor the scores, compare candidates and document the decision.
Where AI can support the candidate shortlisting process
AI is already part of many hiring workflows. Willo’s 2026 Hiring Trends Report found that 64.9% of hiring teams had increased their use of AI, while 78.7% said final hiring decisions must remain human-led.
The shortlisting framework gives AI a defined role because the criteria and evidence rules already exist. It can help a reviewer find and organize relevant information without setting the standard or deciding who should progress.
Use AI to surface evidence against predefined criteria
Give AI the criterion, the evidence rule and the candidate’s original response. Ask for a summary that points the reviewer back to the relevant passages.
Say a candidate’s sales-strategy response runs long and buries the most relevant detail halfway through. AI can surface where they explain the evidence they used, the action they took and the result they achieved. The reviewer can then return to those passages and apply the scorecard anchor.
Ask AI to “find the best candidates,” and you have not told it what good means for the role. Give it the criterion, evidence rule and source response instead. That keeps the output tied to the framework your team agreed before reviewing anyone.
Check AI-assisted summaries against the original response
Summaries compress information. That makes review faster, but it can also remove context, uncertainty or a qualification that matters to the score.
Spot-check the summary against the candidate’s original answer before relying on it. Increase the level of checking when the evidence is ambiguous, the criterion is essential or the decision is close.
An AI-generated summary should point to evidence. It should never become the only evidence in the record.
Keep decisions about who progresses human-led
Objective eligibility questions can follow rules set by the employer. Comparisons that require judgment still need a person to weigh the evidence, consider the context and take responsibility for the decision.
Keep people responsible for how criteria are weighted, how borderline cases and reviewer disagreements are resolved, and who ultimately progresses. AI can prepare the evidence and show where to look. Your hiring team makes the call.
That is the boundary this workflow preserves. AI can help reviewers find and organize evidence, but a person remains responsible for interpreting it and deciding who progresses.
For a closer look at tools that support this work, read Willo’s guide to AI candidate screening tools.
How many candidates should you shortlist?
There is no universal target. Your shortlist is the group that meets the agreed standard and can reasonably move into the next stage.
Decide how many candidates you can realistically take into the next stage before reviewing applications, but do not force the shortlist to match that number. If twelve candidates meet the essential criteria and you can only interview six, use the weighted criteria and stronger evidence rules you agreed in advance. Avoid changing the standard after seeing the results.
If too few candidates clear the essentials, lowering the standard is rarely the first answer. Check whether the requirement is genuinely essential, whether candidates had a fair opportunity to provide the evidence and whether the sourcing strategy reached the right pool. Any correction to the framework should be applied consistently to the full candidate set.
A niche senior role may produce two qualified candidates. A high-volume entry-level role may produce fifteen. Both outcomes can be valid when the same decision rules were applied consistently.
How to maintain shortlist quality when recruiting at volume
High-volume recruitment makes consistency harder. More candidates usually mean more reviewers, longer batches and more chances for the scoring standard to shift.
Miller’s team uses set application windows to make that standard workable. Other teams may use different workflows, but volume should change how the review is organized, not which arbitrary fraction of candidates receives proper consideration.
Calibrate reviewers before scoring begins
Calibration means checking that reviewers apply the scoring standard in the same way. Give them a small set of sample responses to score independently, then compare the results and discuss the evidence behind any disagreement before the full batch begins.
Use that discussion to agree how the team will treat partial answers, transferable experience and missing evidence. If scoring begins beforehand, a later change to the bar may mean rescoring everyone reviewed earlier. The exercise should not reveal a preferred candidate or encourage reviewers to match a senior person’s score.
Score independently before comparing notes
Have each reviewer record their initial evaluation before the panel discusses a candidate. If reviewers compare notes too early, the first confident opinion can anchor everyone else’s score. Independent scoring preserves separate judgments and makes genuine disagreement visible.
When the panel meets, compare the evidence and the scoring rule behind each rating. The discussion can then focus on why the scores differ instead of whose opinion sounds most confident.
Check for reviewer drift between batches
The bar can gradually loosen or tighten as reviewers move through a long candidate list. Review in manageable batches, then score a small shared sample at agreed intervals and compare the results with the examples used at the start.
If the scores have moved, correct the interpretation and revisit affected candidates. Do not leave the first and final batch operating under different standards.
Record the evidence and reason behind every decision
A shortlist becomes easier to review when every outcome contains four things:
- The criteria assessed.
- The evidence reviewed.
- The result against each scoring anchor.
- Why the candidate did or did not progress, and what happens next.
This record helps a hiring manager understand the shortlist without reconstructing the process from comments, memory and separate files.
Tunstall provides one example of what greater screening capacity can look like. The company screened more than 700 candidates with Willo in six months, increasing the number it could review effectively by nearly 75%, according to Willo’s published case study.
Before closing a high-volume review cycle, check that:
- Every candidate was assessed against the current criteria.
- Every scored criterion points to candidate evidence.
- Reviewers used the agreed anchors.
- Missing evidence was separated from weak evidence.
- Borderline cases followed the second-review rule.
- Criteria changes were applied to every affected candidate.
- Every decision records why the candidate did or did not progress, and what happens next.
How Willo helps you apply the shortlisting process at scale
A shortlisting process is only useful if your team can run it consistently for every role and candidate. That becomes difficult when criteria, responses, scores and reviewer notes sit in separate places. Your team spends more time organizing evidence, and the link between the candidate’s answer and the decision becomes harder to follow.
Willo is candidate screening software that helps hiring teams collect comparable evidence and review it against their own definition of what good looks like. The hiring team remains responsible for deciding who progresses.

Build the role around the criteria your team defines
Start by creating the role around the information and criteria your team already has. Willo uses those inputs to build a hiring Blueprint, which you can edit before assessing candidates. In Willo Insights, must-haves, differentiators and contextual information correspond to the essential, weighted and contextual categories used above.
- Define the Blueprint: Add the role requirements and criteria reviewers should use when assessing candidates.
- Review the criteria: Edit, remove or add criteria so the Blueprint reflects the team’s actual standard.
- Connect questions to the role: Build the assessment around the evidence candidates need an opportunity to provide.

For instance, a sales manager Blueprint could make full-cycle B2B ownership a must-have and sales strategy development a differentiator. The assessment then uses the same criteria and priorities as the scorecard developed earlier.
Now every question, score and evidence summary refers back to the same definition of what good looks like.
Collect comparable evidence in the format each criterion requires
Some roles need a candidate to explain their reasoning. Others need a file, a written response or an answer to an objective eligibility question. Forcing every criterion into one format can weaken the evidence.
- Use structured prompts: Give candidates the same question and response conditions for each criterion.
- Choose the response type: Collect video, audio, text, multiple-choice or file-upload responses according to the evidence required.
- Keep candidate records together: Responses, transcripts, scorecards and reviewer comments remain connected to the candidate record.
For instance, the team could use a file upload for a short sales planning exercise and a structured video or text response for the candidate’s reasoning. Reviewers receive comparable inputs without pretending the two criteria require the same type of proof.
You reach the review stage with evidence connected to each role criterion, rather than a collection of responses that still need a scoring framework.
Review the score and supporting evidence before deciding who progresses
With a large candidate pool, finding the relevant evidence can become as difficult as judging it. Reviewers need to see where that evidence appears without handing the decision to an automated ranking.
- Score against the Blueprint: Insights scores candidate responses against the employer-defined criteria in your Blueprint.
- Show the supporting evidence: Reviewers can inspect the analysis and return to the relevant candidate response.
- Keep the decision human-led: The hiring team reviews the evidence, resolves uncertainty and decides who should progress.
For instance, when a candidate receives a lower result for sales strategy, the reviewer can inspect the related evidence and decide whether the response falls below the agreed standard or needs a second look.
You end up with a focused shortlist and the evidence behind it. Your team can screen at scale while still understanding why each candidate did or did not progress.
Book a Willo demo to see how Blueprints, structured responses and evidence-led review work together.


.webp)


