Best practices for evaluation criteria

How to word and rank criteria so the scoring actually separates one candidate from another.

Updated

A badly written criterion doesn't throw an error. It returns a score that looks reasonable and separates nobody. These are the rules that make the most difference.

One thing per criterion

If a criterion mixes experience, languages and sector, the AI has to settle three questions with a single traffic light, and ends up scoring everyone amber.

Split them. Three criteria let you see why a candidate fits, not just how much.

State the threshold, not the adjective

"Senior", "solid experience" or "good level" mean different things to different people, and to the AI too. Give the number or the condition:

  • Solid experience leading teams
  • 4+ years leading teams, where leading means holding a lead, manager or C-level position

That's exactly what the platform's own criteria do: notice their levels talk about "4+ years… clearly shown in work history", not about "lots of experience".

Look after the amber level

Green and red are easy; amber is the one that decides whether the list is useful. It's the "looks right but isn't clear" bucket, and it should say what exactly isn't clear: ambiguous dates, a stack mentioned without context, an unexplained company jump.

A well-defined amber is your manual review queue. A vague amber is half the list.

Rank by what decides a hire

Order carries weight: the first one counts most. Before confirming, ask yourself which criterion would make you discard someone in two seconds, and put that on top.

Nice-to-haves go at the bottom, or straight off: a criterion that rules nobody out only flattens the scoring.

Prefer what can be read off a profile

The AI reads the candidate's profile: experience, headlines, skills. It can work out how long they've used a stack, or whether they came through consulting. It cannot know their salary expectation or whether they'd relocate.

If the data isn't in the profile, no criterion will surface it.

Test on a few candidates

Run the list with a low cap, look at how ten profiles you already know scored, and adjust the criteria before scaling. There's almost always a threshold to fix on the first pass, and fixing it before pulling 200 candidates costs no credits.

The pattern that works

[what to validate] + [concrete threshold] + [what counts as evidence]

For example:

4+ years using dbt, SQL and Python professionally, demonstrated in the work history and not only in the skills section.

Related articles