Rubric Design
How to write interview rubrics an AI can score consistently
A rubric a machine can score is specific, observable, and ordered. The structure we use, the anti-patterns that break scoring, and how to test one first.
5 min readBy Herman Ko
A rubric an AI can score consistently is one where every criterion names an observable behaviour, every level is anchored in evidence a candidate could actually produce, and no two criteria measure the same thing. Get that right and consistency is free; get it wrong and consistency just means being wrong the same way every time.
That second half is the part worth sitting with. A machine applying a rubric identically to every candidate is only an improvement if the rubric encodes something real. The rubric is a human artefact, and it's where the hiring team's judgement actually lives.
What makes a rubric machine-scorable?
Four properties, and they're the same four that make a rubric useful to a human panel. The difference is that a machine will expose the flaws immediately instead of quietly papering over them.
- Observable. The criterion describes something the candidate says or does in the interview, not a trait they possess. "Explains a technical decision to a non-technical stakeholder" is observable. "Strong communicator" is not.
- Single-construct. One criterion, one thing. "Technical depth and collaboration" is two criteria wearing one coat, and any score against it is uninterpretable.
- Evidence-anchored. Each level says what a candidate at that level would have said. Not "good understanding" but "names the specific tradeoff and the condition under which they'd reverse it."
- Ordered. Moving from level 2 to level 3 should require strictly more, not different. If the levels describe different flavours rather than increasing amounts, you have a taxonomy, not a scale.
How many criteria should a rubric have?
Four to six for a screening interview. Fewer than four and you're not differentiating; more than six and a 20-minute conversation spreads so thin that no criterion gets enough evidence to score honestly.
The discipline is the one that makes an Ideal Candidate Profile work: force yourself to rank, and accept that everything below the cut isn't assessed at this stage. Same failure mode, described in how to write an ICP that actually screens candidates. A rubric with fifteen criteria is a rubric with no priorities.
How do you write the levels?
Write the top level first, from a real answer you'd be delighted by. Then write the bottom level from a real answer you've actually heard and been unimpressed by. Fill the middle last, and only with what genuinely sits between them.
Working example, for a criterion called handles ambiguous requirements:
- Level 1 — Describes waiting for clarification. No account of what they did while blocked.
- Level 2 — Asked someone. Names who and what they asked, but the resolution came from elsewhere.
- Level 3 — Made a defensible assumption, states the assumption explicitly, and says how they flagged it.
- Level 4 — All of level 3, plus names the signal that would have told them the assumption was wrong and what they'd have changed.
Each level is a sentence about what the candidate said, and a reader could sort two transcripts against it without knowing the role. That's the bar. If two careful people would put the same answer in different bands, the levels aren't tight enough yet.
What are the anti-patterns?
Five, in roughly the order we encounter them.
- Adjective ladders. Levels defined as poor / fair / good / excellent. This is a scale with no content; every rater fills it with their own standard.
- Trait criteria. "Cultural fit," "drive," "executive presence." Unfalsifiable from a transcript, and historically the place where bias enters a process wearing a lanyard.
- Pedigree proxies. Criteria that reward where someone worked rather than what they demonstrated. If a school name can move the score, the rubric is scoring the resume again.
- Compound criteria. Two constructs in one line. Split them or drop one.
- Criteria the interview can't reach. If nothing in a 20-minute conversation could produce evidence for it, the criterion belongs to a later stage. Move it rather than guess.
How do you test a rubric before running it live?
Score people you already have opinions about. Take six past candidates — two you hired and rated well, two you passed on, two you argued about — and have two colleagues score them against the draft rubric independently.
You're looking for two things. Where the two humans disagree, the levels are ambiguous and need rewriting. Where they agree with each other but disagree with the actual outcome, the rubric is measuring something other than what your team values — a more interesting problem, and worth resolving before a live candidate sees it. Expect to rewrite half the rubric, which is cheaper before a hundred interviews than after.
What does the AI add once the rubric is good?
Consistency and evidence. The same rubric is applied to every candidate, every time, with no good-mood interviews and no Friday-afternoon interviews, and every score links to the transcript moment that earned it.
It also enforces the discipline: a vague criterion produces a visibly vague score with thin evidence attached, so rubric weaknesses surface in week one rather than at quarter end. As CATALO put it: "Lantern gave us a structured, consistent way to assess every candidate against the same criteria. That's been a real step forward for fairness." The structure comes first; consistency is what the system contributes.
Common questions
Can we reuse one rubric across similar roles?
Partly. Criteria often transfer within a role family; level definitions usually need adjusting, because "strictly more" means something different for a junior and a senior version of the same job. Reusing levels unchanged is how a senior rubric quietly becomes a years-of-experience filter.
How often should a rubric be revisited?
When the role changes, when hiring managers start disagreeing with scores, or when a criterion stops differentiating — if everyone lands on level 3, that criterion is no longer doing work. Recalibrating against actual hiring outcomes is the useful loop.
Who should own the rubric?
The hiring manager, with the recruiter drafting. The manager owns the standard because they own the outcome; the recruiter owns the structure because they've seen what breaks. A rubric written entirely by one of them tends to fail in the other's direction.
Lantern applies whatever rubric you write to every candidate identically, with the evidence attached. If you'd like a second read on a rubric before you run it, bring it to a walkthrough.