
Disagreement is data, not conflict
When two reviewers score the same recording three points apart, the rubric is ambiguous — not one of the reviewers. Treat every gap as a defect in the wording of a criterion and fix the wording.
The one-hour routine
Pick three recorded answers: one clearly strong, one clearly weak, one genuinely borderline. Everyone scores independently, then compares. Spend the time on the borderline case, because that is where every real decision lives.
Write the anchors down as you agree them: what a 2 looks like, what a 4 looks like, in the language of the role.
Anchor on behaviour, not impression
Criteria phrased as qualities — 'strong communicator', 'commercial' — collapse into gut feel. Criteria phrased as observable behaviour survive challenge: 'states the assumption before answering', 'quantifies the trade-off they chose'.
Recalibrate when the evidence drifts
Re-run the routine when the role changes, when a new manager joins the panel, or when score distributions shift between cohorts. Because every interview is recorded against the same scene, drift is something you can audit rather than guess at.


