Structured hiring

Structured interviews beat gut feel, and it is not close

The research has been settled for forty years. Structured interviews roughly double the predictive power of unstructured ones — and most teams still run unstructured ones.

Priya Raghunathan · · 4 min read

Ask a hiring manager how the interview went and you will usually get an answer about the person: sharp, a bit quiet, good energy, not sure about them. Ask what question produced that impression and the answer gets vaguer. That gap — between a confident judgement and the evidence that produced it — is the entire problem with unstructured interviewing, and it has been measured repeatedly since the 1980s.

What the research actually says

The finding that keeps replicating is unglamorous: interviews predict job performance moderately well when they are structured, and barely at all when they are not. Structure roughly doubles the correlation. Structured here has a specific meaning, and it is not "we had a list of topics":

  • every candidate gets the same questions, in the same order
  • the questions are tied to things the job actually requires
  • answers are scored against a rubric written before anyone interviewed
  • scores are recorded independently, before interviewers talk to each other

Drop any one of those and most of the gain goes with it. The fourth is the one teams abandon first and miss the significance of — a debrief where the most senior person speaks first is not four opinions, it is one opinion with three echoes.

Why unstructured interviews feel better

Because they are more enjoyable, and because confidence is not calibrated to accuracy. An unstructured interview is a conversation, and conversations produce a strong sense of having understood someone. The classic demonstration of this is grim: interviewers given deliberately random answers still formed confident impressions and still rated candidates as if the answers meant something.

There is a second reason, less often admitted. Unstructured interviews are faster to prepare. Writing a rubric takes an afternoon. Turning up and having a chat takes no preparation at all, and the cost lands somewhere else — six months later, on someone else's team.

What a rubric costs to write

Less than you think, and once per role rather than once per candidate. Four to six criteria is the working range. Fewer and you are not discriminating between candidates; more and interviewers start scoring the same underlying thing three times and calling it corroboration.

For a backend engineering role, a rubric that survives contact with reality looks something like:

  • System design — can they reason about failure, not just draw boxes
  • Coding — does the code work, and is it code someone else can change
  • Incident judgement — what they did when it broke, not what they would do
  • Communication — can they explain a decision to someone who disagrees

Weight them. If system design is twice as important as communication for this req, say so in the rubric rather than in the debrief, where it will be argued about by whoever cares most.

Anchor the scale, or it will drift

A 1-to-10 scale with no anchors is four different scales in a four-person panel. Write one line per level for each criterion — what a 4 looks like, what a 7 looks like, what a 9 looks like. Interviewers will still disagree, but they will disagree about the candidate rather than about the meaning of 7, and that is a disagreement worth having.

Ask about what happened, not what would happen

Hypotheticals measure imagination and fluency. "How would you handle a disagreement with a product manager?" gets you the answer the candidate believes you want. "Tell me about the last time you disagreed with a product manager — what did you do, and how did it end?" gets you an account with details that can be probed, and the follow-up is where the signal is.

The follow-up matters more than the question. A candidate who says they improved reliability has told you nothing until you ask what the error rate was before, what they changed, and how they knew it worked.

The honest objection

Structure is accused of filtering out unusual candidates — the self-taught engineer, the career-changer, the person who interviews badly but works brilliantly. It is worth taking seriously, and the evidence points the other way. Unstructured interviews are where similarity bias does its best work, because with nothing to score against, interviewers fall back on how much the person reminds them of people who have succeeded before. Structure is what protects the candidate who does not look like the last person you hired.

What structure genuinely costs you is the pleasure of the conversation, and the feeling of having made a call rather than read a report. That is a real loss. It is also the thing that was never predicting anything.

Where to start

Take one open role. Write four criteria and weight them. Write one behavioural question per criterion, with the follow-up you would ask if the answer were thin. Have every interviewer score independently before the debrief.

You will notice something within a handful of candidates: the debrief gets shorter and more specific, and the disagreements become tractable. That is not because the tool is clever. It is because everyone is finally arguing about the same thing.

Related

Get started

Screen every candidate like you had a full panel.

Without scheduling a single call. Set up one job, send one link, and read the report tomorrow morning.

Start freeBook a demo

No card. 1 job free forever. Cancel by closing the tab.