How to Build Interview Questions and Scorecards With AI

Coding Liquids tutorial cover featuring Sagnik Bhattacharya for How to Build Interview Questions and Scorecards With AI.
Coding Liquids tutorial cover featuring Sagnik Bhattacharya for How to Build Interview Questions and Scorecards With AI.

Start from four to six competencies drawn from what the job must achieve. Ask AI for two or three behavioural and situational questions per competency, each with follow-up probes, then for a 1-4 rating scale with a concrete description of each score. Every interviewer scores alone before discussing. AI builds the kit; people ask the questions and do the scoring.

The anchors are the part most small employers skip, and they're what makes the kit work. Without a written description of what a 2 looks like compared with a 3, two interviewers hearing the same answer will score it differently, and the debrief turns into whoever argues hardest. AI is fast at drafting those anchors. What it shouldn't do is score candidates itself: automated judgements about people are exactly where bias hides, and recruitment is a legally sensitive use of AI.

Follow me on Instagram@sagnikteaches

From job outcomes to competencies

Don't start with a list of generic competencies ("teamwork", "communication"). Start with what the person must achieve, then ask what capability each outcome needs. Here's an illustrative mapping for a five-person financial planning firm hiring a paraplanner:

Connect on LinkedInSagnik Bhattacharya
What the job must achieveCompetencyEvidence you're looking for
Suitability reports that pass compliance review first timeTechnical accuracyCatches errors in figures, product details and rationale before anyone else does
Research that supports the adviser's recommendationAnalysis and researchCompares options on the criteria that matter to the client, not just cost
Reports clients can actually understandWritten clarityExplains complex products in plain language without losing accuracy
Cases moving on time despite provider delaysOrganisation and follow-throughTracks many cases, chases without being chased, flags risks early
Advisers who trust the paraplanner's workJudgement and escalationKnows when to decide and when to raise something with the adviser

Five competencies is plenty for one interview. If you have more, some belong in the work sample instead. If you already have a role brief from the hiring process, the must-haves there should map almost one-to-one onto these; if they don't, fix the brief first. The tutorial on setting up an AI-assisted hiring process covers the brief.

Subscribe on YouTube@codingliquids

Prompting for questions that separate candidates

Generic prompts give you question-bank questions that every candidate has rehearsed. Give the AI the competency, the evidence, and what a weak answer sounds like:

Write interview questions for a paraplanner at a small financial
planning firm.
Competency: Technical accuracy
Evidence we want: catches errors in figures, product details and
rationale before anyone else does.
A weak answer sounds like: "I'm very detail-oriented and always
double-check my work."

Give me:
- 2 behavioural questions (about a real past situation)
- 1 situational question (a realistic scenario from this job)
- for each, 3 follow-up probes that get past rehearsed answers
Rules: open questions only; no questions a candidate could answer
with a general claim; nothing about age, family, health, religion,
nationality or other personal characteristics.

An illustrative extract of the response:

Behavioural 1: Tell me about a time you found a mistake in a report or
case file after it had been checked by someone else. What was it and
how did you find it?
  Probes: What made you look again at that part? What did you do next,
  and who did you tell? What did you change in how you work afterwards?

Behavioural 2: Describe how you check a suitability report before it
goes to the adviser. Walk me through a recent one.
  Probes: Which checks do you never skip? Which figures do you verify
  against source documents? What did your last check catch?

Situational: A provider illustration shows a different charge from the
one in your draft report, and the adviser needs the report today. What
do you do?
  Probes: Who do you contact first? What do you tell the adviser? What
  would make you hold the report back?

What you'd check before using them: the situational question is good because it has no perfect answer, so it reveals judgement. Behavioural 2 is the weakest: "walk me through how you check" invites a rehearsed process description. The probe "what did your last check catch?" rescues it, so keep the probe and consider making that the main question. And read each question for anything that could be heard as personal: none here, but generated sets sometimes include questions like "how do you manage stress at home?", which you should cut.

Questions to cut, even when AI suggests them

Generated question sets are usually decent, but a few types keep turning up that either waste interview time or create risk. Cut them and use the replacement:

  • "Where do you see yourself in five years?" Everyone has a rehearsed answer and it predicts little. Ask instead: "What would you want to be doing better in a year's time, and how would you get there?"
  • "What's your greatest weakness?" Invites a humblebrag. Try: "Tell me about feedback you received that you disagreed with at first. What did you do?"
  • Anything touching personal life: family plans, childcare, age, health, where someone is "originally from", religion. Illegal to base decisions on in many places and never relevant. If availability matters, ask everyone the same neutral question: "The role has an early start twice a month; is that workable for you?"
  • Brainteasers and trick hypotheticals: "How many golf balls fit in a bus?" tells you nothing about writing a suitability report.
  • Double questions: "Tell me about a time you missed a deadline and how you communicate with advisers generally." Candidates answer one half. Split them.

Read the final set aloud once. Anything you'd feel awkward being asked yourself probably needs rewording.

Probes: where the real evidence comes from

Most candidates prepare a good first answer. The probes find out whether it's theirs. Keep a short list of all-purpose probes on the scorecard, whatever the competency:

  • "What did you do, specifically, rather than the team?"
  • "What was the result? Is there a number you can put on it?"
  • "What would you do differently?"
  • "What did someone else think of how you handled it?"
  • "What happened next time?"

An illustrative exchange shows why they matter:

Candidate: "In my last role I introduced a checklist for suitability
reports that reduced errors across the team."
Probe: "What did you do, specifically?"
Candidate: "My manager asked for it; I drafted the first version from
the compliance findings over six months, and we used it on every case."
Probe: "What was the result?"
Candidate: "Files returned by compliance went from about one in five to
one in twelve over the next half-year, by our own tracking."

The first answer alone would score a 2 or a 3 depending on the interviewer's mood. After two probes there's specific, checkable evidence of a 4. That's the difference between an interview and a conversation.

Rating anchors two interviewers apply the same way

Ask AI to draft anchors for each competency on a 1-4 scale. Four points rather than five removes the comfortable middle score; people have to decide whether an answer was below or above the bar.

For the competency "Technical accuracy" (paraplanner), write a 1-4
rating scale. For each score, describe what the candidate's evidence
sounds like in 1-2 sentences. Make the difference between 2 and 3
especially clear, since that's the hiring bar. Base it on observable
evidence, not personality.

An illustrative result, lightly edited:

ScoreWhat the evidence sounds like
1 - No evidenceGeneral claims only ("I'm careful"). Can't give an example of catching an error, or the example is trivial.
2 - Some evidenceGives an example of finding an error, but it was found by chance or by someone else's prompt. Checks are described vaguely.
3 - Clear evidence (hiring bar)Describes specific checks they run every time, with a recent example of a real error they caught. Knows which figures to verify against source documents.
4 - Strong evidenceAs 3, plus they've improved how others check work (a checklist, a template, training) with a result they can describe.

What to edit in AI anchors: they often slip personality words into the top score ("demonstrates a passion for accuracy"). Replace them with something observable. And make sure the 3 describes someone you'd genuinely hire, since that's how interviewers will use it.

The scorecard itself

One page per candidate per interviewer. Weights reflect what matters most in the job; agree them before the first interview, not after you've met a candidate you like. An illustrative filled-in scorecard for one paraplanner candidate:

CompetencyWeightScore (1-4)Evidence noted (quotes, not impressions)
Technical accuracy30%4Built team checklist; compliance returns "one in five to one in twelve"
Analysis and research25%3Compared platforms on charges and fund range; less on service quality
Written clarity20%3Explained a drawdown option clearly; one jargon slip, corrected unprompted
Organisation and follow-through15%2"I keep on top of it"; example of a missed provider deadline, fix unclear
Judgement and escalation10%3Would hold the report and tell the adviser before the client call
Weighted total3.15Probe organisation further at second stage or in references

The weighted total is (4 × 0.30) + (3 × 0.25) + (3 × 0.20) + (2 × 0.15) + (3 × 0.10) = 1.20 + 0.75 + 0.60 + 0.30 + 0.30 = 3.15. Agree a rule for the total, such as "3.0 or above and no score of 1 on the top two competencies", but treat it as a guide for the debrief rather than an automatic answer.

The evidence column is the important one. "Seemed organised" is an impression; "example of a missed provider deadline, fix unclear" is evidence someone else can weigh.

A work sample built on a fictional client

For a paraplanner, a short work sample tells you more than any question about technical accuracy. The problem is material: you can't hand candidates a real client file. AI solves that neatly. Ask it to create a fictional but realistic case: a couple in their fifties, their income, pensions, savings, objectives and attitude to risk, plus two provider illustrations with slightly different charges, and one deliberate inconsistency between the fact-find and an illustration.

Then check the case yourself. Generated cases often contain arithmetic that doesn't add up (a pension pot that doesn't match its contributions) or products that don't quite exist. Fix the figures, and make sure the deliberate inconsistency is the only one. The task for candidates: "In 90 minutes, draft the recommendation summary and flag anything in the file you'd query with the adviser."

Mark it with the same anchors. A candidate who spots the planted inconsistency and explains why it matters scores a 3 or 4 on technical accuracy regardless of how they interviewed. Keep the task time-boxed, tell candidates exactly how long it should take, and consider paying for it if it runs longer than an hour or two. The tutorial on paraplanning with AI is useful background if the role will involve AI tools, and testing candidates' AI skills in an interview covers adding that to the kit.

Running the debrief without anchoring on the loudest voice

The order of the debrief matters as much as the scorecard:

  1. Everyone submits scores before talking. A shared form or a sealed spreadsheet tab. Once one person says "I loved her", the rest adjust without noticing.
  2. Go competency by competency, not candidate by candidate. Compare everyone's evidence on technical accuracy, then on analysis, and so on.
  3. Discuss any gap of more than one point. Each interviewer reads their evidence aloud. Usually one person heard a probe answer the other missed.
  4. Change a score only for evidence, never for persuasion. Note why it changed.
  5. Decide, and record the reason in one line.

AI can help with the admin here. Paste in the submitted scorecards and ask for a comparison table: candidates as rows, competencies as columns, each cell showing both interviewers' scores and quoted evidence, with gaps over one point highlighted. Tell it explicitly not to rank, recommend or summarise who did best. An illustrative extract:

             Tech accuracy     Analysis        Organisation
Candidate A  4 / 4             3 / 3           2 / 3  (gap: discuss)
Candidate B  3 / 2 (discuss)   4 / 4           3 / 3
Candidate C  2 / 2             3 / 2 (discuss) 4 / 4

That table turns a 90-minute debrief into a 40-minute one, and the conversation stays on the gaps.

What to keep out of AI tools during interviewing

Four boundaries worth writing into the process:

  • No AI scoring or ranking of candidates. Not from transcripts, not from notes, not from recordings. If you recruit in the EU, the EU AI Act treats AI used to evaluate candidates as high-risk, with obligations due from 2 December 2027, and AI that infers people's emotions in the workplace, which the Commission's guidance says includes recruitment, has been prohibited since 2 February 2025.
  • No personality or emotion analysis from video, voice or writing, wherever you hire. The claims are hard to validate and easy to challenge.
  • Recording only with consent, asked in advance, with a real option to decline and no effect on the outcome. Many small firms decide typed notes are enough.
  • Business accounts only for candidate notes and scorecards, with a set retention period. Delete what you don't need once the role is filled.

For how bias can creep into decisions like these even without AI scoring, see AI bias in small business decisions.

Reusing the method for a different role

The kit is role-specific, but the method transfers in an afternoon. Take an illustrative marketing agency hiring an account executive. Its outcomes are different: clients who renew, campaigns delivered on the agreed dates, and reports that tell clients something useful. So the competencies become client communication, organisation across several accounts, commercial awareness, and turning data into a recommendation.

The prompt stays the same shape; only the inputs change. A situational question AI produced for this role, lightly edited: "A client emails at 5pm asking why their campaign results dropped this week, and you don't know yet. What do you send tonight, and what do you do tomorrow morning?" The anchors change too. For client communication, a 3 might read "acknowledges quickly, sets a clear time for a full answer, doesn't guess at causes", while a 4 adds "proposes what they'll check and gives the client one useful thing to do meanwhile".

What doesn't change: four or five competencies, two questions each with probes, a 1-4 scale with the hiring bar at 3, independent scoring, and a mock before the first real interview. Keep a master prompt document with the three prompts from this tutorial, and each new role takes about two hours to prepare instead of a day.

Calibrate the kit with a mock interview first

Before the first real candidate, run a 30-minute mock with a colleague (ideally someone who does a similar job) answering as themselves. Both interviewers score independently using the anchors, then compare.

An illustrative calibration from the planning firm: on written clarity, one interviewer gave a 2 and the other a 4 for the same answer. The anchor for 3 said "explains complex products clearly"; one interviewer took "clearly" to mean no jargon at all, the other to mean accurate and well-structured. The fix was to rewrite the anchor with an observable test: "explains a product so that a client with no financial background could repeat the key point back; uses a technical term only when it's explained". On the second mock, they scored within one point.

Also time the mock. If ten questions with probes ran to 55 minutes in a 45-minute slot, cut two questions now rather than rushing real candidates. Keep the final kit, prompts and anchors together; for the next role, you'll only need to change the competencies and let AI draft the rest again. The same competencies can then feed the new starter's first review, and writing the job description with AI shows how to keep advert, interview and review aligned from the start.

Scorecard and interview questions people ask

How many questions should a 45-minute interview have?

About eight to ten core questions, two per competency, leaving time for probes, the candidate's own questions and a clean finish. If you plan more, you'll rush the follow-ups, which is where the useful evidence comes from. Cut the weakest questions after your mock interview rather than squeezing them in.

Should every candidate get exactly the same questions?

The core questions, yes, in the same order, so you're comparing like with like. Probes will differ because they follow what each candidate says. You can add one or two questions specific to someone's application, such as a gap in evidence on a must-have, as long as you'd ask any candidate with the same gap.

Is it fair to send candidates the questions in advance?

Sending the topics, or even the questions, is increasingly common and tends to produce better, less nervous answers. It favours preparation over quick thinking, which suits most small-business roles. If you do it, send them to everyone at the same time, and lean on probes to get past rehearsed answers.

Can AI summarise my interview notes for the debrief?

It can collate everyone's scores and quoted evidence into one comparison table, which saves time. Don't ask it who performed best, to adjust scores or to recommend a candidate, and don't feed it recordings for analysis. Keep the judgement with the interviewers, and keep the notes in a business account.

Further reads

Sources: AI Act service desk page on Article 5 and the Commission's guidelines on prohibited AI practices; project fact sheet on EU AI Act high-risk timelines. Checked September 2026.

Want an interview kit you can reuse for every hire?

On a 1:1 call we'll turn your role brief into competencies, questions and anchors, run a mock scoring together, and leave you with prompts you can reuse for the next role.

Book a 1:1 call with me