Give every candidate the same short, realistic task from the job, 20 to 30 minutes with an AI tool you provide, and watch how they work: how they brief the tool, whether they check its output against the source, what they refuse to paste in, and how they fix errors. Score it on a written rubric.
Asking "do you use ChatGPT?" tells you almost nothing, because almost anyone will say yes. What separates a useful hire is judgement: noticing when the output is wrong, and knowing when AI shouldn't be used at all. So the most revealing tests include a deliberate mistake in the source material. Fast, polished output that repeats the mistake is a worse sign than slower work that catches it.
Decide what "AI skills" means for this job
An office coordinator and a marketing assistant need different things from AI, so start by writing one sentence: "In this role, AI will mostly be used to…" Then pick what to test from that.
| Role | What good looks like | Test with | Don't bother testing |
|---|---|---|---|
| Office coordinator or admin | Turns messy notes into clear schedules and replies; catches inconsistencies | Job notes, a customer email, a price list | Image generation, advanced automation |
| Estimator or quoting | Uses AI to draft and structure quotes, never to set prices unchecked | A site description and a partial price list with a wrong total | Marketing copy |
| Marketing assistant | Drafts in the firm's voice, checks every claim, spots made-up facts | Three real job photos' descriptions and a brand note | Spreadsheet formulas |
| Customer service | Drafts calm, accurate replies and knows when to hand to a person | An angry email and your policy page | Data analysis |
| Site supervisor | Knows what not to put into AI and how to check an answer | Two interview questions, no exercise | Anything longer |
If the role is new, writing a job advert for an AI-savvy admin assistant helps you describe it before you test for it.
The five things worth scoring
- Briefing. Do they give the tool context, a clear task and constraints, or type one vague line and hope?
- Checking. Do they compare the output against the source material, especially names, dates and numbers?
- Data judgement. Do they notice information that shouldn't go into an AI tool, and leave it out?
- Finishing. Is the final result fit to send: right tone, right length, errors fixed by hand where needed?
- Knowing the limits. Can they say what the tool got wrong, and when they wouldn't use it?
Speed is deliberately missing. Fast candidates who skip the checking are the ones who later send a customer the wrong appointment date with great confidence.
Build the exercise around a planted error
Use real work with names and details changed. Then plant three traps, because the traps are what separate careful people from confident ones:
- A contradiction between two source documents (a date or a quantity), which AI tools tend to smooth over rather than flag.
- A piece of information that shouldn't be pasted in, such as a door code or bank details sitting in the notes.
- A number that doesn't add up, such as a line total that's wrong.
Provide the AI tool yourself: ideally a business account such as ChatGPT Business or Claude Team, which by default keeps your content out of model training, or a personal account with the model-training switch turned off. Everyone gets the same tool, the same material and the same time.
Example: an office coordinator for an electrical contractor
An illustration. Say a seven-person electrical contractor is hiring an office coordinator. The pack contains notes from three of tomorrow's jobs, a complaint email from a customer, and a one-page price list. The planted traps: the complaint says the electrician visited on Tuesday while the job notes say Wednesday; one job's notes include the customer's alarm code; and the notes quote a consumer unit replacement at a price that doesn't match the price list.
The traps work best when they sit in ordinary-looking material. Extracts from the pack might read (illustrative):
JOB NOTES - Thursday
Job 2 08:00 Consumer unit replacement, 3-bed house. Customer at
work: key in the key safe, code 4471. Alarm code 2580,
panel by the back door.
Quoted: $1,150 (standard, per price list)
PRICE LIST (extract)
Consumer unit replacement, up to 10 ways ............ $1,050
Additional circuit ................................... $140
JOB HISTORY - last week
30 Sep (Wed): kitchen socket fault, Mr P. Fault found on ring main,
made safe, return visit to be booked.
COMPLAINT EMAIL (received this morning)
"Your electrician came on Tuesday and left half my kitchen sockets
dead. Nobody has called to rebook. I've been using an extension lead
for a week..."
Nothing is labelled as a trap. A candidate who pastes the whole pack into the assistant hands it two access codes; one who reads first sees them. The price only looks wrong if you check it against the list, and the date conflict only appears if you read the complaint and the job history together.
The brief handed to candidates:
CANDIDATE BRIEF (25 minutes)
You are the office coordinator. You can use the AI assistant on this
laptop as much or as little as you like. We are interested in how you
work, not only the result.
Using the attached pack:
1. Draft a reply to the customer's complaint email. It should be ready
for the owner to approve and send.
2. Write tomorrow's schedule for the two electricians: one short
message each, with addresses, times and anything they need to know.
3. List anything in the pack that looks wrong, inconsistent or
risky, and what you would do about it.
Think aloud if you're comfortable doing so. At the end we'll spend
ten minutes talking through your choices.
Sit where you can see the screen, or have them share it on a video call. Make brief notes as they go: what they typed, what they pasted, what they changed.
A scoring rubric you can copy
Score each area from 1 to 4 straight after the exercise, before you discuss it with anyone else. Checking and data judgement count double, because they're the skills that protect the business.
| Area | 1: Weak | 2: Patchy | 3: Solid | 4: Strong |
|---|---|---|---|---|
| Briefing (x1) | One vague line | Task stated, no context | Context, task and format given | Also sets constraints ("don't promise a refund") and refines the draft |
| Checking (x2) | Accepts output as is | Skims it | Catches one planted error | Catches the contradiction and the wrong total, and says which source they trusted |
| Data judgement (x2) | Pastes the alarm code in | Pastes it, then realises | Leaves it out | Leaves it out and flags that it shouldn't be in the notes at all |
| Finishing (x1) | Sends raw AI output | Light edits, tone still off | Accurate, appropriate tone | Could go out today with no changes |
| Knowing the limits (x1) | "It's always right" | Vague caution | Names what the tool got wrong | Names it and says where they'd not use AI in this job |
Maximum weighted score: 28. A reasonable bar for a role that will use AI daily is 20, with no score of 1 in checking or data judgement. Agree the bar before you see anyone.
The briefing row is easier to score once you've seen both ends of it. A briefing that scores 1:
reply to this complaint [pastes email]
And one that scores 4:
I'm the office coordinator at a small electrical contractor. Draft a
reply to the complaint below for my manager to approve. Apologise for
the delay in rebooking, offer the first available return visit, and
keep it under 120 words. Don't accept liability for the fault or
promise any refund or discount; the manager decides that. Use a
calm, plain tone. Leave [DATE] and [TIME] as placeholders.
[pastes the complaint, with the customer's name removed]
The second one sets constraints the business would care about (no promised refund, a manager's approval), removes the name, and leaves placeholders rather than letting the tool guess a date. A candidate who then says "that's too formal, make it warmer" and edits the result earns the "refines the draft" part of a 4.
Settle one edge case before the first interview: the candidate who doesn't use the assistant at all. Some capable people will do the whole task by hand, often well. Don't mark them down for that alone, but the role needs the skill, so add one line to everyone's brief: "Please use the assistant for at least part of task 1." If someone still declines, score the other four areas as normal, mark briefing as not seen, and ask in the discussion how they would have briefed the tool for the complaint reply. Their total is then out of 24, and the equivalent of a 20-out-of-28 bar is about 17.
Follow-up questions, and what good answers sound like
Use the ten-minute discussion after the exercise, or these questions alone for roles that don't need the full test. Strong answers are specific; weak answers are general enthusiasm.
- "Tell me about a time an AI tool gave you something wrong. How did you notice?" Strong: a specific example and the check that caught it. Weak: "It's never really been wrong for me."
- "What would you never put into a chatbot at work?" Strong: names categories such as customer personal details, passwords and codes, and anything confidential, and mentions checking which account they're using. Weak: "Nothing too private, I suppose."
- "AI gives you a total for a quote. How do you check it?" Strong: recalculates it in a spreadsheet or by hand; knows these tools can slip on arithmetic. Weak: "I'd ask it to double-check."
- "When would you choose not to use AI in this job?" Strong: a delicate complaint, anything legally sensitive, or a first reply to a very upset customer that needs a personal touch. Weak: "I'd use it for everything, it saves time."
- "Your manager asks you to use AI for something you think is risky. What do you do?" Strong: says so, explains the risk, suggests a safer way. Weak: either "I'd just do it" or "I'd refuse."
- "How do you keep your AI skills current?" Strong: a small, realistic habit. Weak: a list of tool names with nothing about how they use them.
Rehearsed answers give way under one follow-up. A candidate asked the first question might say, "I always fact-check everything AI gives me." Ask: "What was the last thing you checked, and what did you check it against?" Someone with the real habit answers at once, with something like "a delivery date it gave me for a supplier order; I checked the order confirmation and it was a week out." Someone repeating good advice drifts back to generalities: "Just making sure it's right, really." Two follow-ups like that on any question above tell you more than the first answer did.
If you want AI to help write the rest of your interview questions and scorecards, building interview questions and scorecards with AI covers that; just keep the scoring itself human.
Worked example: scoring two candidates
Continuing the illustration. Candidate A finished in 14 minutes with a polished complaint reply and tidy schedules. They pasted the entire notes file, alarm code included, and the reply apologised for "Wednesday's visit", taking the job notes' date without noticing the conflict. They didn't spot the price mismatch. Weighted score: briefing 3, checking 1 (x2), data judgement 1 (x2), finishing 3, limits 2 = 12.
Candidate A's complaint reply, as drafted (illustrative):
Dear Mr P,
I'm so sorry about Wednesday's visit and the inconvenience of having
no kitchen sockets. As a gesture of goodwill we'll refund the
call-out charge in full, and an electrician will be with you first
thing tomorrow to complete the repair.
It reads well, which is the problem. The day is the one from the job history, not the customer's, and nobody checked which was right. The refund is a promise the owner never authorised. And "first thing tomorrow" commits an electrician to a slot the pack shows is already taken, since Job 2 starts at 08:00. Each came from the assistant filling a gap in a one-line brief. A coordinator who sends drafts like this opens a new complaint for every one they close.
Candidate B took the full 25 minutes. They removed the alarm code before pasting, asked the tool to list inconsistencies between the documents, confirmed the date conflict against the original email themselves, and wrote "check which day before sending" as a note for the owner. They recalculated the consumer unit line by hand. Their reply was plainer but accurate. Weighted score: briefing 3, checking 4 (x2), data judgement 4 (x2), finishing 3, limits 3 = 25.
B's schedule message for Job 2 shows the same judgement in the output itself:
"Job 2, 08:00: consumer unit replacement, 3-bed house. Customer at work. For the key safe and alarm codes, ring the office before you set off. The price in the notes ($1,150) doesn't match the price list ($1,050); the office is checking, so don't quote either figure on site."
The codes never went into the assistant or into a text message, and the price problem reached the electrician as an instruction rather than a surprise on the customer's doorstep.
Candidate A looks better in the first five minutes and would cost more in the first five months. The rubric makes that visible before the job offer, not after.
If two people interview, score separately and compare afterwards. Disagreements on the first candidate are common and useful. One interviewer might give Candidate A a 2 for checking, because A asked the tool to "double-check the reply"; the other gives a 1, because nothing was compared against the source. The rubric's wording settles it: skimming is a 2, accepting output as it stands is a 1, and asking the tool to check its own work without opening the source is closer to the second. Write the agreed reading beside the rubric row, so later candidates are scored on the same basis as the first.
Live, take-home, or both?
| Format | Shows you | Watch out for | Use when |
|---|---|---|---|
| Live, 20 to 30 minutes | How they actually work, including checking | Nerves; keep it short and friendly | Almost always; it's the most reliable |
| Take-home, 1 hour maximum | A more polished result | You can't see the process, or who did it | Marketing or writing roles, followed by a live discussion |
| Questions only | Attitudes and judgement | Easy to rehearse good-sounding answers | Roles where AI is a minor part of the job |
Keep any take-home task short and clearly an exercise. Don't set real work you intend to use; candidates notice, and it's unfair to them.
For a marketing role, the same planted-error idea translates easily. An artisan bakery hiring a part-time marketing assistant might set a one-hour take-home: write three product descriptions for the website and one social post, from a brand note, an ingredients sheet and three photos. The trap sits in the brand note, which says the seeded loaf is "suitable for a gluten-free diet", while the ingredients sheet lists wheat flour for that loaf. AI tools will happily repeat the brand note's claim in fluent copy. In the live follow-up, ask: "Which claims in your descriptions did you check, and against what?" The candidate worth hiring has already spotted it, or spots it the moment you ask; the one to worry about defends the copy because "that's what the brief said".
Keep it fair and defensible
- Same task, same tool, same time for every candidate, and the same person scoring where possible.
- Offer five minutes with the tool first. You're testing judgement, not whether they've seen this particular interface before.
- Make adjustments on request, such as extra time or a screen reader, for candidates with a disability, and ask about this when you invite them.
- Don't mark people down for lacking a paid subscription or for preferring a different tool; the skills transfer.
- Keep the rubric sheets with your other interview notes, so you can explain a decision if asked.
- Take advice from an HR adviser if you're unsure how a test fits the rules that apply to your recruitment.
Whoever you hire, the test also tells you where their training should start. Carry their weakest rubric area straight into how you onboard new hires onto your AI tools and rules, and check it against the AI literacy your staff need generally.
Questions about AI skills tests in interviews
Should candidates use their own AI tool in the test?
Provide one tool for everyone instead. Candidates with a paid subscription would otherwise have an advantage unrelated to the job, and you can't control what happens to your test material in a personal account. Use a business account that doesn't train on content by default, or a personal account with the model-training switch turned off, and give each candidate five minutes to get familiar with it first.
Is it worth testing AI skills for a role that barely uses AI?
Test lightly. For a site role, two questions are enough: what they would never put into a chatbot, and how they would check an AI answer about something that matters, such as a regulation or a product specification. You're checking for safe judgement, not fluency. A full exercise only makes sense where AI will be part of the daily work.
Should I ban AI from written application tasks?
Banning it is hard to enforce and tells you little. A better approach is to allow it, say so, and then discuss the submission in the interview: ask the candidate to talk through one choice they made and one thing they changed from the first draft. People who understand their own work explain it easily; people who pasted it in usually can't.
Further reads
- How to Set Up an AI-Assisted Hiring Process for a Small Team — Where this test fits in the whole hiring process.
- How to Hire Your First Customer Service Person Alongside AI — Hiring for a role that works next to AI tools.
- How to Write Job Descriptions With AI That Attract Good Hires — Describe the AI part of the role accurately.
- Can AI Screen CVs Fairly? What Small Employers Need to Know — Before you let AI shortlist the candidates you'll test.
- How to Train Staff to Use AI in a Small Business — Close any gaps the test reveals after hiring.
- AI Bias in Small Business Decisions: Hiring, Pricing and Credit — Keep AI out of the parts of hiring where it skews decisions.
- AI or a New Hire? How to Decide Before You Recruit — Break a planned role into tasks, see which ones AI can absorb, cost hire against AI over a year, and set a clear trigger for recruiting anyway.
- Will AI Replace My Employees? An Honest Answer for Small Firms — An honest answer for owners: what the evidence shows, a task-by-task method to score each role, a bookkeeping firm example and what to tell staff.
- How to Write Job Adverts With ChatGPT Without Biased Language — A four-step method for bias-free job adverts with ChatGPT, with a requirements sheet, drafting and audit prompts, sample outputs and a full before and after.
- Questions to Ask a Recruitment Agency About Its AI Screening — Stage-by-stage questions on an agency's AI screening, how to grade the answers, and how a grooming salon found good candidates its agency's filter had rejected.
- How to Hire a Part-Time Marketing Assistant Who Uses AI — Decide what 12 hours a week should produce, write the job description, test candidates on fixing AI output, and set up accounts before day one.
- AI Tools and AI Development: The Complete 2026 Guide — the AI hub, including every tutorial in the AI-for-business series.
Sources: vendor documentation on business-plan data handling for ChatGPT Business and Claude Team (checked September 2026). The exercise and scores are illustrative.