Test each AI feature on your own closed roles: run it on past vacancies where you know who was placed, and check whether it ranks those people highly, explains why, treats comparable candidates alike and saves measurable recruiter time. Anything you can't test, explain to a client or switch off shouldn't touch live candidates.
You already have the answer key: your placement history. Most agencies evaluate recruitment AI by watching a demo on the vendor's data, or by asking recruiters whether it "feels useful". Neither tells you whether it would have found the person you actually placed. An afternoon with ten closed roles does.
Start with an inventory of what's switched on
Recruitment platforms have added AI quickly, often switching features on by default. List every one before testing anything. Bullhorn, for example, lists AI search and match, its Amplify chat assistant, data enrichment, CV formatting and call transcription in its Pro tier, and agentic screening and matching in Max; both tiers are priced by quote. Manatal includes AI recommendations and AI scoring from its $15-per-user plan (billed annually). Workable's screening agent is metered at one credit per applicant screened. A filled-in inventory for a mid-sized agency might look like this (illustrative):
| Feature | On? | Who uses it | What data it sees | Extra cost |
|---|---|---|---|---|
| CV parsing | Yes | Everyone, automatically | Every CV received | Included |
| Candidate-to-job match scores | Yes | 3 of 8 recruiters | Full database | Included in tier |
| AI candidate summaries for clients | Yes | 2 recruiters | CV and notes | Included in tier |
| AI outreach drafting | No | – | – | Add-on |
| Call transcription | Yes | 1 recruiter | Recorded calls | Included in tier |
Two findings are typical. Match scores are on for everyone, but only three recruiters look at them, so nobody knows if they're right. And call transcription is running for one recruiter without anyone having decided how candidates are told. Fix the second before testing the first.
Budget about a day for the six tests below, done by one experienced recruiter with an administrator's help: roughly 90 minutes for parsing, two hours for matching, an hour for the fairness pairs, 90 minutes for summaries, 30 minutes for logs and controls, and an hour for cost and time. It's worth doing in a quiet week, and worth repeating, so keep the test CVs and roles in a folder you can reuse.
Test 1: parsing on 30 real CVs
Why: every other AI feature works on the parsed record. If parsing drops a job title or a date, matching and summaries inherit the error.
How: pick 30 CVs from the last month, deliberately including awkward ones: two-column layouts, scanned PDFs, CVs with tables, career changers, contractors with many short roles. For each, check six fields against the original: name, current title, current employer, employment dates, key skills, location preference.
Pass mark: at least 95% of fields correct (171 of 180). Below 90%, recruiters are retyping and the downstream features can't be trusted.
What tends to fail: contractors' overlapping assignments merged into one job; the most recent role missed on two-column layouts; skills lifted from a "courses I'd like to take" section.
The contractor failure is worth seeing field by field, because it's easy to miss on a quick glance. An illustrative CV lists three assignments: "Project engineer, water treatment upgrade, Jan–Jun 2025", "Commissioning engineer, pumping station, Apr–Sep 2025" (overlapping, part time) and "Project engineer, reservoir works, Oct 2025 to date". The parsed record showed one job: "Project engineer, Jan 2025 to date", with the employer taken from the first assignment. Every field looked plausible, which is why it passed a skim. The commissioning experience, the one thing that made this candidate unusual, had vanished from the record that matching reads. Score a record like this as three field errors (title, employer, dates), not one, and note it as a contractor pattern for the vendor.
Test 2: does matching find the people you placed?
Why: a match score that doesn't rank your placements highly isn't measuring what makes a good candidate for your clients.
How: take ten closed roles with a known placement. Re-run the matching against the candidate pool as it stood at the time (or as close as you can get) and note where the placed candidate ranks.
| Role | Applicants and database matches | Placed candidate's rank | In top 10? |
|---|---|---|---|
| Site manager | 140 | 3 | Yes |
| Payroll administrator | 210 | 7 | Yes |
| Field service engineer | 95 | 22 | No |
| Operations manager | 180 | 1 | Yes |
| Estimator | 60 | 41 | No |
| Five further roles | – | 2, 4, 5, 9, 12 | 4 of 5 |
That's 7 of 10 placed candidates in the top 10 (illustrative). Now look at the misses. The field service engineer came from a different industry with transferable skills; the estimator's CV used the job title "quantity surveyor". Both are cases where a recruiter's judgement beat keyword similarity. Pass mark: placed candidate in the top 10 for at least 7 of 10 roles, with misses you can explain. If the misses have no pattern, the scoring is noise.
One trap makes this test look better than it is. The placed candidate's record has usually been updated since the placement: their CV now includes the very job you placed them in, with its title and skills. Re-run matching on that record and they'll rank near the top almost automatically. In one illustrative run, the operations manager ranked first on the current record but 14th on the CV as it stood when they applied. Use the CV version from the application date, which most systems keep as an attachment, and if you can't find it, drop that role from the test rather than count an inflated rank.
Test 3: comparable candidates, comparable scores
Why: a feature can find good candidates overall and still treat some groups unfairly, and in recruitment that's both a legal and an ethical problem.
How: make paired copies of five real CVs where only one thing changes: the name, a two-year career gap, a graduation year 20 years earlier, or a part-time current role. Label every copy clearly as a test record so nobody contacts it, and delete the copies when you've finished. Run each pair against the same job. The scores should be identical or within a point or two.
One agency's result (illustrative): four of five pairs scored within one point. The pair with a two-year career gap dropped from 82 to 64. The gap was labelled "caring responsibilities", which in practice correlates with sex. That's the kind of pattern you need to find before a candidate or a client does. Across real outcomes, a common rule of thumb is to investigate when one group's pass rate at a screening stage is below four-fifths of the highest group's. The sum is quick: if 60 of 120 applicants without a career gap pass the AI screen (50%) and 28 of 80 with a gap do (35%), the ratio is 35 ÷ 50, or 0.70. That's under 0.8, so the screen needs a closer look even if nobody meant it to treat the groups differently. The full method is in bias checks every recruiter should run on AI CV screening.
Pass mark: no pair differs by more than a few points on a factor unrelated to the job. Any bigger gap means switching off score-based filtering until the vendor explains and fixes it.
Test 4: summaries and outreach a client would sign off
Why: AI candidate summaries go to clients under your name. One invented claim damages trust in every summary you send.
How: generate summaries for 20 candidates and check every factual claim against the CV and your notes. Here's one output (illustrative) for an operations manager:
Experienced operations manager with 11 years in logistics. Led a team
of 12 across two depots, delivering a 15% reduction in fuel costs.
Holds a recognised project management qualification. Available on
four weeks' notice and seeking a salary of around 58,000.
Checked against the CV, two claims were wrong. The CV said "part of a 12-person operations team", not "led"; and the fuel-cost figure appeared nowhere. It seems to have been generated to make the achievement sound concrete. The notice period and salary were correct, from the recruiter's notes. Pass mark: zero invented claims in 20 summaries. This is a feature where 95% isn't good enough; a single fabricated achievement in front of a client is a reputational event. Writing candidate summaries clients read covers a review routine.
If you switch on outreach drafting, test it the same way before any message reaches a candidate. An illustrative draft to a passive candidate: "Hi [first name], I noticed you've been at your current company for six years, so you're probably ready for a new challenge. This role pays well above what you're likely on now." Two problems. It presumes the candidate wants to leave, and it guesses at their salary, which reads as intrusive whether or not it's right. The version worth sending names the role, one specific reason their background fits and the salary range the client has approved, and asks whether they'd like to hear more. Score 20 drafts on two questions: would a recruiter send it unedited, and does it say anything about the candidate that isn't on their public profile or CV?
Test 5: explanations, overrides and logs
- Explanations: for any score, can a recruiter see why, in terms they could repeat to a candidate or client? "Matched on: 6 of 8 required skills, 4 years in similar role" passes. A bare percentage fails.
- Overrides: can a recruiter move a candidate up or reject the AI's ranking, and does the system record who did it and why?
- Logs: if a candidate asks why they weren't put forward, can you reconstruct what the AI scored and what the recruiter decided, with dates?
- Off switch: can an admin turn each feature off for the whole account, or for a particular desk?
Any "no" here is a limit on how you can use the feature, not just a nice-to-have. A feature that can't be explained or overridden shouldn't filter candidates out; at most it can suggest.
Test the logs with a mock request rather than a look at the settings. Pick a candidate from a closed role who wasn't put forward and write, from the system alone, the answer you'd give if they asked why. A passing reconstruction reads something like (illustrative): "Applied 3 March. AI match score 58, below the recruiter's shortlist of 12 candidates scoring 71 to 90. Reviewed by the desk recruiter on 5 March; not shortlisted because the role required a site-management qualification not shown on the CV. No automated rejection." If you can't get past the first sentence without asking the recruiter what they remember, the logs fail, however detailed the vendor's audit page looks.
Test 6: cost and time per placement
Metered AI makes the cost visible per role. Workable's September 2026 pricing is a useful model: screening uses one credit per applicant, sourcing two, and a completed candidate chat ten; every plan includes 3,000 free credits, with further credits at $0.095 to $0.12 depending on pack size. A role with 250 applicants costs 250 credits to screen, so the free allowance covers about 12 such roles, and each role beyond that costs roughly $25–$30 to screen at those rates.
Time is the other half. Time a sample of five roles with and without the feature, from applications closed to shortlist sent. If screening drops from about four hours to about 90 minutes per role, and your recruiters' time is worth $40 an hour, that's around $100 saved per role against $25–$30 of credits. Time the steps, not just the total, because that shows where the saving really comes from. For one illustrative payroll administrator role: reading and sorting applications fell from 3 hours 40 minutes to 40 minutes, but checking the AI's top 25 added 30 minutes that didn't exist before, and writing the shortlist email stayed at 20 minutes either way. The net saving is still large, but the checking step belongs in the sum, and it's the first thing that gets skipped when a desk is busy. The per-role sums are set out further in how much time AI saves a recruitment agency per role.
Questions the vendor must answer in writing
- Is our candidate data used to train or improve models, including the vendor's own? Can we opt out?
- Which AI providers process candidate data, and where?
- What exactly does the match score measure, and has the vendor tested it for bias? Can we see a summary of the results?
- How are candidates told that AI is used, and can a candidate ask for a human review?
- If we place candidates with employers in the EU: the EU AI Act lists AI used for recruitment and candidate screening as high-risk, with those obligations now deferred to 2 December 2027. How does the vendor plan to meet them? Any chatbot that talks to candidates must already tell them it's AI, a transparency duty that has applied since 2 August 2026.
- Can we export our data, including AI scores and logs, if we leave?
This is not legal advice; if the answers leave you unsure, ask your data-protection adviser or an employment lawyer.
Putting the results on one page
A completed scorecard for the illustrative agency above:
| Feature | Test result | Decision |
|---|---|---|
| CV parsing | 94% of fields correct; contractors' roles merged | Keep; recruiters check contractor records |
| Match scores | 7 of 10 placements in top 10; misses explainable | Keep as a suggestion, never as a filter |
| Match fairness | Career-gap pair dropped 18 points | No score-based rejection until the vendor responds |
| Client summaries | 2 invented claims in 20 | Recruiter checks every claim; vendor ticket raised |
| Call transcription | Accurate; no candidate notice | Paused until notice wording agreed |
Repeat the tests after every major update from the vendor, and at least once a year. AI features change underneath you, and a feature that passed in spring can behave differently by autumn. If the results suggest your current system can't do the job, whether to bolt AI screening onto your ATS or switch is the next decision.
Questions recruiters ask about testing AI
How many past roles do I need to test matching properly?
Ten closed roles with a known placement is enough to spot a feature that doesn't work, and twenty gives a steadier picture. Choose roles across your main desks and seniority levels, and include at least two where the placed candidate had an unusual background, because those are the cases where keyword-driven matching usually fails and where good recruiters add most value.
Do we need candidates' consent to test AI on their CVs?
Testing on records you already hold, for the purpose of checking your own tools, is usually covered by the reasons you hold the data, but check your privacy notice and ask your data-protection adviser. Use closed roles, keep the test inside the software, don't export CVs to other tools, and delete any test copies once you've scored the results.
Should we switch off AI features that fail a test?
Yes, or restrict them to uses where the failure doesn't matter. A matching feature that misses placed candidates can still help with duplicate detection, for example. What you shouldn't do is leave a failing feature on and rely on recruiters to ignore it; under time pressure people follow the ranking in front of them. Record what you switched off and why.
Further reads
- AI Bias in Small Business Decisions: Hiring, Pricing and Credit — The wider picture on bias in AI-assisted decisions.
- How to Write Job Adverts With ChatGPT Without Biased Language — Fixing bias earlier, in the advert, before screening starts.
- Questions to Ask a Recruitment Agency About Its AI Screening — The questions your clients may start asking you.
- AI for Recruitment Business Development: Winning New Clients — Using the time saved to win new clients.
- AI Literacy Requirements: What Your Staff Need to Know — What recruiters need to understand about the tools they use.
- AI CV Screening for Recruitment Agencies: Setup and Safeguards — Five set-up steps and the safeguards to have in place before an agency lets AI score CVs, with a rubric, prompt, sample output and back-test.
- How Staffing Agencies Automate Candidate Follow-Up With AI — The seven moments where candidates go quiet, which to automate first, message templates, AI reply sorting and the consent rules that keep texts welcome.
- AI Recruiting Tools for Small Businesses: What's Worth Paying For — Nine kinds of AI recruiting tool, what each costs, which a small business should pay for, and three to skip, with examples from real hiring tasks.
- How to Set Up an AI-Assisted Hiring Process for a Small Team — A hiring pipeline for small teams where AI writes, summarises and schedules, people decide, and every rejection is read by a human first.
- How to Build Interview Questions and Scorecards With AI — Turn the job's real outcomes into competencies, questions, probes and rating anchors with AI, so every interviewer scores the same answer the same way.
- Can AI Screen CVs Fairly? What Small Employers Need to Know — How AI CV screening goes unfair, what the rules expect of small employers, and a six-step set-up with tests you can run in an afternoon.
- AI Tools and AI Development: The Complete 2026 Guide — the AI hub, including every tutorial in the AI-for-business series.
Sources: Bullhorn small agency pricing page; Manatal pricing page; Workable press release on AI recruiting agents and credit pricing (14 September 2026); EU AI Act (Annex III and Article 50) with the Digital Omnibus deferral dates.