How to Evaluate the AI Features in Your Recruitment Software

Coding Liquids tutorial cover featuring Sagnik Bhattacharya for How to Evaluate the AI Features in Your Recruitment Software.
Coding Liquids tutorial cover featuring Sagnik Bhattacharya for How to Evaluate the AI Features in Your Recruitment Software.

Test each AI feature on your own closed roles: run it on past vacancies where you know who was placed, and check whether it ranks those people highly, explains why, treats comparable candidates alike and saves measurable recruiter time. Anything you can't test, explain to a client or switch off shouldn't touch live candidates.

You already have the answer key: your placement history. Most agencies evaluate recruitment AI by watching a demo on the vendor's data, or by asking recruiters whether it "feels useful". Neither tells you whether it would have found the person you actually placed. An afternoon with ten closed roles does.

Follow me on Instagram@sagnikteaches

Start with an inventory of what's switched on

Recruitment platforms have added AI quickly, often switching features on by default. List every one before testing anything. Bullhorn, for example, lists AI search and match, its Amplify chat assistant, data enrichment, CV formatting and call transcription in its Pro tier, and agentic screening and matching in Max; both tiers are priced by quote. Manatal includes AI recommendations and AI scoring from its $15-per-user plan (billed annually). Workable's screening agent is metered at one credit per applicant screened. A filled-in inventory for a mid-sized agency might look like this (illustrative):

Connect on LinkedInSagnik Bhattacharya
FeatureOn?Who uses itWhat data it seesExtra cost
CV parsingYesEveryone, automaticallyEvery CV receivedIncluded
Candidate-to-job match scoresYes3 of 8 recruitersFull databaseIncluded in tier
AI candidate summaries for clientsYes2 recruitersCV and notesIncluded in tier
AI outreach draftingNo––Add-on
Call transcriptionYes1 recruiterRecorded callsIncluded in tier

Two findings are typical. Match scores are on for everyone, but only three recruiters look at them, so nobody knows if they're right. And call transcription is running for one recruiter without anyone having decided how candidates are told. Fix the second before testing the first.

Subscribe on YouTube@codingliquids

Budget about a day for the six tests below, done by one experienced recruiter with an administrator's help: roughly 90 minutes for parsing, two hours for matching, an hour for the fairness pairs, 90 minutes for summaries, 30 minutes for logs and controls, and an hour for cost and time. It's worth doing in a quiet week, and worth repeating, so keep the test CVs and roles in a folder you can reuse.

Test 1: parsing on 30 real CVs

Why: every other AI feature works on the parsed record. If parsing drops a job title or a date, matching and summaries inherit the error.

How: pick 30 CVs from the last month, deliberately including awkward ones: two-column layouts, scanned PDFs, CVs with tables, career changers, contractors with many short roles. For each, check six fields against the original: name, current title, current employer, employment dates, key skills, location preference.

Pass mark: at least 95% of fields correct (171 of 180). Below 90%, recruiters are retyping and the downstream features can't be trusted.

What tends to fail: contractors' overlapping assignments merged into one job; the most recent role missed on two-column layouts; skills lifted from a "courses I'd like to take" section.

The contractor failure is worth seeing field by field, because it's easy to miss on a quick glance. An illustrative CV lists three assignments: "Project engineer, water treatment upgrade, Jan–Jun 2025", "Commissioning engineer, pumping station, Apr–Sep 2025" (overlapping, part time) and "Project engineer, reservoir works, Oct 2025 to date". The parsed record showed one job: "Project engineer, Jan 2025 to date", with the employer taken from the first assignment. Every field looked plausible, which is why it passed a skim. The commissioning experience, the one thing that made this candidate unusual, had vanished from the record that matching reads. Score a record like this as three field errors (title, employer, dates), not one, and note it as a contractor pattern for the vendor.

Test 2: does matching find the people you placed?

Why: a match score that doesn't rank your placements highly isn't measuring what makes a good candidate for your clients.

How: take ten closed roles with a known placement. Re-run the matching against the candidate pool as it stood at the time (or as close as you can get) and note where the placed candidate ranks.

RoleApplicants and database matchesPlaced candidate's rankIn top 10?
Site manager1403Yes
Payroll administrator2107Yes
Field service engineer9522No
Operations manager1801Yes
Estimator6041No
Five further roles–2, 4, 5, 9, 124 of 5

That's 7 of 10 placed candidates in the top 10 (illustrative). Now look at the misses. The field service engineer came from a different industry with transferable skills; the estimator's CV used the job title "quantity surveyor". Both are cases where a recruiter's judgement beat keyword similarity. Pass mark: placed candidate in the top 10 for at least 7 of 10 roles, with misses you can explain. If the misses have no pattern, the scoring is noise.

One trap makes this test look better than it is. The placed candidate's record has usually been updated since the placement: their CV now includes the very job you placed them in, with its title and skills. Re-run matching on that record and they'll rank near the top almost automatically. In one illustrative run, the operations manager ranked first on the current record but 14th on the CV as it stood when they applied. Use the CV version from the application date, which most systems keep as an attachment, and if you can't find it, drop that role from the test rather than count an inflated rank.

Test 3: comparable candidates, comparable scores

Why: a feature can find good candidates overall and still treat some groups unfairly, and in recruitment that's both a legal and an ethical problem.

How: make paired copies of five real CVs where only one thing changes: the name, a two-year career gap, a graduation year 20 years earlier, or a part-time current role. Label every copy clearly as a test record so nobody contacts it, and delete the copies when you've finished. Run each pair against the same job. The scores should be identical or within a point or two.

One agency's result (illustrative): four of five pairs scored within one point. The pair with a two-year career gap dropped from 82 to 64. The gap was labelled "caring responsibilities", which in practice correlates with sex. That's the kind of pattern you need to find before a candidate or a client does. Across real outcomes, a common rule of thumb is to investigate when one group's pass rate at a screening stage is below four-fifths of the highest group's. The sum is quick: if 60 of 120 applicants without a career gap pass the AI screen (50%) and 28 of 80 with a gap do (35%), the ratio is 35 ÷ 50, or 0.70. That's under 0.8, so the screen needs a closer look even if nobody meant it to treat the groups differently. The full method is in bias checks every recruiter should run on AI CV screening.

Pass mark: no pair differs by more than a few points on a factor unrelated to the job. Any bigger gap means switching off score-based filtering until the vendor explains and fixes it.

Test 4: summaries and outreach a client would sign off

Why: AI candidate summaries go to clients under your name. One invented claim damages trust in every summary you send.

How: generate summaries for 20 candidates and check every factual claim against the CV and your notes. Here's one output (illustrative) for an operations manager:

Experienced operations manager with 11 years in logistics. Led a team
of 12 across two depots, delivering a 15% reduction in fuel costs.
Holds a recognised project management qualification. Available on
four weeks' notice and seeking a salary of around 58,000.

Checked against the CV, two claims were wrong. The CV said "part of a 12-person operations team", not "led"; and the fuel-cost figure appeared nowhere. It seems to have been generated to make the achievement sound concrete. The notice period and salary were correct, from the recruiter's notes. Pass mark: zero invented claims in 20 summaries. This is a feature where 95% isn't good enough; a single fabricated achievement in front of a client is a reputational event. Writing candidate summaries clients read covers a review routine.

If you switch on outreach drafting, test it the same way before any message reaches a candidate. An illustrative draft to a passive candidate: "Hi [first name], I noticed you've been at your current company for six years, so you're probably ready for a new challenge. This role pays well above what you're likely on now." Two problems. It presumes the candidate wants to leave, and it guesses at their salary, which reads as intrusive whether or not it's right. The version worth sending names the role, one specific reason their background fits and the salary range the client has approved, and asks whether they'd like to hear more. Score 20 drafts on two questions: would a recruiter send it unedited, and does it say anything about the candidate that isn't on their public profile or CV?

Test 5: explanations, overrides and logs

  • Explanations: for any score, can a recruiter see why, in terms they could repeat to a candidate or client? "Matched on: 6 of 8 required skills, 4 years in similar role" passes. A bare percentage fails.
  • Overrides: can a recruiter move a candidate up or reject the AI's ranking, and does the system record who did it and why?
  • Logs: if a candidate asks why they weren't put forward, can you reconstruct what the AI scored and what the recruiter decided, with dates?
  • Off switch: can an admin turn each feature off for the whole account, or for a particular desk?

Any "no" here is a limit on how you can use the feature, not just a nice-to-have. A feature that can't be explained or overridden shouldn't filter candidates out; at most it can suggest.

Test the logs with a mock request rather than a look at the settings. Pick a candidate from a closed role who wasn't put forward and write, from the system alone, the answer you'd give if they asked why. A passing reconstruction reads something like (illustrative): "Applied 3 March. AI match score 58, below the recruiter's shortlist of 12 candidates scoring 71 to 90. Reviewed by the desk recruiter on 5 March; not shortlisted because the role required a site-management qualification not shown on the CV. No automated rejection." If you can't get past the first sentence without asking the recruiter what they remember, the logs fail, however detailed the vendor's audit page looks.

Test 6: cost and time per placement

Metered AI makes the cost visible per role. Workable's September 2026 pricing is a useful model: screening uses one credit per applicant, sourcing two, and a completed candidate chat ten; every plan includes 3,000 free credits, with further credits at $0.095 to $0.12 depending on pack size. A role with 250 applicants costs 250 credits to screen, so the free allowance covers about 12 such roles, and each role beyond that costs roughly $25–$30 to screen at those rates.

Time is the other half. Time a sample of five roles with and without the feature, from applications closed to shortlist sent. If screening drops from about four hours to about 90 minutes per role, and your recruiters' time is worth $40 an hour, that's around $100 saved per role against $25–$30 of credits. Time the steps, not just the total, because that shows where the saving really comes from. For one illustrative payroll administrator role: reading and sorting applications fell from 3 hours 40 minutes to 40 minutes, but checking the AI's top 25 added 30 minutes that didn't exist before, and writing the shortlist email stayed at 20 minutes either way. The net saving is still large, but the checking step belongs in the sum, and it's the first thing that gets skipped when a desk is busy. The per-role sums are set out further in how much time AI saves a recruitment agency per role.

Questions the vendor must answer in writing

  • Is our candidate data used to train or improve models, including the vendor's own? Can we opt out?
  • Which AI providers process candidate data, and where?
  • What exactly does the match score measure, and has the vendor tested it for bias? Can we see a summary of the results?
  • How are candidates told that AI is used, and can a candidate ask for a human review?
  • If we place candidates with employers in the EU: the EU AI Act lists AI used for recruitment and candidate screening as high-risk, with those obligations now deferred to 2 December 2027. How does the vendor plan to meet them? Any chatbot that talks to candidates must already tell them it's AI, a transparency duty that has applied since 2 August 2026.
  • Can we export our data, including AI scores and logs, if we leave?

This is not legal advice; if the answers leave you unsure, ask your data-protection adviser or an employment lawyer.

Putting the results on one page

A completed scorecard for the illustrative agency above:

FeatureTest resultDecision
CV parsing94% of fields correct; contractors' roles mergedKeep; recruiters check contractor records
Match scores7 of 10 placements in top 10; misses explainableKeep as a suggestion, never as a filter
Match fairnessCareer-gap pair dropped 18 pointsNo score-based rejection until the vendor responds
Client summaries2 invented claims in 20Recruiter checks every claim; vendor ticket raised
Call transcriptionAccurate; no candidate noticePaused until notice wording agreed

Repeat the tests after every major update from the vendor, and at least once a year. AI features change underneath you, and a feature that passed in spring can behave differently by autumn. If the results suggest your current system can't do the job, whether to bolt AI screening onto your ATS or switch is the next decision.

Questions recruiters ask about testing AI

How many past roles do I need to test matching properly?

Ten closed roles with a known placement is enough to spot a feature that doesn't work, and twenty gives a steadier picture. Choose roles across your main desks and seniority levels, and include at least two where the placed candidate had an unusual background, because those are the cases where keyword-driven matching usually fails and where good recruiters add most value.

Do we need candidates' consent to test AI on their CVs?

Testing on records you already hold, for the purpose of checking your own tools, is usually covered by the reasons you hold the data, but check your privacy notice and ask your data-protection adviser. Use closed roles, keep the test inside the software, don't export CVs to other tools, and delete any test copies once you've scored the results.

Should we switch off AI features that fail a test?

Yes, or restrict them to uses where the failure doesn't matter. A matching feature that misses placed candidates can still help with duplicate detection, for example. What you shouldn't do is leave a failing feature on and rely on recruiters to ignore it; under time pressure people follow the ranking in front of them. Record what you switched off and why.

Further reads

Sources: Bullhorn small agency pricing page; Manatal pricing page; Workable press release on AI recruiting agents and credit pricing (14 September 2026); EU AI Act (Annex III and Article 50) with the Digital Omnibus deferral dates.

Want your recruitment software's AI tested properly?

On a 1:1 call we'll list the AI features switched on in your system, choose the closed roles to test them against, and set pass marks your team and your clients would accept.

Book a 1:1 call with me