How to Vet an AI Consultant's Case Studies and References

Coding Liquids tutorial cover featuring Sagnik Bhattacharya for How to Vet an AI Consultant's Case Studies and References.
Coding Liquids tutorial cover featuring Sagnik Bhattacharya for How to Vet an AI Consultant's Case Studies and References.

Ask for two or three case studies that match your size and problem, then test each claim: the starting point and the result, the dates, the tools, and who did the work. Call at least two references, choosing one yourself from their client list rather than only the ones offered, and ask what broke after the consultant left.

Most case studies aren't false. They're selected: the best project out of a dozen, measured at the best moment, described in the kindest terms. So the job is less about catching lies and more about finding out whether the result would repeat for an organisation like yours, and whether it lasted once the consultant had gone.

Follow me on Instagram@sagnikteaches

The vetting checklist, grouped by what you're testing

Work through the groups in order. The first two can be done from documents in an evening; the last three need conversations.

Connect on LinkedInSagnik Bhattacharya

A. Is the case study specific?

  • The client is described by size and type. Why: a result for a 200-person firm tells you little about a four-person office. Verify: if the client isn't named, ask for the headcount, sector and what they used before.
  • The problem is stated in numbers. Why: "drowning in admin" can't be compared with anything. Verify: ask what was measured before the work started, and how.
  • The tools are named. Why: "an intelligent assistant integrated with their systems" could be a spreadsheet formula or a custom app. Verify: ask which products it runs on and who holds the accounts.
  • The dates are given. Why: AI tools change fast; a 2024 project may not be repeatable in the same way now. Verify: ask when it started, when it went live and when the result was measured.
  • The result has a baseline and a measurement period. Why: "cut admin by 60%" means nothing without "from what, measured over which weeks". Verify: ask for both, in writing.
  • The cost and duration are at least roughly stated. Why: a result that took six months and a large budget isn't comparable with a four-week project on a small one. Verify: ask for a cost band and how long it took from start to go-live.

B. Is the evidence real?

  • You can see it working. Why: screenshots and slides can be mock-ups. Verify: ask for a live walkthrough or a screen recording of the actual system, with client details hidden if needed.
  • The figures come from the client's records. Why: consultants sometimes report their own estimates as results. Verify: ask whether the client measured the result or the consultant did.
  • The consultant's role is clear. Why: "we delivered" may mean they subcontracted the build, or joined a bigger team. Verify: ask who did the work day to day, and whether that person would work on your project.

C. Did it last?

  • It's still in use six to twelve months on. Why: many automations are switched off quietly when the person who understood them leaves. Verify: ask the reference directly.
  • Someone at the client can maintain it. Why: a system only the consultant understands is a dependency, not a result. Verify: ask what documentation and handover the client received; the AI consultant handover checklist lists what good looks like.

D. Does it match you?

  • Similar size, budget and data sensitivity. Why: a project that worked with a full-time operations manager may fail with a part-time administrator. Verify: compare the reference's team with yours.
  • Similar tools, or a reason to switch. Why: a consultant whose every case study uses one platform will probably recommend it to you. Verify: ask whether they've built the same thing on the tools you already have.

E. Are the references genuine?

  • At least two, one of your choosing. Why: offered references are the happiest clients. Verify: ask for a list of recent clients and pick one yourself.
  • An arm's-length relationship. Why: a reference who is a friend, a business partner or a former colleague isn't independent. Verify: ask the reference how they came to work with the consultant.

Taking apart one case study, claim by claim

Here's an invented case study of the kind that often appears on consultants' websites:

Subscribe on YouTube@codingliquids

How we helped a community centre cut admin by 60% with AI. The centre's small team was overwhelmed by bookings. We implemented an intelligent booking assistant integrated with their systems, delivering a 60% reduction in admin time and 95% customer satisfaction within weeks. "It's changed how we work," says the centre manager.

It reads well and says almost nothing you can check. Each claim becomes a question:

ClaimWhat's missingQuestion to ask
"Cut admin by 60%"Baseline, method, periodHow many hours a week before and after, and who measured them?
"Intelligent booking assistant"The actual toolsWhich products does it run on, and is it built or configured?
"Integrated with their systems"Which systems, and howWhat does it connect to, and what happens when one of those changes?
"95% customer satisfaction"Survey size and a before figureHow many people answered, and what was the score before?
"Within weeks"When measured; still true?What are the figures now, a year on?
The manager's quoteWhether you can talk to themMay I call the manager?

Send the questions in writing before any call, and note which ones get specific answers. A consultant with a real result usually replies with figures and offers the reference without being pushed. One who answers "every client is different" to all six has told you the case study was marketing copy.

A realistic twist worth knowing about: more consultants now draft case studies with AI tools, which is fine for the writing but risky for the numbers, because an assistant filling out a draft will happily supply a plausible percentage. If a figure appears in the case study but the consultant can't say where it came from, treat it as unverified.

Arithmetic that exposes an inflated result

Many inflated claims fall apart under a quick sum you can do on the back of the case study. Four common ones:

  • "Saved 20 hours a week" in a three-person office. Three people working 37.5 hours is 112.5 hours a week. If about a third is admin, that's roughly 37 hours, so the claim is that more than half of all admin disappeared. It's possible when one big manual job was automated, but ask exactly which 20 hours and how they were counted.
  • "Response times cut by 90%." From ten hours to one, perhaps, but for which messages? If the AI answered the easy half instantly and the hard half still waited a day, the average falls sharply while the customers who mattered waited just as long. Ask for the median and for the slowest 10%.
  • "400% ROI in three months." Ask for the two numbers behind it. If the project cost $3,000, the claim is $12,000 of benefit in three months, or $4,000 a month. For a small organisation that's close to a full-time salary. Did someone's paid hours actually fall, or is it time valued at an hourly rate that never left the bank?
  • "Automated 80% of enquiries." "Automated" can mean answered and closed, or merely acknowledged by an auto-reply. Ask how many of the 80% needed no human follow-up at all.

None of these sums proves a claim wrong. They tell you which question to ask first, and a consultant who welcomes the question is a better sign than a perfect number.

Choosing references yourself instead of accepting the list

Every consultant offers references who'll speak well of them. That's expected, and those calls are still worth making. But add one reference you chose, so at least one call isn't curated.

Three ways to find it. Ask for the names of the last five to ten clients and pick one. Ask for "a project that was difficult, and the client from it", which is the most revealing request you can make; a consultant who has one and shares it is showing you how they handle problems. Or look at who has publicly recommended or worked with the consultant, and contact one of them directly and politely.

Some clients genuinely can't be named, especially where confidentiality agreements apply. That's reasonable, but then ask the consultant to arrange a call with that client rather than accepting no reference at all. If you're going to share your own data during the project, the same confidentiality questions will apply to you; whether to sign an NDA before sharing data with a consultant covers that side.

A 20-minute reference call script, with sample notes

Reference calls go wrong when they're friendly chats that end with "so you'd recommend them?". Ask for specifics instead. A script that fits in twenty minutes:

1. How did you come to work with [consultant]?
2. What was the problem before, and how did you measure it?
3. What exactly did they build or set up? Who did the work day to day?
4. What result did you see, and when did you measure it?
5. Is it still in use? Who looks after it now?
6. What broke or needed fixing after they finished, and how did
   they respond?
7. Did the price and timeline match the quote? What changed?
8. What would you do differently if you hired them again?
9. Is there anything you'd want to know in my position?

Write notes during the call, not afterwards. Here are illustrative notes from a good reference:

Reference 2 (community hall manager, 4 staff, chosen from the client list). Call length 18 minutes.
Relationship: met through a sector network, no other connection.
Before: about 25 booking enquiries a week by email, around 5 minutes each.
Built: booking form, shared calendar, automatic confirmation and draft invoice. The consultant built it herself.
Result: about 2 minutes per enquiry, measured over the two months after launch against the two months before.
Still running: yes, 14 months on. One fix when the calendar app changed a setting; fixed within two days, one hour charged.
Price and timeline: fixed fee as quoted; a week late, because we were slow giving access.
Differently: "We'd have written our booking rules down first. She had to drag them out of us."
Would rehire: yes.

Compare that with the notes from a weak reference: "Lovely to work with, very knowledgeable, it's been great, definitely recommend." Friendly, and empty. When a reference can't answer questions 2, 4 and 5, either they weren't close to the project or there isn't much to report.

Public reviews and badges: what they do and don't prove

Public signals help you shortlist, but they answer different questions from a reference call:

  • Review sites with verification. Clutch checks a reviewer's identity through LinkedIn, Google or a company email address, compares their online footprint with what they submitted, and has an editor check every review; when a firm submits clients as references, its analysts may interview them by phone. That proves the reviewer is real and was a client. It doesn't make the sample random, since firms usually choose whom to put forward.
  • Vendor tiers and certifications. A Zapier, Make or HubSpot partner tier shows the consultant has passed the vendor's requirements and has a track record with that product. It says nothing about work with organisations like yours, and partners earn from the vendors they recommend.
  • Marketplace ratings. Star ratings on freelance platforms skew high. The written reviews for jobs similar to yours are the useful part.
  • Endorsements on social profiles. Often written by connections, sometimes in exchange for one back. Treat them as introductions, not evidence.

The same logic applies when the thing being sold is software rather than consulting; how to check an AI software vendor's customer references adapts these checks for vendors.

A church office runs the checks on two consultants

To see the whole checklist working, follow a church office through it; the office and both consultants are invented. The office has a part-time administrator and wants help automating hall-hire bookings: enquiry replies, the calendar and invoices. Two consultants reached the shortlist through the process in where to find a good AI consultant.

Consultant X sent a polished case study: the community-centre story above, almost word for word. Group A exposed it quickly. Asked for the baseline, X said the 60% "was based on the first week's feedback". Asked for tools, X named a platform the office had never heard of, which X's firm resold. The offered reference, the centre manager, was friendly but couldn't say whether the system was still in use: "I think we went back to email after our volunteer coordinator left." And asked who did the build, X said a partner firm had handled "the technical side".

Consultant Y sent two plainer case studies with dates, tools and before-and-after hours. The administrator picked a reference from Y's client list: the community hall manager whose notes appear above. The system was still running 14 months on, with one fix handled promptly, and Y had done the work personally.

The administrator recorded the results against the five groups, which made the choice easy to explain to the church council:

Checklist groupConsultant XConsultant Y
A. SpecificFail: no baseline; 60% came from first-week feedbackPass: hours before and after, dates, tools named
B. Evidence realPartly: demo of the platform, not of the client's systemPass: screen recording of the live booking flow
C. LastedFail: probably abandoned after a volunteer leftPass: running 14 months, one fix after a calendar change
D. Matches usWeak: relies on a platform X resellsPass: built on tools the office already uses
E. References genuineOffered reference only; vague on outcomesOne offered, one chosen by us; both specific

The checks took the administrator about three hours across a week: an evening reading and writing questions, two reference calls and one follow-up email. X's case study didn't contain a single false sentence. It just couldn't survive any of the questions. Y got the job, with a fixed fee for the first phase.

Reference-check answers that should end the conversation

Some responses to the checklist are reasons to stop, however good the rest looks:

  • "All our clients are under NDA", with no offer of an anonymised figure or an arranged call.
  • Results only ever as percentages, never with the hours, costs or volumes they're percentages of.
  • No working system to show, only slides and mock-ups.
  • A reference who turns out to be a friend, relative or business partner, and wasn't described as one.
  • Vagueness about who does the work, when the consultant plans to pass it to someone you haven't met.
  • Irritation at the questions. You'll be asking harder ones during the project.

Any one of these is enough to walk away. For warning signs that show up earlier, in proposals and sales calls, see AI consultant red flags. And if a consultant passes everything above, the references you gathered become useful again later: they tell you what "done well" looked like for someone like you, which is the standard to hold your own project to.

Further reads

Sources: Clutch help pages on how reviews are verified; Zapier, Make and HubSpot partner directory and tier pages.

Want a second opinion on a consultant's case studies?

On a 1:1 call we'll go through the case studies and proposal you've been sent, check whether the claimed results add up for an organisation your size, and list the questions to put to the references.

Book a 1:1 call with me