Judge it on six things, in writing: whether the agreed scope was delivered, what changed against your baseline numbers, whether staff still use it at day 90, what your team can run and change without the consultant, the quality of their advice (including what they told you not to buy), and how they behaved when things went wrong.
The timing matters as much as the method. At handover, everything looks good because everything is new. Real value shows at 30 to 90 days, once the automation has met a busy week, a staff change and a few odd cases. So check deliverables at handover, check outcomes and adoption at day 30, and make the final judgement at day 90. Most advice on how to evaluate a consultant focuses on hiring one; this is about the other end of the engagement.
Six measures of a consultant's value, weighted
A single number like "hours saved" misses too much. A consultant can save hours with something your team can't maintain, or deliver modest savings while stopping you from an expensive mistake. Score each of these from 0 to 5 and weight them by what mattered to you when you hired:
| Measure | Question it answers | Suggested weight |
|---|---|---|
| Scope delivered | Did we get what the proposal promised, meeting its acceptance criteria? | 20% |
| Outcomes against baseline | What changed in time, errors, response times or revenue, measured the same way as before? | 30% |
| Adoption at day 90 | Are people still using it, or have they drifted back to the old way? | 15% |
| Team independence | Can we run, adjust and pause it ourselves using the documentation? | 15% |
| Quality of advice | Did they steer us well, including away from things we didn't need? | 10% |
| Conduct | Were they honest about problems, on time and clear about costs and any tool commissions? | 10% |
Adjust the weights before you score, not after. If the main reason you hired was to learn how to do this yourselves, independence might deserve 30%. If it was a single time-critical automation, outcomes might deserve 40%.
Three check-ins: handover, day 30 and day 90
Spread the judgement over three short sessions rather than one verdict. Each looks at different things:
- At handover (about an hour): go through the promised-versus-delivered table, run the acceptance tests yourself, and confirm you hold every account, export and document. Don't release the final payment until this is done.
- At day 30 (about 30 minutes): first outcome numbers against the baseline, early adoption, and a list of workarounds staff have invented. Send the consultant anything that's inside the original scope while any support period is still running.
- At day 90 (about an hour): the full scorecard, the money sum, and the independence test. This is the judgement that counts, and the one to use if you're asked for a reference.
Put the three dates in the diary on the day you sign. Without them, day 90 tends to pass unnoticed, and any problems surface months later when nobody can remember what was agreed.
Start from the promise: a promised-versus-delivered table
Take the proposal or statement of work and list every deliverable and acceptance criterion on its own line. Next to each, write what was delivered and the evidence. It sounds bureaucratic; it takes about half an hour and prevents the most common unfairness in both directions: blaming a consultant for something never agreed, or crediting them for something promised but never quite finished.
A chat assistant can draft the table from the proposal and your notes, provided you stop it filling gaps:
Below are (1) the consultant's proposal and (2) my notes on what
was delivered. Build a table with one row per promised deliverable
or acceptance criterion. Columns: promised (quote the proposal),
delivered, evidence from my notes, status.
Status must be MET, PARTLY MET, NOT MET or NO EVIDENCE.
Use NO EVIDENCE whenever my notes don't mention it.
Don't infer delivery from invoices or from the proposal itself.
[paste proposal text]
[paste your notes]
In an illustrative run, the draft was mostly right but made two telling errors. It marked "one-hour staff training session" as MET because the final invoice listed it, although the notes said the session was cancelled and replaced by a recorded video. And it marked "measurable reduction in first-reply time" as MET without any number, because the consultant's handover email said replies were "much faster". The fixes: change the training line to PARTLY MET and decide whether the video does the job, and refuse MET on any outcome without a figure. Those are the same two mistakes people make when they judge from memory.
If the proposal had no acceptance criteria, write the ones you'd reasonably have expected, and be fair about it. Next time, put them in before work starts; scoping an AI project with deliverables and acceptance criteria shows how. The difference looks like this:
- Vague: "Improve guest communication with AI." Measurable: "Median first reply under 30 minutes between 8am and 10pm over four weeks; no message with a door code sent without a booking-ID check."
- Vague: "Automate cleaning coordination." Measurable: "Every confirmed booking creates a cleaning job within 15 minutes; cleaners are notified by 6pm the day before; no missed cleans over 30 days."
Outcomes against the baseline, not against the consultant's dashboard
Outcomes should be measured the way you measured before the project: the same stopwatch, the same spreadsheet, the same definitions. If you never recorded a baseline, reconstruct one honestly from timesheets, email timestamps or a week of timing the old way on a sample. Measuring time saved after rolling out AI has practical methods.
Be wary of the consultant's own activity numbers. A campsite owner might receive a monthly report headed "9,800 operations run", which sounds impressive. On Make, every scheduled check uses a credit even when nothing new has arrived, and three scenarios checking every 15 minutes account for about 8,600 of those operations on their own. Operations, tasks and "AI actions" measure activity, not value. The figures that count are the ones you'd notice without the report: hours back, errors avoided, replies sent faster, bookings won.
Also watch the calendar. A tour operator who judges at week two, in the quiet season, may see almost nothing: at 30 enquiries a week, saving 4 minutes each is two hours. The same automation at 120 enquiries a week in peak season saves eight. Judge across a representative period, or scale the quiet-period numbers to your real annual volumes.
Adoption: is anyone still using it at day 90?
An automation nobody uses delivers nothing, however well it was built. Check adoption with evidence, not by asking people whether they like it:
- Usage logs. Run histories in Zapier or Make, message counts in the tool, or for Microsoft 365 Copilot, the usage report in the admin centre, which shows active users over 7, 28, 90 and 180 days.
- Share of work going through it. Of last week's 140 guest messages, how many were answered from an AI draft?
- Workarounds. Who still does it the old way, and when? Workarounds are often the most useful finding, because each one points to a gap.
Low adoption isn't automatically the consultant's fault. If staff were never given time to learn the new process, that's on the business. But a consultant who handed over a tool without showing the people who'd use it, or designed something that adds steps for staff, owns part of the problem.
What your team can do without them
The quiet test of a good engagement is whether you still need the consultant for routine changes. Try three things yourself, using only the documentation: pause and restart the main automation, change something small (a message template, a price, a staff member's contact details), and find out why yesterday's run failed. If any of these needs a call to the consultant, note why. A missing page in the documentation is easy to fix. A design that hard-codes a cleaner's phone number inside an automation, so a staff change needs a rebuild, is a design fault worth raising. The AI consultant handover checklist lists what you should hold by the end.
Advice that saved money counts too
Part of what you pay a consultant for is judgement, and some of the best judgement is "don't". If they talked you out of a $149-a-month website chatbot because most guest questions arrived through the booking platform's messaging, not the website, that's $1,788 a year that belongs in the value column. So does a recommendation to switch on a feature in software you already had rather than buying a new tool, or a warning about a plan that would have locked you in for twelve months. Write these down with a rough figure; they're easy to forget because nothing visible happened.
The opposite applies too. Advice that led you to pay for things you didn't use, or that overlooked a cheaper route you found later, belongs on the other side of the ledger. Judge the advice by what was knowable at the time, though, not with hindsight about a price change or a product launch nobody could have predicted.
Separating their performance from your side of the bargain
Every AI project depends on the client too: access to systems, clean enough data, decisions made on time, and staff time for testing. Before scoring, list the delays and changes that came from your side. Things that were yours to provide:
- system access and API keys, on the dates agreed;
- test data and examples;
- decisions on wording, rules and exceptions;
- staff time for testing and training;
- any change of scope after work started.
Things that were theirs: realistic estimates, sound design, testing before go-live, documentation, clear communication and honesty about limits. If you lost nine working days waiting for your own lock-system provider to issue an API key, a late finish isn't a mark against the consultant. If they never tested a booking made at midnight, it is.
Scoring a holiday-let project: a filled-in scorecard
Here's the method applied to an illustrative holiday-let manager with 18 properties and two staff, who paid a consultant a fixed fee over six weeks for three deliverables: AI-drafted guest replies with automated pre-arrival messages, an automated cleaning schedule, and monthly AI summaries of each property's bookings for the owners.
The baseline, measured before work started: guest messaging took 11 hours a week, with a median first reply of 3 hours 40 minutes in the daytime; cleaning coordination took 4 hours a week, with two missed cleans in the previous quarter; owner statements took 6 hours a month.
What happened: in week 4, the consultant found the accounting exports too inconsistent for reliable owner summaries, recommended dropping that deliverable, and took $600 off the fee. In week 2 of live running, a date-mapping error sent one guest the previous guest's door code; the consultant fixed it the same day and added a check that matches booking IDs before any code is sent. At day 90:
- guest messaging took 5 hours a week, median first reply 25 minutes between 8am and 10pm;
- cleaning coordination took 1.5 hours a week, with no missed cleans;
- staff answered about 80% of messages from AI drafts; one person still typed weekend replies from scratch because the drafts lacked the weekend key-safe instructions;
- the manager could change templates alone but needed the consultant when a cleaner left, because the cleaner's number was hard-coded (fixed free under the 30-day support clause);
- running costs were about $85 a month in tools and AI usage.
| Measure | Weight | Score (0 to 5) | Why | Weighted |
|---|---|---|---|---|
| Scope delivered | 20% | 4 | Two of three delivered; third dropped honestly with a fee reduction | 0.80 |
| Outcomes against baseline | 30% | 5 | 8.5 hours a week saved; replies more than eight times faster | 1.50 |
| Adoption at day 90 | 15% | 4 | About 80%; weekend gap identified and fixable | 0.60 |
| Team independence | 15% | 3 | Templates fine; staff changes needed the consultant | 0.45 |
| Quality of advice | 10% | 5 | Talked the manager out of a $149-a-month chatbot | 0.50 |
| Conduct | 10% | 4 | Door-code error fixed fast and openly, but should have been caught in testing | 0.40 |
| Total | 100% | 4.25 of 5 |
And the money. The 8.5 hours a week saved come to about 442 hours a year; at an illustrative $25 an hour that's $11,050. Take off $1,020 a year in running costs and the net is about $10,030, or roughly $836 a month, so the reduced fee of $2,600 paid back in a little over three months, before counting the $1,788 a year of chatbot fees avoided. For the general method behind these sums, see calculating AI ROI with a worked example.
A rough reading of totals: 4 or above is a strong engagement; 3 to 4 is acceptable with specific fixes; below 3 means the gap deserves a conversation, with the evidence in hand.
When the answer is "not enough value"
If the scorecard comes out low, keep it factual and give the consultant a chance to respond. A sequence that stays fair to both sides:
- Write down the gaps against the acceptance criteria, with your evidence and your own side's delays listed honestly.
- Ask for a remediation plan for anything inside the original scope: what they'll fix, by when, at no extra cost.
- Check the contract for acceptance terms, a support or warranty period, and whether the final payment depends on acceptance. The clauses to check in an AI consulting contract covers the ones that matter here.
- Decide whether to finish with someone else. If the relationship has broken down, make sure you hold the accounts, exports and documentation first; rescuing a stalled AI project or exiting it cleanly walks through that.
- Leave an honest reference. Other small businesses rely on them, and a balanced account (what went well, what didn't, how problems were handled) is more useful than stars.
In practice it can go like this. Picture a letting agency whose tenant-query chatbot scores 2.6: the build matched the proposal, but only about 30% of tenant questions go through it at day 90, and nobody recorded how long queries took before. Working through the steps, the agency admits it never gave staff time to point tenants to the chatbot, so part of the low adoption is its own. The consultant, for their part, agrees the bot was trained on an out-of-date tenancy handbook, fixes that within scope in a week, and adds a short script for staff to use on the phone. A month later adoption is near 60%. Not every low score ends that well, but a factual list on both sides makes a good ending far more likely.
Keep the scorecard either way. It becomes the brief for your next project: the acceptance criteria you wish you'd written, the baseline you'll measure first, and the questions you'll ask the next consultant before you hire.
Further reads
- Is Paying for AI Help Worth It? Your Time vs an Expert's Fee — Whether outside help makes sense before the next project.
- Is an AI Consulting Retainer Worth It After Your First Project? — Decide whether to keep the consultant on after a good result.
- What an AI Consultant Can't Do for You, and What You Must Own — The parts of success that were always yours to own.
- Who Owns the AI Workflows a Consultant Builds for You? — Check you own what you paid for.
- Should You Pay an AI Consultant on Results? Success Fees Explained — When paying on results would have changed the picture.
- How to Vet an AI Consultant's Case Studies and References — Judge the next consultant's past results before hiring.
- Is a 1:1 AI Consultation Worth It for a Small Business? — The real cost of a 1:1 AI consultation, how many hours it must save to break even, and three cafés that get three different answers.
- Do Small Medical Practices Need an AI Consultant? — When a small medical practice can handle AI itself, the signs outside help will pay for itself, what a consultant shouldn't do, and a scoped brief to adapt.
- How to Calculate the ROI of an AI Automation Before You Build It — A nine-step pre-build ROI method, a copyable worksheet, a roofing contractor's quote follow-ups costed, and a removals firm's idea that failed the test.
- How to Choose an AI Consultant: 20 Questions to Ask First — Twenty questions to put to any AI consultant, what strong and weak answers sound like, and a scoring sheet filled in for a farm shop.
- After an AI Consultation: Turn the Advice Into a 30-Day Plan — A one-page AI action plan template, a filled-in 30-day plan for a boutique hotel, the owner's weekly check and how to judge the result at day 30.
- Scope Creep in AI Projects: How to Handle Change Requests — A testable baseline, a one-page change request, sizing rules and a worked change log for keeping an AI project to its agreed scope.
- AI Tools and AI Development: The Complete 2026 Guide — the AI hub, including every tutorial in the AI-for-business series.