Your pilot worked if it hit targets written down before it started: one main number, such as minutes per task, beating your baseline by a set margin; a quality bar met on at least 50 real cases; no worsening in a guardrail such as complaints; and staff still using it in the final week. Criteria set afterwards don't count.
Criteria usually fail in one predictable way: they can't fail. "Save time" and "staff like it" will pass whatever happens. Criteria hold up when they're numbers, measured the same way before and during the pilot, fixed before launch, and judged by someone other than the person who championed the tool. The method below takes about 90 minutes to apply to a pilot you're about to run.
Step 1: write the one question the pilot must answer
A pilot is an experiment with one question. Write it as a sentence with a job, a change and a condition: "Does AI drafting cut the time to answer routine patient messages, without more mistakes reaching patients?" Everything else in your criteria exists to answer that question honestly.
One pilot runs through every step below, at an illustrative independent optician with six staff. Two people on the front desk answer routine messages: glasses ready for collection, appointment changes, contact lens reorders and questions about opening hours. The pilot gives them a business AI workspace holding the practice's standard wording, price list and policies, so they can draft replies and check them before sending. Clinical questions are excluded and go to an optometrist as before.
If you haven't yet planned how the pilot itself will run, running a first AI pilot project covers the charter, the set-up and the weekly review. The steps here deal only with the part that decides the verdict.
Step 2: take the threshold from your baseline, not a vendor's claim
Your main number needs a starting value and a target. The starting value comes from measuring the current process for two normal weeks; setting a baseline before you introduce AI covers how. The optician's baseline: 140 routine messages over two weeks, at a median of 6.5 minutes each, with the two weeks' medians at 6.1 and 6.9 minutes.
That spread matters. Week-to-week noise was about 0.8 minutes, so a pilot showing 6 minutes proves nothing. Set the target well outside the noise. The optician chose 4.5 minutes or less, including reading and correcting the draft: a cut of roughly 30%, big enough that hitting it couldn't be luck. A vendor's claim of "cut reply time by 70%" was ignored, because it described someone else's messages and someone else's staff.
Use the median rather than the average when a few items take far longer than the rest. One complicated complaint that takes 40 minutes shouldn't decide whether routine messages got faster.
Step 3: define "correct" before anyone grades
Quality criteria fall apart when "good enough" is decided case by case. Write a grading rubric before the pilot starts, so every draft is sorted into the same four boxes:
| Grade | Definition | Example at the optician |
|---|---|---|
| A | Sent as drafted | "Your glasses are ready to collect. We're open 9 to 5:30, Monday to Saturday." |
| B | Wording or tone edited only, under 30 seconds | Softened "you must" to "please" |
| C | A fact corrected: date, price, product or person | Draft said daily lenses; the patient wears monthlies |
| D | Rewritten or discarded | Draft answered a different question from the one asked |
Grading is quick once the boxes are defined: about 20 seconds a draft, logged as it's sent. The optician's log had one line per draft:
Date Type Grade Minutes Note
14 Oct Collection A 2.5
14 Oct Lens reorder C 6.0 guessed daily lenses
14 Oct Appointment move B 3.5 softened tone
15 Oct Opening hours A 1.5
To check the rubric is clear enough, have a second person grade a random ten drafts without seeing the first grades. If the two of you disagree on more than two of the ten, the definitions are too loose: agree what separates B from C for the cases you disagreed on, write it into the rubric, and do it before week 3, when grading starts to count.
The pass mark then becomes precise: at least 80% of graded drafts are A or B, and no C or D draft reaches a patient uncorrected. The first half measures usefulness; the second measures safety. Keep them separate, because a tool can be useful and unsafe at once.
Step 4: add guardrails that must not get worse
A guardrail is a number you're not trying to improve but can't allow to slip. Speed gains are worthless if they come with confused customers or tired staff. Choose one or two that would reveal the most likely harm:
- Clarification contacts: patients who call or reply to ask what a message meant. The optician's baseline was 6 per fortnight; the guardrail was no more than 8.
- Complaints about communication, from your complaints log.
- Time to first reply, if faster drafting might tempt staff to batch messages and reply later.
- Overtime or after-hours checking, if the review load might shift the work rather than remove it.
Pick guardrails you can count from records you already keep. A guardrail that needs a new survey won't get measured in week three.
Step 5: settle the sample size and the dates
Small samples make pilots look better or worse than they are. The arithmetic is simple: with 20 graded drafts, each mistake moves your error rate by 5 percentage points, so one unlucky day can swing the verdict. With 50, each mistake is 2 points. To tell a 90% success rate from a 95% one with any confidence, you want around 100 graded cases or more.
The optician expected about 70 routine messages a week, so four weeks gave roughly 280 drafts, and grading every draft in weeks 3 and 4 gave well over 100. Two more date rules protect the verdict:
- Judge the later weeks. Weeks 1 and 2 are learning time, and people often work differently when something is new and being watched. The optician's criteria applied to weeks 3 and 4 only.
- Compare like with like. If your baseline was a quiet month and the pilot runs in your busiest, speed may fall for reasons unrelated to AI. Note anything unusual in either period on the criteria sheet.
Step 6: agree the verdicts before the results arrive
Decide now what each combination of results will mean, so the decision meeting applies rules rather than debating feelings:
| Verdict | When it applies | What happens |
|---|---|---|
| Pass | Main number, quality bar, guardrails and usage all met | Roll out to the rest of the job, with the same checks |
| Extend | Main number met, one other criterion narrowly missed with a clear cause | Two more weeks with one named change; allowed once only |
| Stop | Main number missed by more than half the margin, or C/D drafts reached patients more than once | Switch off, write up what was learned |
| Redesign | Results good for one type of case, poor for another | Narrow the scope to the good type and run a short second pilot |
The "extend once only" rule is the one that saves you. Without it, a borderline pilot drifts on for months, costing staff time and subscriptions while never quite being judged.
Step 7: put it on one page, sign it and date it
The criteria only protect you if they were fixed before the results existed. Write them on one page, date it, and have the judge and the person running the pilot both sign it. The optician's:
AI PILOT SUCCESS CRITERIA Agreed: [date] Signed: [x2]
Question: does AI drafting cut the time to answer routine patient
messages without more mistakes reaching patients?
Scope: routine messages only (collections, appointment changes,
lens reorders, opening hours). Clinical questions excluded.
Main number: median minutes per message, including checking.
Baseline 6.5 (two weeks). Target 4.5 or less, weeks 3-4.
Quality: at least 80% of graded drafts A or B; no C or D draft
sent uncorrected. Grade every draft in weeks 3-4.
Guardrail: clarification calls or replies no more than 8 per
fortnight (baseline 6).
Usage: AI drafting used for at least 75% of eligible messages
in week 4, from the message log.
Verdicts: pass / extend once (2 weeks, one change) / stop /
redesign, as defined in the table.
Judge: practice owner. Runs pilot: senior receptionist.
Before signing, it's worth asking a chat assistant to attack the sheet. It's good at finding loopholes you're too close to see. A prompt that works:
Here are the success criteria for a four-week AI pilot in a small
optician's practice. [paste the sheet]
List every way this pilot could meet all the criteria while
actually failing, and every way it could miss them while actually
working. For each, suggest one change to the wording.
An illustrative reply, abridged:
"1. Usage is measured from the message log, which may count drafts generated rather than drafts sent; define usage as messages sent using an AI draft. 2. Staff choose which messages go through AI, so they could pick easy ones; compare the mix of message types with the baseline. 3. The median hides a long tail; also report the slowest 10%. 4. A sudden drop in message volume would make the guardrail easy to meet; express clarification contacts per 100 messages."
What you'd do with it: points 1, 2 and 4 are real gaps and went into the sheet. Point 3 was fair but unnecessary for routine messages, so it was noted and left out. The assistant also suggested adding a patient satisfaction survey, which was dropped for the reason given in step 4: nobody would run it in week three. Take the loopholes, leave the extra work.
Criteria that can't fail, rewritten so they can
Most weak criteria can be rescued by asking "what number, measured how, by when?" Five common ones, before and after:
- "Save staff time" becomes "median minutes per routine message, including checking, falls from 6.5 to 4.5 or below in weeks 3 and 4".
- "Staff like it" becomes "AI drafting is used for at least 75% of eligible messages in week 4, from the message log". Usage is what liking looks like when nobody is asking.
- "Better replies" becomes "at least 80% of graded drafts are A or B, and no C or D draft reaches a patient".
- "Customers are happy" becomes "clarification calls stay at or below 8 per fortnight, against a baseline of 6".
- "It's worth the money" becomes "hours saved in weeks 3 and 4, valued at the staff hourly cost, exceed the tool's monthly cost plus the hours spent maintaining it".
If you use a tool with its own reporting, borrow its evidence but not its definitions. The Microsoft 365 admin centre's Copilot usage report shows active users over 7, 28, 90 or 180 days, which is useful for your usage criterion. An automation's run history, such as Zapier's Zap history with its success, filtered and errored statuses, gives you volumes and failure counts. Some customer-service tools that bill per "resolution" count one when the customer simply goes quiet for a set period, which is not the same as the customer being satisfied.
How the optician's pilot was judged
Here are the illustrative results for weeks 3 and 4, against the signed sheet:
| Criterion | Target | Result | Met? |
|---|---|---|---|
| Median minutes per message | 4.5 or less | 4.2 | Yes |
| Drafts graded A or B | 80% or more | 88 of 104 (85%) | Yes |
| C or D drafts sent uncorrected | None | None (12 C and 4 D caught) | Yes |
| Clarification contacts per fortnight | 8 or fewer | 7 | Yes |
| Usage in week 4 | 75% or more | 85% | Yes |
A pass. But reading the graded drafts, not just the totals, showed that 10 of the 12 C grades were contact lens reorders, where the draft guessed the lens type. So the roll-out came with one change: lens reorders were taken out of AI drafting until the practice's lens list, by patient reference, could be looked up rather than guessed. Criteria tell you whether to proceed; the graded cases tell you how.
At 2.3 minutes saved on about 70 messages a week, the practice freed roughly 2.7 hours a week, worth about $60 a week in front-desk time at an illustrative $22 an hour, against about $50 a month for two business seats. Measuring time saved after an AI roll-out explains how to keep checking that the saving holds once the pilot's attention has moved on.
Reading results that come back mixed
Results often come back mixed rather than as a clean pass. Five patterns, and what each usually means:
- Faster, but quality slipped. Staff are trusting drafts without reading them. Tighten the check and look at which case types produce C grades before deciding.
- Quality fine, but no faster. Checking takes as long as writing did. Common for short messages; the job may be too small for drafting to help, and a template might do better.
- Good results, low usage. One enthusiast carried the pilot. Talk to the non-users before rolling out; their reasons are your roll-out plan.
- Poor early, improving steadily. The instructions were being refined as it ran. This is a legitimate reason for the single extension, if the trend is clear in the weekly numbers.
- Good for one case type, poor for another. The redesign verdict exists for this. An illustrative language school piloted AI suggestions for placement-test levels; teachers accepted almost all suggestions for adults and overrode about one in four for teenagers. It kept the adult tests and dropped the rest.
If you want an earlier warning than a four-week verdict, piloting AI in shadow mode lets you grade the AI's output against what staff actually did, before anything reaches a customer.
When the numbers pass but the pilot still failed
A pilot can hit every target and still be a poor decision. Before rolling out, check for these four quiet failures:
- The time moved rather than disappeared. The front desk got faster because the practice manager started checking drafts after closing time. Count the checker's time too.
- The easy cases were chosen. If staff decided which messages went through AI, they may have picked the simple ones. Compare the mix of case types with the baseline.
- People worked harder because they were measured. If week 4's figures were far better than anything in the baseline, not just the AI-assisted part, some of the gain may be attention rather than the tool. Re-measure a month after roll-out.
- The cost was never counted. Include subscriptions, set-up hours and the weekly upkeep of instructions and reference files in the verdict, not just the time saved.
Finally, schedule the follow-up now. A pilot proves the tool worked for four weeks with attention on it; reviewing an AI tool after 90 days tells you whether it still works when nobody is watching, which is the result that matters for the renewal date.
Judging an AI pilot fairly: follow-up questions
Who should judge whether the pilot passed?
Someone who didn't choose the tool and isn't paid by its vendor. In a small business that's often the owner judging a pilot a manager ran, or a manager judging the owner's pilot. The judge applies the criteria sheet as written. The person who ran the pilot presents the numbers and can argue for an extension, but shouldn't decide the verdict on their own enthusiasm.
Can I use the vendor's own success metrics?
Use them as extra information, not as your criteria. Vendor dashboards count what the product does, such as drafts generated, conversations handled or resolutions, and their definitions can be generous. A resolution may simply mean the customer went quiet. Your criteria should measure what changed in your business: minutes per task, errors that reached customers, complaints, and whether staff kept using it.
Can I change the criteria halfway through?
Only for a reason you'd have accepted before the start, such as discovering that your baseline measured something different, and only by writing down the change, the date and the reason. Never loosen a target because results are falling short. If the criteria were genuinely wrong, it's usually cleaner to stop, rewrite them and restart the clock than to patch them mid-pilot.
Should an outside supplier's final payment depend on the pilot passing?
It can, if the criteria are ones they can influence and you've both signed them before the pilot starts. Tie the payment to measures within their control, such as accuracy on an agreed test set or the system running reliably, rather than to outcomes that depend on your staff's adoption. Put the criteria and the grading method in the contract so there's nothing to argue about at the end.
Further reads
- AI KPIs for Small Businesses: 12 Metrics Worth Tracking — Twelve metrics to choose your main number and guardrails from.
- AI Proof of Concept vs Pilot vs Production: What Changes? — Know what a pilot should prove compared with the stages either side.
- How to Measure Customer Reaction After Introducing AI — Build a guardrail that captures how customers really respond.
- Why Your AI Pilot Stalled, and How to Get It Live — What to do when a pilot never reaches its decision date.
- How Much Does an AI Pilot Project Cost? — Budget the pilot so its cost is part of the verdict.
- How to Scope an AI Project: Deliverables and Acceptance Criteria — Turn pilot criteria into acceptance criteria for a full project.
- How Long Does AI Marketing Automation Take to Set Up and Pay Off? — Setup times and payback windows by automation type, a four-number payback formula and a delicatessen's first seven months, month by month.
- How to Measure Whether Copilot Is Paying for Itself — A seven-stage method for checking whether Microsoft 365 Copilot licences pay back, using Microsoft's own reports, a timed sample and one honest sum per person.
- How to Run a Two-Week AI Tool Trial Before You Commit — A day-by-day plan for trialling an AI tool on real work, with a test-case list, a daily log, a filled-in scorecard and the cancellation steps people forget.
- How to Roll Out Microsoft 365 Copilot to a Small Team — Start everyone on free Copilot Chat, license three heavy users for four weeks, then let the usage report decide who gets a seat. A café group shows how.
- How to Pilot Your First AI Agent Without Risking Customers — Run your first AI agent in shadow mode, then behind approvals, then live in a narrow window. A café's catering agent shows each stage, cost and stop rule.
- How Soon Should AI Pay for Itself? Payback Periods by Project — Payback targets for seven kinds of AI project, the sum with ramp-up months left in, and a toy shop testing four ideas against them.
- How to Write a One-Page AI Business Case for a Small Business — Five stages to a one-page AI business case, a finished example from a small charity, and a drafting prompt that stops AI inventing figures.
- Should You Pay an AI Consultant on Results? Success Fees Explained — Six fee structures priced at three outcomes, a toy shop that nearly paid a consultant for Christmas, and the metrics that invite gaming.
- How to Write a Request for Proposal for an AI Project — Eight sections every AI project RFP needs, a members' club's RFP filled in, a pricing table that makes replies comparable, and a scoring matrix.
- After an AI Consultation: Turn the Advice Into a 30-Day Plan — A one-page AI action plan template, a filled-in 30-day plan for a boutique hotel, the owner's weekly check and how to judge the result at day 30.
- AI Tools and AI Development: The Complete 2026 Guide — the AI hub, including every tutorial in the AI-for-business series.
Sources: Microsoft 365 admin centre documentation (Copilot usage report); Zapier help centre (Zap history run statuses); Intercom and Zendesk pricing documentation (how resolutions are counted).