How to Run Your First AI Pilot Project in a Small Business

Coding Liquids tutorial cover featuring Sagnik Bhattacharya for How to Run Your First AI Pilot Project in a Small Business.
Coding Liquids tutorial cover featuring Sagnik Bhattacharya for How to Run Your First AI Pilot Project in a Small Business.

Run your first AI pilot on one workflow, with one tool and two or three people, for four to six weeks. Set success and stop criteria before you start, measure the current process for two weeks, review results for 15 minutes each week, then decide to keep, fix or stop. Budget $50 to $100 a month in tools plus a few staff hours.

The point of a pilot isn't to prove AI works. It's to produce a clear decision about one job, based on numbers from your own business. Write down the result that would make you stop before week one, because after a month of effort everyone in the room will want the verdict to be "keep".

Follow me on Instagram@sagnikteaches

The worked example throughout is an illustrative optician practice with two optometrists, a practice manager and three dispensing and reception staff. About 60 non-clinical emails a week reach the shared inbox: "are my glasses ready?", "do you have appointments on Saturday?", contact lens reorders, frame repairs and lens price questions. Replying takes the team around six hours a week.

Connect on LinkedInSagnik Bhattacharya

Stage 1: Choose the workflow and write a charter (1 to 2 hours)

Pick one job that happens daily or weekly, is mostly text, and can be checked by a person before anything reaches a customer. Then write a one-page charter. It fixes the scope, which is the single best defence against a pilot that drifts into "let's see what AI can do".

Subscribe on YouTube@codingliquids
PILOT CHARTER
Name:            AI-drafted replies to non-clinical email enquiries
The job:         Staff paste a non-clinical enquiry; AI drafts a reply;
                 staff check, edit and send. Nothing is sent automatically.
Out of scope:    Anything about symptoms, prescriptions, eye health or
                 complaints. These go to an optometrist or the manager.
People:          Two dispensing staff (weeks 1-2), plus reception (weeks 3-6)
Tool:            Business AI plan, one shared project holding instructions,
                 price list, lens turnaround times and opening hours
Data rule:       No dates of birth, prescriptions or clinical notes pasted in
Dates:           Baseline 2-13 June; pilot 16 June - 25 July
Decision date:   28 July, 1-hour meeting
Owner:           Practice manager
Success / stop:  See criteria table

The out-of-scope line matters more in an optician's than in most businesses. Health information is treated as a special, more protected kind of personal data under data-protection law such as the GDPR, and a wrong answer to a symptom question is a safety issue. Keeping clinical content out of the pilot entirely removes the hardest questions. If your pilot would touch health, financial or children's data, check the plan with your data-protection adviser first.

Stage 2: Set success and stop criteria (30 minutes)

Decide now what result would make you keep it, and what would make you stop. Writing these before the pilot starts stops everyone from reading the results generously later.

MeasureBaseline (to be measured)Success at week 6Stop straight away if
Minutes per reply, including checkingAbout 63.5 or lessHigher than baseline at week 3
Drafts sent with no or minor editsn/a60% or moreBelow 30% at week 3
Wrong information reaching a patientRare; log anyNone seriousAny AI-drafted reply to a clinical question is sent
Hours to first replyAbout 20 working hours8 or lessn/a

Keep it to three or four measures, with at least one about quality. For more on choosing thresholds that survive scrutiny, see setting success criteria for an AI pilot.

Stage 3: Measure the current process for two weeks (10 minutes a day)

Without a baseline, "it feels faster" is all you'll have at the decision meeting. For two normal weeks, staff note the time they start and finish ten replies a day and tally the total emails. Email timestamps give you time to first reply without any extra work. Our guide to setting a baseline before introducing AI covers sampling and what to do about unusual weeks.

In the optician's baseline, the team handled 118 non-clinical emails in two weeks. The sampled replies averaged 6.1 minutes each, and the median time to first reply was just under 20 working hours, because the inbox was cleared in batches twice a day.

The log itself can be a shared spreadsheet with five columns. A few rows from one dispensing assistant's Tuesday might read:

Email typeStartedSentMinutesNote
Glasses ready?09:1209:153Checked lab tracker
Lens price question09:1509:2611Had to ask manager about thinner-lens upgrade price
Contact lens reorder11:4011:455
Frame repair14:0214:097Interrupted by a walk-in

The note column earns its place. The 11-minute price question shows the real cost wasn't typing but finding the answer, which tells you the price list must be in the AI's files. Interruptions stay in the figures, because they'll happen during the pilot too; just keep logging the same way.

Stage 4: Set up the tool with guardrails (half a day)

Use a business plan, not personal accounts. ChatGPT Business and Claude Team both list at $25 a seat a month billed monthly, with a minimum of two seats, and neither trains on business content by default. Both let you create a shared project that holds standing instructions and reference files, so every member of staff drafts from the same price list and rules.

The instructions are where most of the quality comes from. A starting version:

You draft email replies for an optician practice. A member of staff
will check and send every reply.

Use ONLY the price list, lens turnaround table and opening hours in
this project's files. If the answer isn't in those files, write
[STAFF TO CHECK: ...] instead of guessing.

If the email mentions symptoms, eye pain, sudden changes in vision,
headaches, a prescription query or a complaint, do not draft a reply.
Write only: "CLINICAL OR COMPLAINT - pass to optometrist/manager."

Tone: warm, brief, plain English. Sign off with "The team at [practice]".
Never promise a date unless it is in the turnaround table.

Test it on 20 past emails before the pilot starts, including five awkward ones. If it drafts a reply to anything clinical, tighten the instruction before going further.

The awkward emails are usually mixed ones. One from the optician's test set read: "Hi, my new glasses keep giving me a headache by lunchtime. Also, is the second-pair offer still on? I'd like prescription sunglasses." The first version of the instructions produced this (illustrative):

"Thanks for getting in touch! Yes, our second-pair offer is still running until the end of the month, and prescription sunglasses are included. Just pop in or reply with a time that suits and we'll get you booked in. The team at [practice]"

It's friendly, accurate on the offer, and dangerous: the headache was ignored because the email was mostly about a sale. The fix was one sentence in the instructions: "If ANY part of the email mentions a symptom, treat the whole email as clinical." Re-run the full 20 after any change like this, not just the email that failed, because a tighter rule can start catching harmless emails too. Here it did: "my frames give me a headache to look at, they're so dated" was flagged as clinical. That's an acceptable error in the safe direction.

Stage 5: Run it for four to six weeks with a weekly review

Before day one, spend 15 minutes briefing the people involved. Most pilot problems start with a misunderstanding about what staff are allowed to do, so cover the same points every time:

KICK-OFF BRIEFING (15 minutes)
- What we're testing: AI drafts for non-clinical emails only
- What we're NOT testing: anyone's speed or performance
- You are responsible for every reply you send, as now
- If a draft is wrong, fix it and log it; bad drafts are useful data
- Never paste in: dates of birth, prescriptions, clinical notes
- Clinical or complaint emails: don't use the tool, pass them on
- Log: time started, time sent, edit level, one line if it went wrong
- Questions or worries: the practice manager, any time

Agree what each edit level means with one real draft in front of everyone, or two people will log the same draft differently. Take a reply to a frame repair enquiry. The AI's draft said: "We can repair most frames in store. Please bring them in and we'll take a look; repairs usually take 3-5 working days." The member of staff changed "3-5 working days" to "about a week, as we send soldering work to our lab" and sent it. That's a minor edit: one fact corrected, structure kept. A heavy edit is rewriting more than half the draft; discarded is starting again from a blank reply. Log the minor edit's reason too, because "repair turnaround wrong" appearing four times in a week is next week's file change.

The second line of the briefing deserves emphasis. If staff think the pilot is really measuring them, they'll quietly avoid the tool on hard emails or work faster than usual while being timed, and both distort the results.

Keep the pilot small at the start and widen it only when the numbers support it.

WeekWhat happensWhat you log
1Two staff use it for every non-clinical emailMinutes per reply, edit level (none, minor, heavy, discarded)
2First instruction changes, based on the worst draftsSame, plus what was changed and why
3Reception joins; checkpoint against the stop criteriaSame
4Normal running; no new changes unless something breaksSame
5-6Steady state, measured like the baselineTen timed replies a day, total volume, time to first reply

The weekly review takes 15 minutes and follows the same four items every time: this week's numbers against the baseline; the three worst drafts and why they went wrong; one change to the instructions or files; anything that nearly went wrong. Freezing changes in weeks 4 to 6 matters, because you can't measure a setup that keeps moving.

In the optician's pilot, the worst drafts in week 1 shared a cause. The AI told a patient their varifocals would be ready "within 7 days", taking the figure from a general line in the price list, when varifocals from the lab take longer. Adding a turnaround table by lens type fixed it, and that class of error didn't recur. If you'd rather see the AI's output for a while before staff act on it at all, piloting in shadow mode is a more cautious alternative for this stage.

Stage 6: The decision meeting (1 hour)

Put the baseline and week 5-6 figures side by side, then apply the rules you wrote in stage 2:

  • Keep if the success criteria are met and no stop criterion was triggered. Decide who owns it from now on and when you'll next review it.
  • Fix if there's a clear time saving but quality falls short, and you can name the specific change that would fix it. Allow one extension of up to four weeks, once.
  • Stop if a stop criterion was hit or there's no measurable saving. Write down what you learnt, cancel the seats, and pick the next candidate.

The optician's week 5-6 figures came out at 3.2 minutes per reply including checking, 68% of drafts sent with no or minor edits, no serious errors, and a median time to first reply of about 6 working hours, because drafting was now quick enough to do between patients. On about 60 emails a week, that's roughly 2.9 hours saved weekly. The decision was keep, with the practice manager as owner and a review in three months.

"Fix" is the decision that needs the most discipline, so it helps to see one. Picture an illustrative four-person plumbing and heating firm that piloted AI drafts of quote follow-up emails. At week six, time per follow-up had fallen from 9 minutes to 4, comfortably past the target, but only 41% of drafts went out with no or minor edits against a 60% target. The weekly logs showed why: most heavy edits were the plumber rewriting the parts and labour lines, because the AI had no parts price list and was estimating. That's a named, specific cause, so the decision was fix: load the parts list into the project, extend by four weeks, same criteria. If the extension had ended at 50%, the rule written in stage 2 would have made it a stop, however close it looked.

Whatever you decide, write a half-page pilot note and file it with the charter: the figures, the decision, the three biggest problems and how they were fixed, and what you'd do differently. It takes 20 minutes and it's the most valuable thing the pilot produces after the decision itself, because your second pilot can start from it instead of from scratch. It also protects you from the most common afterthought, six months on, of nobody remembering why the tool was set up the way it was.

The optician's note, filled in, might run to these five points:

  • Result: 6.1 to 3.2 minutes per reply; 68% sent with no or minor edits; first reply 20 to 6 working hours; no serious errors. Decision: keep.
  • Biggest problems: varifocal turnaround (fixed with a lens-type table), mixed emails hiding symptoms (fixed with the "any part" rule), repair times (fixed by adding the lab's turnaround to the files).
  • What surprised us: the time to first reply improved more than minutes per reply, because drafting fitted between appointments.
  • Do differently next time: include mixed emails in the first 20 test cases; agree edit levels on day one, not in week two.
  • Owner and review: practice manager; next review in October; files to update whenever prices change.

What a first pilot costs in money and hours

Using the optician's numbers as an illustration:

  • Tools: two business seats at $25 a month for two months is $100 at list price, with no annual commitment.
  • Setup: about six hours of the practice manager's time for the charter, instructions, files and testing.
  • Baseline: about ten minutes a day per person for two weeks.
  • Reviews: six 15-minute meetings for three people, about 4.5 staff hours.
  • The learning dip: the first week is usually slower than normal while people get used to it. Plan for it.

Add those hours up and set them against the saving. Setup (6 hours), baseline logging (ten minutes a day for three people over ten working days, about 5 hours) and reviews (4.5 hours) come to roughly 15.5 staff hours plus the $100 in seats. At 2.9 hours saved a week, the staff time is paid back after roughly five weeks of the new routine, and after that the running cost is the two seats. If your own sum shows a payback of a year or more, the workflow was probably too small to pilot on its own; pair it with a second job that uses the same tool rather than abandoning the idea.

Pilots that need a developer or an outside consultant, or that connect several systems, cost considerably more. How much an AI pilot project costs breaks down those larger budgets.

Five ways first pilots fail without anyone noticing

  • Scope creep. "While we're at it, could it write the rota?" Every addition resets the clock. Note ideas for later; don't add them.
  • No end date. A pilot without a decision meeting becomes a permanent half-used tool. Book the meeting when you write the charter.
  • Rubber-stamp checking. By week four, reviewers may approve drafts without reading them. If the edit rate drops to zero overnight, sit with a reviewer for ten minutes and watch.
  • Only easy cases. Staff sometimes skip the tool for awkward emails, which flatters the numbers. Log every eligible email, including those done by hand, and ask why.
  • The champion leaves. If one person holds all the know-how, a holiday stalls everything. The instructions and files in the shared project are your insurance.

If your pilot has already stalled, why AI pilots stall and how to get them live works through recovery. And if you're unsure whether you need a pilot or a smaller test first, proof of concept versus pilot versus production explains what changes at each step.

Questions about running a first AI pilot

How many people should take part in a first pilot?

Two or three is usually right. One person can't show whether the setup works for anyone other than its builder, and more than three makes the weekly review slow and the results harder to compare. Choose people who do the job every day, including one sceptic, because their objections expose problems an enthusiast will work around without mentioning.

Do customers need to know about the pilot?

If a member of staff reads, edits and sends every reply, the message is theirs and many firms don't mention the drafting tool. If AI output reaches customers without a person checking it, or if a customer might reasonably think they're talking to a person when they aren't, tell them. Check any sector rules that apply to you, and take advice if you're unsure.

Can I run the pilot on a free AI plan?

You can test ideas on a free plan with made-up examples, but don't pilot with real customer messages there. Free and personal plans have different data terms from business plans, and they often lack shared projects and admin controls. Two business seats for a six-week pilot cost about $100 at list price on monthly billing, which is a small price for keeping the pilot clean.

Further reads

Sources: ChatGPT Business and Claude Team pricing pages; OpenAI and Anthropic help centres on Projects. Checked September 2026.

Planning your first AI pilot?

On a 1:1 call we'll pick the workflow, write the charter and success criteria together, and set up the tool and review routine so the pilot ends in a clear decision.

Book a 1:1 call with me