Your pilot has probably stalled for one of five reasons: nobody owns it, there's no agreed pass mark, it sits outside the daily workflow, a visible mistake broke trust, or nobody planned the steps from trial to everyday use. Find the one that applies, fix only that, and set a go-live date with a pass mark attached.
The fix is rarely a better tool. Many stalled pilots worked well enough on the examples someone tried; what's missing is the unglamorous scaffolding around them. Below is a one-hour diagnosis, the fix for each cause, a worked restart for a kitchen fitting firm, a 30-day plan, and the criteria for stopping cleanly if the pilot deserves to end.
Diagnose it in an hour
Get the two or three people involved in a room, or on a call, and go through these questions honestly. The first "no" is usually the cause; if several are "no", fix them in this order. Bring whatever numbers exist, even rough ones: how often the AI was used each week since the start, and any examples of output people rejected. Usage that fell steadily points to workflow; usage that fell off a cliff points to a specific incident.
Two illustrative weekly usage records show the difference. A pilot logging 22, 19, 15, 12, 9 and 7 uses over six weeks is leaking a few people a week, typically because the AI route is slightly more effort than the old one and loses a little more ground on every busy day. A pilot logging 22, 24, 21, 23, 6 and 4 was fine until week five, so ask what happened in week four. Usually someone can name it: a wrong figure sent to a customer, an embarrassing draft read aloud in the office, a login that stopped working.
| Question | If the answer is no |
|---|---|
| Can everyone name the one person who owns this pilot, and do they have time set aside for it? | Cause 1: no owner |
| Is there a written pass mark that says what "good enough to go live" means? | Cause 2: no pass mark |
| Does using the AI take fewer steps than the old way, inside tools people already open? | Cause 3: outside the workflow |
| Do the people using it trust its output as much as they did in week one? | Cause 4: broken trust |
| Is there a list of what must happen before it's used by everyone, every day? | Cause 5: no route to live |
| On your test examples, does the AI actually reach the quality you need? | Less common: the task doesn't suit AI, or needs a different design |
| At full volume, does it still cost less than the time it saves? | Less common: the economics don't work |
The last two are worth asking because they're the only causes where the right answer might be to stop. The first five are all fixable, usually in weeks rather than months.
The five causes and how to fix each
1. Nobody owns it
Symptoms: the pilot comes up in conversation ("how's the AI thing going?") but nobody can say what happened last week. Questions go unanswered. Problems get mentioned but not fixed.
Fix: name one owner, ideally the person whose work the pilot changes most, and give them a protected hour a week. Their job is small but specific: collect problems, fix the instructions or escalate, track the numbers against the pass mark, and report once a month. An owner with no time isn't an owner.
The monthly report can be four lines. From an illustrative events caterer's pilot, where AI drafted menu proposals from enquiry forms: "Used on 18 of 21 enquiries. 15 proposals sent with light edits. One error caught: a dish described as nut-free when its sauce contains almonds. Changed the instructions so the AI never states allergen status; the office adds it from the allergen chart." That's enough for the business owner to see the pilot is moving, and to notice straight away if a month's report doesn't arrive.
2. There's no pass mark
Symptoms: opinions differ on whether it's "working". Someone always wants to try one more thing. The pilot has quietly passed its planned end date.
Fix: write the pass mark now, with numbers and a period. For example: "at least 8 in 10 outputs usable with light edits, average time per task under 5 minutes including checking, and no errors reaching customers, for three consecutive weeks." Then agree the decision date. Setting success criteria that hold up has more examples and the traps to avoid.
Most stalled pilots have a goal rather than a pass mark. Two more rewrites, both illustrative:
| Pilot | What was written at the start | A pass mark you can decide on |
|---|---|---|
| Bookkeeping practice: AI categorising receipts | "See whether it saves time on receipts" | At least 95 of a 100-receipt weekly sample in the right category, and under 30 seconds of checking per receipt, for three weeks running |
| Garden centre: AI product descriptions | "Get the website descriptions done faster" | 40 descriptions a week published, each with under 3 minutes of editing, and no wrong hardiness or size claims in a weekly spot check of 10 |
The second column is why these pilots drifted: "saves time" and "faster" are true after almost any trial, so they never force a decision. The third column can fail, which is exactly what makes it useful.
3. It lives outside the daily workflow
Symptoms: usage started high and drifted down. People say they "forget". Using the AI means opening another app, logging in again, or copying and pasting between systems.
Fix: count the steps. If the AI route takes more clicks than the old way, it will lose on a busy day, however good the output. Move the AI into a tool people already open, make it the default route, and where possible retire the old one so there's only one way to do the job. This is the most common fix and the least glamorous.
Counted out for an illustrative property management firm's pilot, where AI summarised tenants' maintenance emails for the job system:
OLD WAY (3 steps) PILOT ROUTE (7 steps)
1. Read email 1. Read email
2. Type summary into 2. Open the chat assistant in a browser
job system 3. Sign in (it timed out daily)
3. Assign contractor 4. Copy the email, paste, add the prompt
5. Copy the summary
6. Paste into the job system
7. Assign contractor
The summaries were good; the route was worse. The restart used the assistant built into the firm's email, which added a summary button to the message itself, bringing the route back to four steps. Usage recovered within a fortnight, with no change to the prompt at all.
4. One visible mistake broke trust
Symptoms: usage fell sharply after a specific incident. People now check every output so thoroughly that the time saving has gone, or have stopped using it without saying so.
Fix: treat the mistake properly rather than hoping it's forgotten. Find exactly what went wrong, add a rule or a check that prevents that kind of error, re-run your test examples to confirm, and tell the team what changed. People forgive a tool that visibly learns; they don't forgive one that might do it again. The "tell the team" part can be a short message. From an illustrative car servicing garage whose AI-drafted reminders quoted an old price:
Subject: Service reminders: what went wrong and what's changed
What happened: last Tuesday's reminders quoted $149 for a standard
service. It's been $169 since March. Four customers queried it and
we honoured $149 for them.
Why: an old price list was still saved in the assistant's project,
and it used that one.
What's changed: the old list is deleted. There's now one price
file with a "last updated" date, and the prompt says to use only
that file.
What we checked: re-ran 20 of last month's reminders; all show
$169.
What I'm asking: keep using it, and tell me the same day if a
reminder looks wrong, even if you've fixed it yourself.
Checking AI is doing good work, not just fast work covers building quality checks that catch these before customers do.
5. Nobody planned the route from pilot to live
Symptoms: the pilot "passed", but it's still one or two people using it on their own accounts. Everyone agrees it should be rolled out; nothing happens.
Fix: write the go-live list. For most small-business pilots it looks like this:
- Company accounts or seats for everyone who'll use it, with personal accounts retired.
- Access to the files, inboxes or systems it needs, and no more.
- A one-page playbook: the prompt or set-up, what to check, what never to put in.
- A short training slot for the people who weren't in the pilot.
- A monitoring routine: who looks at what, weekly, for the first two months.
- A fallback: what happens if the tool is down or starts misbehaving.
- The monthly cost at full volume, confirmed against the budget.
What changes from proof of concept to pilot to production explains why each item matters once more people depend on it.
Worked example: a kitchen fitter's job-update pilot
Take an illustrative kitchen fitting firm: the owner, an office manager, a designer and five fitters. In spring it piloted a simple idea: at the end of each fitting day, the fitter records a one-minute voice note; AI turns it into a short progress update for the customer and a snag list for the office; the office manager checks and sends the update.
For three weeks it went well. Around eight in ten fitting days produced an update, customers liked them, and the office manager spent about four minutes per update. By week six, updates were arriving for roughly three days in ten, and by week nine nobody mentioned it. The one-hour diagnosis found three causes at once:
- No owner. The owner assumed the office manager was running it; the office manager assumed it was the owner's project.
- Outside the workflow. Fitters had to open a separate recording app and log in at the end of a long day. They already opened the firm's job app to clock off; the voice note wasn't there.
- Broken trust. In week four an update told a customer their worktops would be fitted "tomorrow", repeating the fitter's hopeful comment, when templating had slipped. The customer took a day off work for nothing and complained. After that, the office manager rewrote every update from scratch.
The restart fixed each one. The office manager became the owner, with an hour a week. The voice note moved into the job app's notes field, recorded as part of clocking off. A rule was added to the AI's instructions: never state dates or times; the office adds those from the schedule. And the pass mark was written down: updates sent for at least 80% of fitting days by 6pm, under four minutes of office time each, and no date errors, for four weeks running. Four weeks after the restart, updates were going out on about 85% of fitting days at around three minutes each, and one date slip was caught by the office before sending. They went live, and the go-live list covered the fitters who'd joined since the pilot.
A 30-day restart plan
- Days 1 to 3. Run the diagnosis. Name the owner. Write the pass mark and the decision date (day 30).
- Days 4 to 10. Fix the cause you found: move it into the workflow, add the missing check, or give the owner their hour. Re-run your test examples to confirm nothing else broke. A pre-mortem on the restarted version takes 30 minutes and often spots the next failure early.
- Days 11 to 25. Use it for real, every time the task comes up. The owner tracks the pass-mark numbers weekly and fixes problems within the week.
- Days 26 to 30. Decide: go live (work through the go-live list), extend once by up to four weeks with a specific reason, or stop.
A specific reason names what the extra weeks will show. "Our supplier's price file only changes at the start of each month, so we need one more cycle to see whether the AI picks up new prices" is specific: at the end you'll know. "People need more time to get used to it" isn't, and it's the sentence most pilots use just before they stall a second time.
If you'd rather go through the diagnosis with someone who hasn't been close to the pilot, that's the kind of thing my AI implementation consultation covers. Either way, the one rule is not to extend twice: a pilot that needs a third attempt at the same pass mark is telling you something.
Signs the pilot should close rather than restart
Stop, rather than restart, if any of these is true: the AI can't reach the pass mark on your test examples even with good instructions; the task turned out to happen far less often than you thought; the cost at full volume exceeds the value of the time saved; or the underlying process has changed so much that you'd be piloting something new. If an outside developer or agency built the pilot, rescuing a stalled AI project or exiting it cleanly covers the contractual side.
The first of those is testable in an afternoon, using past jobs where you already know the right answer. An illustrative tiling firm piloted AI to estimate tile quantities from customers' room photos. It picked 20 finished jobs, fed in the original photos, and compared the AI's estimates with the quantities actually used: 11 of 20 were within 10%, and three were out by more than a third because the photos gave no reliable sense of scale. Better instructions moved it to 12 of 20. With a pass mark of 18, that's a task that needs a measurement from a person, not a better prompt, and the pilot closed.
The cost test catches pilots whose bill looked trivial at trial volume. Suppose a web shop trialled a support assistant charged at $0.99 per resolved outcome, the rate Intercom lists for its Fin agent. At 80 resolutions a month the bill was about $79, which nobody questioned. But most of those conversations were order-status questions that took a person about 90 seconds; at a staff cost of $20 an hour, each one saved roughly 50 cents of time while costing 99 cents. At the full volume of 1,200 a month, that's about $1,190 spent to save about $600. Work out the per-conversation figures before you scale anything priced per use.
When you stop, write a five-line close-out: what you tried, what the numbers showed, why you stopped, what you'd do differently, and what you'll try next. It takes ten minutes and stops the same idea being revived in a year without the lessons. Here's one from an illustrative picture-framing shop:
TRIED: AI chat assistant on the website for framing enquiries,
eight weeks.
NUMBERS: 31 chats in total, 4 turned into orders. 19 needed a
person anyway (sizes, mount colours, "can you see my
photo?").
WHY STOP: Too few chats, and most needed judgement it couldn't
give. The $39 a month cost more than the orders it won.
DIFFERENTLY: Count the enquiries first. We'd have seen the volume
wasn't there.
NEXT: AI-drafted replies to email enquiries, sent by us.
Then pick your next candidate, and use how to run your first AI pilot project to set it up with an owner, a pass mark and a go-live list from day one.
More questions about stuck pilots
How long should an AI pilot run before it goes live?
For a single task in a small business, four to eight weeks of real use is usually enough, provided you set the pass mark at the start. Shorter than four weeks and you haven't seen enough awkward cases; longer than about three months without a decision and the pilot has probably stalled rather than being thorough. Put the decision date in the diary on day one.
Would switching to a different AI tool fix it?
Occasionally, but check the other causes first. Changing tools resets the learning, the prompts and the habits, and if the real problem was ownership or workflow, the new tool stalls the same way. Switch only if your test examples show the current tool can't do the task to the pass mark even with good instructions, and another tool can on the same examples.
Who should own a pilot in a team of five or ten?
The person whose work it changes most, with an hour or so a week set aside for it, and the owner's backing to change how things are done. Owners often keep pilots for themselves and then have no time for them. A named operational owner plus a monthly ten-minute check-in with the business owner is enough governance for a small team.
Further reads
- How to Pilot AI in Shadow Mode Before Customers See It — Restart safely with AI running alongside, not instead.
- Why AI Projects Fail in Small Businesses (It's Rarely the Tech) — The wider reasons AI projects fail, rarely the tech.
- How to Scale AI From One Workflow to the Whole Business — What to do once this pilot is finally live.
- How to Get Your Staff to Actually Use AI Tools — Fix a pilot that stalled because people stopped using it.
- How to Review an AI Tool After 90 Days: Keep, Fix or Cancel — The keep, fix or cancel review for live tools.
- How to Choose and Support an AI Champion in a Small Team — Choose the owner your pilot was missing.
- Do I Need an AI Consultant? 8 Signs It's Time to Get Help — Eight signs a small business needs outside AI help, each with a test you can run this week, plus the signs that it doesn't need a consultant yet.
- AI Implementation Plan for a Small Law Firm: The First 90 Days — A day-by-day 90-day AI plan for a six-person law firm: policy first, two scored pilots, shadow testing, an error log and a go/no-go decision.
- Why AI Adoption Stalls in Accounting Firms, and How to Restart It — The five reasons AI stalls in accounting practices, a twenty-minute diagnosis, and a 30-day restart timed around the deadline calendar.
- How to Run a Two-Week AI Tool Trial Before You Commit — A day-by-day plan for trialling an AI tool on real work, with a test-case list, a daily log, a filled-in scorecard and the cancellation steps people forget.
- How to Pilot Your First AI Agent Without Risking Customers — Run your first AI agent in shadow mode, then behind approvals, then live in a narrow window. A café's catering agent shows each stage, cost and stop rule.
- How Long Does AI Implementation Take for a Small Business? — Realistic planning ranges for small-business AI projects, where the weeks really go, and a week-by-week timeline for a sports equipment shop.
- How Much Does an AI Pilot Project Cost? — What a small-business AI pilot really costs, with a catering company's six-week pilot costed, three budget levels and the commitments that inflate it.
- After an AI Consultation: Turn the Advice Into a 30-Day Plan — A one-page AI action plan template, a filled-in 30-day plan for a boutique hotel, the owner's weekly check and how to judge the result at day 30.
- Scope Creep in AI Projects: How to Handle Change Requests — A testable baseline, a one-page change request, sizing rules and a worked change log for keeping an AI project to its agreed scope.
- 10 AI Implementation Mistakes Small Businesses Make — Ten setup, billing and upkeep mistakes that turn a sensible AI project into a year-long cost, each with how it shows up and the fix.
- AI Tools and AI Development: The Complete 2026 Guide — the AI hub, including every tutorial in the AI-for-business series.