Why AI Projects Fail in Small Businesses (It's Rarely the Tech)

Coding Liquids tutorial cover featuring Sagnik Bhattacharya for Why AI Projects Fail in Small Businesses (It's Rarely the Tech).
Coding Liquids tutorial cover featuring Sagnik Bhattacharya for Why AI Projects Fail in Small Businesses (It's Rarely the Tech).

Most small-business AI projects fail for organisational reasons, not technical ones: nobody owns the project, the goal is vague and was never measured, staff aren't given time to change how they work, the awkward cases were never planned for, and nobody maintains it after launch. The tools usually do roughly what they promised.

That's good news, because those causes are cheaper to fix than technology. Most of them show symptoms by week three if you know what to look for, and an eight-question check before you start prevents the worst. Sometimes, too, stopping a project is the right decision rather than a defeat.

Follow me on Instagram@sagnikteaches

How a typical failure unfolds

Failed projects rarely have a single cause. They follow a chain, and each link makes the next more likely:

Connect on LinkedInSagnik Bhattacharya
  1. The goal is set as "use AI to save time on admin", which nobody can measure.
  2. Because the goal is vague, nobody records how long the task takes now.
  3. The enthusiast who suggested it builds it in spare moments, which makes them the owner by default and without time.
  4. It launches in a busy month. Staff try it when quiet and drop it when pressed, which is exactly when it would help most.
  5. The first odd cases come in. It handles them badly, and there's no plan for what should happen instead.
  6. People quietly go back to the old way. The workflow keeps running, unused.
  7. Six months later, the verdict is "we tried AI and it didn't work for us".

Notice that nothing in the chain is about the model's quality. Break any two of the links early and the project usually survives.

Subscribe on YouTube@codingliquids

Seven non-technical causes, and what each looks like in week three

1. No owner with real time

The person responsible is also the busiest person in the business. By week three: the check-in meeting has been moved twice, and questions from staff go unanswered for days. Fix it by naming an owner and putting two hours a week in their diary, protected like a client meeting.

2. A goal that can't be measured

Without a before-figure, nobody can say whether it's working, so discussions turn into opinions. By week three: one person says it's brilliant and another says it's useless, and there's no way to settle it. Pick one number (time per task, turnaround, errors found) and record it for two weeks first. Setting a baseline before you introduce AI shows how to do this without a stopwatch culture.

A before-and-after from an illustrative small accountancy practice shows the difference. The original goal: "Use AI to help with client onboarding." The rewrite: "Cut the time to prepare a new-client welcome pack from about 75 minutes to 45, measured on the next ten new clients, with no more corrections at partner review than we get now." The second version names the task, the number, the sample and a quality guard, so the week-eight conversation is about a figure rather than who feels more strongly.

3. The wrong first job

The task was too rare to be worth automating, too risky to get wrong, or depended on judgement the AI can't supply. By week three: nearly every output needs heavy rewriting. Sometimes the job needed plain rules rather than AI at all; AI versus rule-based automation helps tell the two apart.

Rarity is the easiest of the three to check with a sum. An illustrative surveying practice spent about ten hours setting up an AI workflow for its annual insurance renewal pack, a job that takes three hours by hand once a year. Even if the workflow cut it to one hour, the set-up would take five years to pay back, and the instructions would be out of date by the second renewal. The same ten hours spent on turning site notes into report sections, about 40 minutes a time and done four times a week, would have paid back in about two months even if AI only halved the time.

4. A process nobody had agreed on

Three people did the task three different ways, and the AI was built around one of them. By week three: two people use it differently and one doesn't use it. This is the problem of automating a broken process, and the fix comes before any tool.

5. No time allowed to change habits

Using a new workflow is slower for the first fortnight. If nobody's workload allows for that, the old way wins every busy day. By week three: usage is high on Mondays and near zero by Thursday. Launch in a quieter period, and tell people explicitly that the first two weeks will be slower.

The dip is easy to underestimate, so put numbers on it. At an illustrative estate agency, writing a property listing the old way took about 15 minutes. With the new AI workflow, the negotiators' own timings came out at roughly 22 minutes a listing in week one (learning the template, checking every line), 16 in week two and 10 by week three. Anyone who judged it on week one saw a tool that made the job slower and went back to the old way. Anyone told in advance "expect week one to be slower, and we'll compare in week three" got to the 10-minute version. The only difference was the sentence at launch.

6. The awkward cases were left for later

The workflow handles the typical case and mangles the rest. By week three: there's a growing list of "it doesn't handle X", and people have built workarounds in their own spreadsheets. List the five most awkward cases before building, and decide what should happen to each, even if the answer is "send it to a person".

Here is that list filled in for an illustrative online homeware shop about to let AI draft replies to returns emails. It took the owner and one colleague 20 minutes, working from the last two months of the returns inbox:

Awkward caseHow often (last 2 months)What should happen
Return requested after the 30-day window9AI drafts a polite refusal quoting the policy; a person approves before it sends
Item arrived broken, with a photo6No AI reply; goes straight to the owner the same day
Exchange wanted, but the other colour is out of stock5Draft offers a refund or a restock alert; a person sends it
One email about two different orders4AI flags it rather than guessing which order; a person splits it
Angry or threatening message2No draft at all; owner handles it

Twenty-six awkward emails in two months is a lot for a shop that assumed returns were "all the same". Without this list, those 26 would have been the first week-three complaints.

7. Nobody maintains it after launch

Software changes, services change, and instructions go stale. By week three: this one doesn't show yet. It shows at month three, when an app update breaks a step and the error alerts go to an inbox nobody reads.

Worse is the break that raises no alert at all. An illustrative case: a small cleaning company's website form feeds new enquiries into its CRM through an automation. Someone tidies the form and renames the "Phone" field to "Best number to call". The automation doesn't fail; it passes an empty value, because the field it was mapped to no longer exists, and a blank phone number isn't an error. Six weeks later the sales lead wonders why callbacks have dropped and finds 31 enquiries with no number. What would have caught it: a five-minute monthly check where the owner opens the last five records the automation created and reads them.

Causes 1, 4 and 5 are about people, which is why change management for small teams is worth reading even if "change management" sounds like something only large firms need.

The times it really is the tech

Sometimes the technology is the problem. It's worth being honest about when:

  • The inputs are too messy. Handwritten notes, poor phone scans and noisy audio all lower accuracy, sometimes below the point where checking is quicker than doing.
  • The connection you need doesn't exist. An app is listed on an automation platform, but not with the trigger or action the workflow needs.
  • The vendor changes something. A model update changes how the tool responds to the same instructions, a feature moves to a pricier tier, or a product is withdrawn. Microsoft's Copilot Pro, for instance, stopped being sold to new customers in October 2025 and support for existing subscribers ended on 1 August 2026. OpenAI closed its Sora video apps in April 2026 and switched off the Sora API on 24 September 2026, so any workflow built on it simply stopped. Keep a note of which step depends on which product, so a withdrawal means replacing one step rather than rebuilding everything.
  • Usage costs grow with volume. Per-task or per-credit pricing that looked trivial in testing becomes significant at full volume.

Even these are usually visible in a proper test on real cases before launch. When a project fails for a technical reason, the underlying failure is often that the test was skipped.

Some "failures" are really expectations set by a demo

A surprising number of abandoned projects were working, just not as well as someone had been led to expect. Vendor demos, social media clips and conference talks show the best run on the cleanest example. A workflow that cuts a task from 40 minutes to 25 is judged a flop because the demo showed it done in 30 seconds.

Three habits keep expectations honest:

  • Write the target down before you start, in minutes or errors, and make it modest. For a first workflow, a third off the time of a frequent task is a good result. If you beat it, good.
  • Count the whole job. Time from starting the task to the finished result being sent, including checking and edits. Generation speed is irrelevant if the checking takes 20 minutes.
  • Compare against your own baseline, not the demo. The only fair question is whether the task is now cheaper or better than it was in your business, with your inputs.

The opposite problem exists too: a project declared a success on launch day because the first three outputs looked impressive. Both are fixed by the same thing, a number recorded before the change and measured the same way after it.

One failed project replayed: a video production company

Consider an illustrative case: a nine-person video production company tries an AI workflow for interview-led films. A senior editor who is keen on AI sets it up: interview recordings are transcribed automatically, and an AI step picks the strongest soundbites and produces a draft paper edit, a text list of clips with timecodes that editors use to assemble a first cut.

The first attempt. The goal is "speed up the edit". Nobody records how long first assemblies take now. The senior editor builds it over three weeks of evenings. At launch, some timecodes drift out of sync, and the AI picks soundbites that sound punchy but ignore the client's key messages, because the client brief was never part of its input. Two of the four editors never open the paper edits. Two months later the senior editor is buried in a large job, a storage folder is renamed, and the workflow stops. Nobody notices for five weeks. The verdict: "AI can't do edits."

The replay, same tools. The producer owns the project, with two hours a week. The goal is specific: cut the time from receiving interview footage to a first assembly from about 6 hours to 4 for a typical 45-minute interview, with a baseline taken across five jobs. The client brief's key messages are added to the AI step's input. Before launch, the workflow runs on five past interviews and its picks are compared with the clips the editors actually used. The output is renamed from "paper edit" to "shortlist", and editors still make the choices. Failure alerts go to the producer, who checks one job a month.

The illustrative result. First assemblies take about 4.5 hours rather than 6, and editors use the shortlists on most jobs. It's a modest gain, but it's real, measured and still running. The technology was identical both times.

An eight-question check before you start

Answer these in writing before any building begins. It takes about 30 minutes, and running a pre-mortem is a good way to go deeper on question 3.

1. Who owns this, and which two hours a week are theirs for it?
2. What single number are we trying to move, and what is it today?
3. What result, by what date, would make us stop?
4. Does everyone who does this task do it the same way? If not, which way wins?
5. What are the five most awkward cases, and what should happen to each?
6. Once it's live, who checks the output, and how often?
7. Whose account does it run on, who pays, and where do error alerts go?
8. When is the review, and who will be in the room?

If you can't answer questions 1, 2 and 4, don't start yet. Those three gaps account for most of the chain above.

Filled in, the check is short. Here are illustrative answers from a five-person landscaping firm planning AI-drafted follow-ups on quotes:

1. Office manager. Tuesday and Thursday, 8 to 9am, in the diary.
2. Quotes followed up within seven days. Today: 11 of the last 40.
3. Stop if fewer than 30 of the next 40 get a follow-up by week 8,
   or if two customers complain about the tone.
4. No. The owner phones, the office manager emails, the two team
   leaders don't follow up. Agreed: email on day 5, a call on day 12
   from whoever priced the job.
5. Accepted verbally on site (no email); revised quote requested
   (hold); quote over $15,000 (owner calls personally); customer
   said "after winter" (follow up in the spring); customer is
   also a supplier (no automated email).
6. Office manager reads every draft for the first month, then
   five a week.
7. The company's shared admin login, paid on the business card.
   Alerts go to the office manager and the owner.
8. Friday of week 8, owner and office manager, 30 minutes.

Answer 4 is where most of the value sits. Until someone wrote it down, nobody had noticed that the team leaders' quotes were never chased at all, and fixing that is a bigger gain than anything the AI adds to the wording.

Signs an AI project should be stopped, not rescued

Not every project should be rescued. Stop on purpose when:

  • After six weeks of real use alongside the old method, the number hasn't moved.
  • Checking the output takes as long as doing the task by hand.
  • The owner has left and nobody can take it on for the next three months.
  • At real volume, the running cost is more than the time it saves is worth.

Stopping well means writing half a page on what you learned, keeping the test cases, and switching off the automation cleanly so nothing half-works in the background. If a project isn't failing so much as stuck between test and launch, that's a different problem, covered in why your AI pilot stalled and how to get it live.

Questions after a failed AI project

What percentage of AI projects fail?

Widely quoted failure rates come from surveys of large organisations, and they define failure in different ways, so the figures vary widely and don't transfer well to a five-person firm. A more useful question for a small business is whether your project shows the causes described here: no owner, no baseline, no time for staff to change habits, and no plan for awkward cases.

How long should we run a project before judging it?

Give a first workflow six to eight weeks of real use after testing, with the old method running alongside for at least the first two to four. Judging after one week mostly measures the learning curve. Judging after six months without interim checks usually means the project has already drifted. Set the review date before you launch, not after.

Is it worth trying again after a failed AI project?

Usually, if you can name why it failed. Write half a page on which causes applied, keep the test cases, and change the conditions before retrying: a named owner with time, a baseline, and the awkward cases planned for. If the honest reason was that the job suits simple rules or doesn't happen often enough, pick a different job instead.

Further reads

Sources: Microsoft support notice on the retirement of Copilot Pro; Zapier help pages (checked September 2026).

Want to restart an AI project that stalled?

On a 1:1 call we'll work out which causes applied to your last attempt, decide whether the job is still worth doing, and set up the owner, baseline and checks it needs this time.

Book a 1:1 call with me