The Limits of AI: What It Still Gets Wrong in a Small Business

Coding Liquids tutorial cover featuring Sagnik Bhattacharya for The Limits of AI: What It Still Gets Wrong in a Small Business.
Coding Liquids tutorial cover featuring Sagnik Bhattacharya for The Limits of AI: What It Still Gets Wrong in a Small Business.

AI still gets facts, figures and judgement wrong. It invents references, miscalculates totals, misses details buried in long documents, works from out-of-date knowledge, and sounds equally confident whether it's right or not. It also knows nothing about your clients, prices or history unless you give it that information in the conversation.

None of that makes it useless; it tells you where to put it. The practical rule is to use AI where a mistake is cheap to catch, and to add a specific check wherever it isn't. Each limit below comes with where it bites in a small firm, how to spot it and the check that catches it. There's a one-page table at the end to pin up, and a simple way to decide which tasks are safe to hand over.

Follow me on Instagram@sagnikteaches

It invents references that look real

A language model writes by predicting plausible text, and the format of a reference is very easy to predict. So it can produce a case name, a clause number, a standard's section or a statistic with a named source, all in the right format, none of it real. Courts have sanctioned lawyers who filed submissions citing cases that an AI tool made up.

Connect on LinkedInSagnik Bhattacharya

It isn't only a legal problem. An engineering consultancy asking for the clause in a design standard that covers a point, or an insurance broker asking which section of a policy wording excludes flood damage, can get an answer that looks exactly like the real thing.

Subscribe on YouTube@codingliquids

Here's how that plays out for the broker, in an illustrative exchange. Asked with no document attached, "Which section of a standard commercial property wording excludes flood?", a chat assistant might reply:

Flood is typically excluded under Section 4.2(c), "Exclusions: Flood and Surface Water", unless the Flood Extension endorsement (FE-01) has been added to the schedule.

Every part of that is formatted like a real reference, and none of it comes from the insurer's document; it's a plausible pattern. Attach the actual wording and change the question to "Quote the exact sentence that excludes flood, with its page number and heading. If you can't find one, say so." Now the answer points at a page you can open and a sentence you can compare word for word, which takes ten seconds.

Warning sign: any reference you didn't supply yourself. The check: open every reference at its source before it goes anywhere. Asking the AI to quote the exact passage helps, because a quote is easier to verify, but it can invent quotes too. Tools that answer from your own documents and link to the passage are safer, and still worth clicking. AI hallucinations explained for business owners covers why this happens.

It adds up badly unless it actually runs a calculation

A model predicts digits the way it predicts words, so a long total can come out slightly wrong while looking perfectly reasonable. Some assistants can switch to running a real calculation in code, but they don't always do it, and they don't always tell you which they did.

This bites wherever numbers leave the building: premium comparisons across a dozen policies, a fee schedule with expenses, a timesheet summary for billing. Warning sign: a total with no working shown. The check: recompute anything client-facing in a spreadsheet, or ask the tool to show its calculation step by step and then check the steps. Why AI is bad at maths has the full routine for quotes and invoices.

A realistic version, for illustration: a painting and decorating firm pastes 14 line items for a three-bedroom repaint into a chat assistant and asks for a tidy quote with a total. The quote comes back at $6,840. The lines actually add up to $6,480. Nobody in the office queried it, because $6,840 sounded about right for the job; the client did, after adding the lines up on a phone calculator. The firm corrected it and looked careless. A =SUM() in the spreadsheet the lines came from would have caught it in seconds. The safer habit is to let the spreadsheet produce every total and ask the AI only for the wording around it.

It skims long documents

Current models can take in very long documents, but taking them in isn't the same as weighing every page equally. Summaries of long material tend to be tidy and can drop the awkward detail that mattered most.

Say an engineering consultancy asks for a summary of an 80-page tender pack. The summary covers scope, programme and deliverables neatly, and leaves out a penalty clause for late delivery on page 51. Nothing in the summary looks wrong; the problem is what isn't there.

Warning sign: a summary with no page or section references, or one that's much neater than the document. The check: ask targeted questions instead of "summarise this": "List every clause that mentions penalties, retention, liability or insurance, with page numbers." Split very long documents into sections, and have a person read the parts that carry money or liability.

Its knowledge stops at a date

Every model is trained on material up to a cut-off date. After that it doesn't know about price changes, new regulations, a client's change of ownership or a competitor's new service unless it can search the web, and web search can surface an out-of-date page just as easily as a current one.

A recruitment agency asking for current salary ranges for a role, or a broker asking about a change in an insurer's terms, can get last year's answer stated as this year's. The check: supply current facts yourself and let the AI do the writing. When it cites a web source, look at the date on the page.

It doesn't know your business

Out of the box, AI writes for an average business: a generic tone, standard services, typical prices and whatever terms of business are common. Ask it for a client letter and you'll get a competent letter from a firm that doesn't exist.

Warning sign: sentences that could appear unchanged on any competitor's website. The fix: give it your context every time, or once in a saved project: who you serve, what you do and don't do, your house style, two or three past examples you liked. How to give AI your business context shows what to include.

The difference is plain side by side. A two-person dog-grooming salon asking only "Write a message telling customers our prices are going up" might get something like this (illustrative):

Dear valued customer, to continue providing the exceptional service you deserve, we will be making a small adjustment to our pricing. We appreciate your continued loyalty.

Give it the facts and the salon's voice (two groomers, customers know them by first name, a full groom for a medium dog goes from $55 to $60 next month, nail trims and bath-only visits stay the same, first rise in three years) and the draft becomes:

Hi, a quick heads-up from both of us: from [date], a full groom for a medium dog goes from $55 to $60. Nail trims and bath-only visits aren't changing. It's our first price change in three years, and it covers the rising cost of our shampoos and insurance.

Much better, and still not ready. The last clause gives a reason the owner never supplied. It may even be true, but it's the AI filling a gap with something plausible, which is the same habit as the invented reference above. Delete it or replace it with the real reason.

It sounds just as sure when it's wrong

This is the limit that causes the most damage, because it disables the usual human alarm. A colleague who's unsure hesitates. AI doesn't; a wrong answer arrives in the same calm, well-organised prose as a right one.

A field experiment with 758 consultants at BCG, run by researchers including Fabrizio Dell'Acqua, shows the effect. On tasks within the AI's capabilities, consultants using GPT-4 completed 12.2% more tasks, 25.1% faster, with output rated over 40% higher in quality. On one task deliberately chosen to fall outside those capabilities, consultants using AI were 19 percentage points less likely to reach the correct answer than those working without it. The researchers called this the "jagged frontier": tasks that look equally hard can fall on either side of what AI does well, and you can't tell which from the confidence of the answer.

The check: treat fluency as zero evidence. Ask the tool "What would make this answer wrong?" and "What did you assume?", and make sure a person with the relevant knowledge owns every conclusion that matters.

The follow-up question often tells you more than the original answer. Take a small recruitment agency that asks for a recommended fee on a hard-to-fill role, gets a confident percentage, and then asks "What did you assume, and what would make this wrong?". An illustrative reply:

I assumed: (1) a permanent placement, not a contract; (2) the fee is based on first-year salary; (3) no rebate period or discount has been agreed with this client; (4) typical market fee levels, which I can't verify for your sector or for this year. If the client has negotiated preferred-supplier terms, the recommendation doesn't apply.

Assumption 4 admits the headline number was a general guess, and assumption 3 sends the consultant to check the client's terms of business. Neither admission appeared in the first, confident answer.

It can't weigh relationships, fairness or context it wasn't told

Whether to chase a 15-year client for a late invoice or ring them instead, how to break bad news to a candidate who nearly got the job, what to charge a client going through a difficult year: these depend on history and relationships the AI can't see. It will still offer a confident recommendation.

It can also carry bias from its training data into decisions about people, such as ranking CVs or suggesting who gets a discount. AI bias in small business decisions covers hiring, pricing and credit. Use AI to lay out options and draft the wording; keep the decision with the person who knows the relationship.

Automations fail quietly

Every limit above is manageable when a person reads the output. Inside an automation or an agent, nobody does. An email assistant files a complaint under "general enquiries", a workflow sends a renewal reminder to the wrong client, a transaction is coded to the wrong account every month, and none of it raises an alarm.

Here's what that looks like when it happens (an illustration). A small lettings agency's email sorter tags incoming messages as "maintenance", "rent" or "general". A tenant writes: "Following up on my last email. The boiler still isn't working and we have a newborn at home." The sorter reads the routine opening, files it under "general", which someone checks once a day, and the message sits for two days. The agency finds out when the landlord rings, furious, because the tenant went to them directly. Follow-up emails are exactly where this happens, because the urgent detail comes after a harmless first line. A weekly sample of 20 tagged emails, with someone asking "would I have filed this the same way?", would most likely have shown the pattern in week one.

The check: run new automations alongside the manual process before they go live, keep a log of what the AI did, sample a few outputs every week, and set alerts for anything unusual. Piloting AI in shadow mode explains how to run it silently first.

The limits on one page

LimitWhere it bitesWarning signCheck that catches it
Invented referencesLegal research, standards, policy clauses, statisticsA reference you didn't supplyOpen every one at source
Bad arithmeticQuotes, fee schedules, premium comparisons, timesheetsA total with no workingRecompute in a spreadsheet
Skimmed documentsTenders, contracts, specifications, bundlesNo page references; too tidyTargeted questions with page numbers
Stale knowledgePrices, rules, market rates, client newsNo date on the sourceSupply current facts yourself
No knowledge of your firmClient letters, proposals, marketingCould be any firm's textSaved context and examples
False confidenceAdvice, analysis, recommendationsNone: that's the problemA knowledgeable person owns the conclusion
No judgement of peopleChasing, pricing, hiring, bad newsA firm recommendation about a personAI drafts options; a person decides
Silent automation errorsEmail sorting, record updates, remindersNothing, until a client complainsShadow run, logs, weekly samples

Deciding which tasks are safe to hand over

Two questions sort almost any task. How much does a mistake cost if nobody notices it? And how easy is a mistake to spot when someone looks?

Easy to spotHard to spot
Cheap if missedHand it over with a quick read. Internal emails, meeting notes, first drafts.Hand it over and sample regularly. Tagging, filing, sorting enquiries.
Costly if missedAI drafts, a person checks every one. Client letters, quotes, job adverts.Keep it human; AI helps with research only. Legal advice, structural calculations, final figures.

Try it on a real list. Say a six-person insurance broker has four candidate tasks. Drafting the covering email for renewal packs is cheap if slightly off and easy to spot, so it goes top left. Tagging incoming emails as claims, renewals or new business is cheap per mistake but easy to miss, so it goes top right with a weekly sample of 20. Summarising the differences between last year's and this year's policy wording is costly if a change is missed, but checkable against the documents, so it goes bottom left, with the broker reading both wordings for the clauses the summary flags. Advising a client which cover to buy stays bottom right: the AI can gather the facts, and the broker gives the advice.

Most of the value for a small firm sits in the top-left and bottom-left boxes: plenty of work where AI drafts quickly and a person checks cheaply. The damage almost always comes from the bottom-right box, where a task looked routine enough to automate and turned out to be judgement. When not to use AI in your business lists the tasks that belong there.

Further reads

Sources: Dell'Acqua and colleagues, 'Navigating the Jagged Technological Frontier' (field experiment with 758 BCG consultants, working paper, 2023); vendor documentation on context windows and knowledge cut-off dates.

Want to know which of your tasks are safe to hand over?

On a 1:1 call we'll sort your candidate tasks by the cost of a missed mistake, decide where AI drafts and a person checks, and set up checks that fit how your team already works.

Book a 1:1 call with me