AI Hallucinations Explained for Business Owners: Causes and Fixes

Coding Liquids tutorial cover featuring Sagnik Bhattacharya for AI Hallucinations Explained for Business Owners: Causes and Fixes.
Coding Liquids tutorial cover featuring Sagnik Bhattacharya for AI Hallucinations Explained for Business Owners: Causes and Fixes.

An AI hallucination is a confident answer that's false: an invented statistic, a quote nobody said, a policy that doesn't exist. Models produce them because they generate plausible text rather than look facts up. You limit the damage by giving the AI your sources, letting it say "not stated", and checking every name, number and quote before anything leaves.

You can't switch hallucinations off; every current model produces them sometimes. What you can control is where they're likely, how fast they're caught, and whether any reach a customer. That starts with understanding why they happen, which is simpler than it sounds, and ends with two prompts and a checking routine you can use today.

Follow me on Instagram@sagnikteaches

Why a model makes things up

A language model doesn't look anything up in a database of facts. It writes by predicting, word by word, what text is most likely to come next, based on patterns in the enormous amount of writing it was trained on. For facts that appear often in that writing, the likeliest continuation is usually the true one. For facts that appear rarely or never, such as a small firm's founding year, the exact wording of a niche regulation, or the spec of your client's new product, the likeliest-sounding continuation can simply be wrong.

Connect on LinkedInSagnik Bhattacharya

There's a second reason. OpenAI published research in September 2025 arguing that the way models are trained and scored rewards guessing over admitting uncertainty. It compares this to a student facing a multiple-choice exam: leaving a question blank scores nothing, so guessing is the better strategy. Models learn the same lesson, so a confident answer tends to come out where "I don't know" would be more honest.

Subscribe on YouTube@codingliquids

Two consequences matter for a business. First, a hallucination isn't a lie; there's no intent, and nothing in the model knows the statement is false. Second, it's written in exactly the same fluent, assured tone as the correct parts. You can't spot one by how it sounds. You can only spot it by checking.

Where hallucinations show up in business work

Kind of outputWhat typically gets inventedHow much it matters
Press releases and marketing copyStatistics, survey results, awards, customer numbersHigh: public and hard to retract
Proposals and tendersPast project details, client names, performance figuresHigh: can breach tender rules or mislead a buyer
Summaries of long documentsA clause, figure or date that isn't in the documentMedium to high, depending on what the summary is used for
Answers about laws, rules and deadlinesArticle numbers, thresholds, dates, requirementsHigh: acted on without checking
Research on competitors and suppliersProducts, prices, people, company historiesMedium
References, sources and linksBook titles, report names, URLs, case citationsHigh if published or relied on
Customer-facing chatbot repliesRefund terms, prices, policies, availabilityHigh: customers act on them
Brainstorms and internal first draftsAnythingLow, as long as nothing is carried forward unchecked

Figures are a special case. As well as inventing numbers, chat assistants are unreliable at calculating them, which is a separate weakness covered in why AI is bad at maths.

What makes hallucinations more likely

Most hallucinations in business use are triggered by the way the request is made. These are the common triggers, roughly in order of how often they cause trouble:

  1. Asking for facts without supplying a source. "What's our competitor's pricing?" invites a plausible guess.
  2. Asking for things that exist to fill a gap. "Give me three statistics that support this point" is an instruction to produce three statistics. It will, whether or not they exist. (There's a worked example of this one after the list.)
  3. Niche or recent subjects. The less that has been written about something, the more the model fills in.
  4. Very long inputs. Details buried deep in a long document are more likely to be missed, merged or misattributed than details near the start or end.
  5. Leading questions. "Why is option A the best choice?" gets reasons for A, invented if necessary.
  6. Demands for completeness. "List every requirement" pushes it to pad the list rather than stop at what it knows.
  7. Long, wandering chats. An early wrong assumption gets repeated and built on. Starting a fresh chat with the correct facts often fixes it.

Trigger 2 is the easiest to see in action. Picture a small bike shop writing its spring newsletter, with no web search switched on, and this prompt:

Give me three statistics showing that more people are cycling to work,
with sources, for our newsletter.

An illustrative reply, typical of what comes back:

1. Commuter cycling has risen by 34% since 2020 (National Cycling
   Trends Report, 2024).
2. 6 in 10 office workers say they would cycle if their employer
   offered secure parking (Workplace Mobility Survey).
3. E-bike sales grew 48% year on year (Industry Sales Index, 2025).

All three look citable. None comes with a link, the report names are generic enough to belong to anything, and there's no sign the model read any of them; treat every figure as invented until you find the original. The rewrite asks for what the model can honestly give:

I want to say in our newsletter that more people are cycling to work.
Don't give me any figures. Tell me which kinds of organisations
publish commuting or cycling data, and what search terms I should use
to find their latest figures myself.

That takes five more minutes of searching, and the shop ends up quoting one real figure with its source named, or none at all, which reads better than three that a customer could look up and fail to find.

Trigger 5 is subtler, because the invented part hides among true ones. An illustrative café owner pastes two quotes from coffee-machine suppliers and asks, "Why is Supplier A the better choice for us?" The reply gives four reasons. Three are in the quotes: a lower monthly lease, free descaling visits and a two-year warranty. The fourth, "faster engineer response times", isn't mentioned anywhere; the question asked for reasons A wins, so the model supplied one more. The neutral version asks for a comparison instead:

Here are two supplier quotes. Compare them in a table on: monthly cost,
contract length, servicing, warranty, engineer response time and
cancellation terms. Where a quote doesn't say, write NOT STATED.
Then give the strongest reason against each supplier.

In the table that comes back, response time reads NOT STATED for both suppliers, which turns an invented advantage into a question to ask both sales reps before signing.

Trigger 4 shows up in long contracts. An illustrative accountancy practice uploads a 70-page software reseller agreement and asks for the key terms. The summary says there's "a 12-month minimum term with automatic renewal". The agreement does contain a 12-month minimum, but in a schedule for a different product the practice doesn't resell; its own product is on a rolling monthly term. The model merged two passages that looked alike. Asking for the clause or schedule number after every term, and summarising the main agreement and each schedule separately, makes that kind of mix-up visible in seconds.

Trigger 7 catches people in long planning sessions. The owner of an illustrative dog-grooming salon spends an evening with an assistant planning a price rise. Early on she mentions "we have three groomers". Forty messages later she says, once, that a fourth has just joined, and the final table of staff costs per appointment still uses three. Nothing new was invented; the model kept building on the first figure it was given. When a chat has run long and the facts have changed, start a fresh one with a short list of current facts at the top.

Fixes that work, from least to most effort

1. Give it the source and fence it in

Paste in or upload the material the answer should come from, and tell it to use only that. This single change removes most business hallucinations.

2. Give it permission to say "not stated"

Models guess partly because nothing tells them a blank is acceptable. Say explicitly that "NOT STATED" is a good answer when the material doesn't cover something.

3. Ask for the supporting quote

Asking it to put the supporting passage after each claim makes unsupported claims easy to see, and makes your checking faster. Here's a prompt that combines the first three fixes:

Use only the material between the lines below. Do not add any facts,
figures, names, dates or quotes that are not in it.
If the material doesn't answer something, write NOT STATED instead of guessing.
After each factual sentence, add the supporting words from the material
in [square brackets].

Task: [what you want written or answered]

-----
[paste the source material here]
-----

Here's what that prompt's output looks like when a small catering company asks for a summary of a venue's supplier terms. An illustrative extract:

Invoices are due 30 days after issue.
  ["Invoices are payable within thirty (30) days of issue."]
Either side can end the agreement with 14 days' written notice.
  ["either party may terminate on fourteen days' written notice"]
The venue must give 30 days' notice of any price change.
  ["prices may be revised annually"]
Late-payment interest: NOT STATED.

The brackets do the work. Lines one and two are supported word for word. Line three isn't: the quoted words say prices can change once a year and nothing about notice, so the 30 days is the model filling a gap with what supplier terms usually say. Without the brackets, that line would have read as settled. The NOT STATED is useful too; it tells the caterer exactly what to ask the venue.

4. Use tools built to answer from your documents

Some tools are designed around your own sources. Gemini Notebook (formerly NotebookLM) answers questions from the documents you upload and cites the passages it used. ChatGPT and Claude both offer Projects, where you attach reference files that every chat in the project draws on. For questions your staff ask repeatedly, building a company knowledge base AI can answer from goes further. Web search modes help for current facts too, but only if someone opens the cited pages.

5. Separate writing from checking

Ask for the draft in one step and the check in another, ideally in a fresh chat so the checker isn't defending its own work. This prompt is quick to run on any draft:

Below are a draft and the source material it was written from.
List every factual claim in the draft: each name, number, date, quote,
and statement about a product, price, law or policy.
For each one, say whether the source material supports it and quote the support.
Mark anything without support as UNSUPPORTED. Do not rewrite the draft.

DRAFT:
[paste draft]

SOURCE MATERIAL:
[paste source]

Run on a landscaping firm's new "About us" page, with the owner's notes as the source, the check might return (illustrative):

"Family-run since 2011"        SUPPORTED   ["started the business in 2011"]
"Over 400 gardens completed"   UNSUPPORTED (notes say "a few hundred")
"Fully insured"                SUPPORTED   ["public liability insurance"]
"Award-winning designs"        UNSUPPORTED (no award in the notes)
"Free 3D design with every quote" UNSUPPORTED (notes: "3D design available")

Three unsupported lines on one short page is typical. The first two are harmless-sounding inflation, and the third is a promise: a customer who asks for the free 3D design is entitled to wonder why they're being charged for it. Each gets rewritten from the notes, and the owner decides whether "a few hundred" is worth replacing with a real count from the job records.

6. A human check on the six things that hurt

Before anything leaves the business, a person checks: names (people, companies, products), numbers, dates, quotes, links and references, and promises (anything that commits you to a price, deadline or policy). Everything else is style. A five-minute fact-check routine for AI output sets this out as a repeatable habit.

A house rule: how much checking each output gets

Checking everything to the same standard wastes time on brainstorms and under-checks the things that matter. A simple tiered rule, written down once, lets everyone decide in seconds:

TierExamplesCheck required
1. Stays internal, nobody acts on itBrainstorms, rough outlines, meeting prepNone beyond common sense
2. Internal, someone acts on itSummaries used for decisions, research notes, supplier comparisonsKey facts traced to a source before acting
3. Goes to a client or customerEmails, reports, proposals, chatbot answersThe six-point check by a named person
4. Public, contractual or financialPress releases, website claims, tenders, anything with prices or legal termsThe six-point check plus a second person, and every statistic linked to its original source

Two refinements make the rule stick. First, the person who ran the AI prompt shouldn't be the only checker at tier 4, because they've already read the draft as correct once. Second, if a tier 2 summary is later pasted into a tier 3 email, it moves up a tier and gets the fuller check. Most slips happen at that hand-over, when something checked lightly for internal use travels further than intended.

A realistic version, as an illustration: a five-person IT support firm asks an assistant to summarise a software vendor's long licensing announcement for the team. It's tier 2, so the engineer checks the headline change and moves on. A week later a colleague pastes two paragraphs of that summary into emails to eight clients, including the line "existing licences renew at the current price until 30 June". The announcement never gave a date; the model had added one. Three clients plan budgets around it, and the firm spends an afternoon sending corrections. Nobody did anything careless by their own tier's standard. The slip was in not re-checking the paragraph when it moved from tier 2 to tier 3.

A PR consultancy's press-release check

An illustration. Say a four-person PR consultancy drafts a product-launch release for a client using AI. The brief contains the client's product sheet and an approved quote from the client's managing director. The draft reads well. It also contains three problems:

  • "According to a recent industry survey, 68% of small firms struggle with…" There's no survey in the brief. The model produced a statistic because releases usually have one.
  • The managing director's quote has gained an extra sentence she never said, smoothly written in her style.
  • The launch date has shifted from the 14th to the 4th.

The account manager runs the checking prompt, which takes about two minutes and flags the survey figure and the extra quote sentence as unsupported. Her own read of names, numbers and dates against the brief takes another eight and catches the date, and a colleague gives it the second read that tier 4 requires. Ten minutes of checking against about 90 minutes to write the release from scratch is still a good trade, but only because the checking happened.

The consultancy changes two things afterwards: the drafting prompt now says "no statistics unless they appear in the brief", and the final quote always goes back to the client for approval, which good practice required anyway. Catching made-up figures in AI-drafted proposals applies the same approach to sales documents.

When a hallucination reaches a customer

This is where the business risk becomes real. In February 2024, a tribunal ordered an airline to compensate a customer after the chatbot on its website described a refund policy that didn't match the airline's actual policy. The airline argued, in effect, that the chatbot was responsible for its own statements; the tribunal rejected that. The practical lesson for any business: what your AI tells a customer will generally be treated as what you told them.

If it happens to you:

  1. Correct it quickly and plainly, in the same channel the customer received it.
  2. Decide what's fair. Where the customer acted reasonably on what they were told, honouring it is often cheaper than the argument. Take advice if the sums are significant.
  3. Find the cause. Missing source material? No human check? An out-of-date document the tool answered from?
  4. Log it, so the same failure is spotted if it recurs. Keeping an AI error log shows a simple format.

For the conversation with the customer, what to do when AI gets something wrong with a customer has wording you can adapt.

More questions about hallucinations

Will newer AI models stop hallucinating?

Newer models tend to hallucinate less on many tests, and vendors are working on it, but no current model is free of the problem. OpenAI's own research describes hallucinations as a persistent issue linked to how models are trained and scored. Plan on keeping your checking routine whatever model you use, and treat any improvement as a bonus rather than a reason to stop checking.

Does turning on web search stop the AI making things up?

It helps for current facts, because the answer is built from pages it has just read and it usually shows the sources. But it can still misread a page, blend two sources, or rely on a poor one. Open the cited pages for anything important and confirm the claim is really there, in the form the AI gave it.

Can I lower the AI's temperature setting to stop hallucinations?

Usually not. Temperature, which controls how varied the wording is, was only ever adjustable through the API, and many current models no longer accept it: Anthropic has deprecated it on newer Claude models and OpenAI's reasoning models don't support it. Where it exists, a lower setting makes answers more consistent, not more truthful. Supplying the source material and allowing 'not stated' matter far more.

Further reads

Sources: OpenAI research summary 'Why language models hallucinate' (September 2025); Gemini Notebook help pages; OpenAI and Anthropic help pages on Projects; published tribunal decision on an airline chatbot (February 2024).

Worried about AI getting facts wrong for clients?

On a 1:1 call we'll look at where your team uses AI for anything factual, set up source-based prompts for those jobs, and agree a checking routine that fits the time you have.

Book a 1:1 call with me