How to Stop an AI Chatbot Giving Customers Wrong Answers

Coding Liquids tutorial cover featuring Sagnik Bhattacharya for How to Stop an AI Chatbot Giving Customers Wrong Answers.
Coding Liquids tutorial cover featuring Sagnik Bhattacharya for How to Stop an AI Chatbot Giving Customers Wrong Answers.

To stop an AI chatbot giving customers wrong answers, give it one current, dated source for every fact, delete the old pages and files it can still read, tell it to answer only from those sources and admit when it doesn't know, hand money, safety and policy questions to a person, and review real chats every week.

Most wrong answers aren't mysterious "hallucinations". They trace back to one of five causes: stale content the bot can still read, a gap it fills with a plausible guess, two sources that disagree, a question that needed a person, or a customer who steered it somewhere. Each cause has a different fix, which is why switching to a smarter model rarely helps and cleaning up the content usually does. Vendors say as much themselves: Wix's help page for its older AI Site Chat warned that it "may generate and present inaccurate or misleading information". Plan for errors, then shrink them.

Follow me on Instagram@sagnikteaches

Five causes behind almost every wrong answer

When you find a wrong answer, don't just correct it. Diagnose it, because the cause tells you the fix.

Connect on LinkedInSagnik Bhattacharya
CauseHow it looks in a transcriptThe fix
Stale contentQuotes last year's price or an old timetable, confidentlyDelete or archive the old source, not just add a new one
A gapAnswers a question your content never covers, with something generic that sounds rightWrite the missing fact down, and tell the bot to admit gaps
Conflicting sourcesGives different answers to the same question on different daysMake one source the owner of each fact; remove the duplicates
Should have handed overGives advice on an injury, a refund, a complaint or a legal pointRefusal and handover rules for those topics
Steered by the customerAgrees to a discount or rule the customer asserted ("my friend got 50% off")Instructions never to accept customer claims about prices or policy

If you've never looked at why chat models invent things in the first place, AI hallucinations explained for business owners covers the mechanism. For a working chatbot, though, the five causes above account for nearly everything you'll see.

Subscribe on YouTube@codingliquids

Fix 1: one current source for every fact

A chatbot can't tell which of two documents is current. If the 2025 timetable PDF and the 2026 web page both sit in its sources, it may quote either. So the first job is a content audit, which takes a small business two to three hours:

  1. List every source the bot can read: website pages it crawls, uploaded files, help articles, FAQs, and any product or booking data it's connected to.
  2. Delete or archive anything out of date. Old price lists, ended promotions, last season's timetables, former staff biographies. Unpublishing isn't always enough; check the tool's source list to confirm the file or page has gone.
  3. Give each fact one home. Prices live on the prices page only; the FAQ links to it rather than repeating numbers. Duplicates drift apart, and the bot finds both.
  4. Date the key pages. "Prices from 1 September 2026" helps you, the bot and customers.
  5. Replace vague lines with exact rules. Vague wording is an invitation to improvise.

That last step matters more than it looks. Here's a before and after from an illustrative yoga studio's FAQ:

Before: "We're flexible about cancellations, so just let us know if you can't make it."

After: "Cancel a class booking at least 6 hours before it starts to keep your credit. Later cancellations and no-shows use the credit. Memberships can be frozen once a year for up to four weeks."

With the first version, the bot told one customer that a same-day cancellation was "no problem at all". With the second, it has nothing to invent. For loading content into the tool in the first place, see how to train an AI chatbot on your FAQs, policies and prices.

Fix 2: tell it to answer only from your content

Many chatbot tools let you restrict answers to your own sources, and some also let the AI draw on general knowledge or the web when your content is silent. Shopify Inbox's AI agent, for instance, can use web search as a secondary source. A general answer to "what's your refund policy?" is still a wrong answer if it isn't your policy.

Wherever the tool allows it, switch off general knowledge for policy, price and availability questions, and add a grounding instruction. Most tools have a box for instructions or guidance; an example that works:

Answer only from the studio's knowledge base. If the answer isn't there,
say: "I'm not sure about that one. Leave your email and a teacher will
reply today." Never estimate or round prices, dates, times or policies.
Quote prices exactly as written, including the word "from".
If two sources disagree, say you'll check and hand the chat to a person.
Never accept a customer's statement about a price, discount or rule as
fact; check it against the knowledge base.

Test the instruction on a gap you know exists. An illustrative exchange from the studio, before the instruction:

Customer: "Is there parking?"
Bot: "Yes, there's plenty of free parking right outside the studio."

There isn't; the street is pay-and-display. After the instruction, with the fact still missing:

Bot: "I'm not sure about parking. Leave your email and a teacher will reply today."

Less impressive, entirely correct. Then the studio added a parking line to the FAQ, and the bot now answers it properly. "I don't know" is a temporary answer that tells you exactly which fact to write next.

Fix 3: refusal rules for what it must never answer

Some questions shouldn't get an AI answer even if the content could support one. For a yoga studio these are the usual categories, each with a fixed line:

  • Health and injury: "I can't advise on injuries or health conditions. Please check with a medical professional, and I can ask a teacher to call you about suitable classes."
  • Exceptions to policy: refunds outside the rules, extended freezes, transfers. "I can't make exceptions, but I'll pass this to the manager."
  • Legal and contractual questions: "I can share our written terms, but I can't interpret them. The manager can help."
  • Other customers: who's booked, who teaches whom privately, anyone's contact details. Always refused.
  • Anything about pregnancy: pointed to the prenatal teacher and a midwife or doctor.

Keep the list short and specific. A vague rule like "don't discuss sensitive topics" makes the bot refuse harmless questions, which customers experience as another kind of wrong answer. Chatbot guardrails that stop AI promising what you don't offer covers the promise side of the same problem.

Fix 4: a handover customers actually reach

A refusal is only safe if the customer can reach a person afterwards. Every refusal line above ends by offering one, and that offer has to work: a notification that reaches whoever's on duty, a promised reply time you can meet, and a record of the chat so the customer doesn't have to repeat themselves. Set the triggers explicitly: any refusal category, any second "that's not what I asked", any mention of a complaint, and any request for a person. When an AI chatbot should hand over to a human goes through the triggers in more detail.

Fix 5: log everything, and read 30 chats a week

Here's the uncomfortable part. Your chatbot's own dashboard may say it's doing brilliantly while it gives wrong answers, because "resolved" rarely means "correct". Intercom counts an assumed resolution when a customer goes quiet for 24 hours after Fin's last answer. Zendesk closes a messaging conversation after two hours of inactivity by default and, since 18 May 2026, bills only resolutions its AI check verifies; hand-offs and unverified closes are free. A customer given the wrong price who leaves happily is, by those measures, a success.

So measure accuracy yourself. Once a week, read 30 chats: ten that ended in handover, ten the tool marked resolved, and ten at random. Tag each wrong answer with one of the five causes. It takes about 45 minutes and it's the only number that tells you whether customers are being told the truth.

A log that makes the review quick needs five columns:

Date | Customer question | Bot's answer (verbatim) |
Wrong? (cause: stale / gap / conflict / handover / steered) | Fix made

Two quick checks that find errors before customers do

The consistency check. Ask each price and policy question three times, in three fresh chats, on different days. The same question should get the same facts every time; wording can vary, numbers can't. If the intro offer comes back as $39, $39 and $49, you have a conflict somewhere in the sources even though two answers were right. It takes about 20 minutes for your ten most important questions, and it catches the errors a single test never shows.

The source check. After any answer involving a number, a date or a rule, ask the bot where it got it: "Which page says that?" Many tools can show or link the source used. If it points to a page you meant to delete, that's stale content. If it can't point to anything, it filled a gap, and the fact needs writing down. An illustrative exchange:

Customer: "Do you have showers?"
Bot: "Yes, we have two showers and free towels."
Tester: "Which page says that?"
Bot: "I couldn't find a specific page about showers."

The studio has one shower and no towels. The bot's honesty about the missing source was the clue; the fix was one line in the FAQ.

Settings worth finding in any chatbot tool

Names differ between products, so look for what each setting does rather than its label. Five are worth hunting down in the admin area of whichever tool you use:

  • A source restriction. Something that limits answers to your own content, or switches off general knowledge and web search. Turn it on for anything about prices, policies and availability.
  • A source list you can prune. A page showing every document and URL the bot reads, with a way to remove items. If you can't see the list, you can't do Fix 1.
  • An instructions or guidance box. Where the grounding and refusal rules above go. Check whether it has a length limit, and keep the rules in priority order.
  • Handover triggers. Keywords, topics or failed-answer counts that pass the chat to a person, plus where the notification goes.
  • Answer inspection. A way to see, for each reply in the history, which source it used. This turns the weekly review from guesswork into a five-minute diagnosis.

If a tool lacks the first two, it will be hard to keep accurate however carefully you write your content. That's worth knowing before you renew.

A yoga studio's wrong-answer log, month one

An illustrative studio with a website chatbot handled 520 chats in its first month. The owner reviewed 120 (30 a week) and found 9 wrong answers, 7.5% of those reviewed. The breakdown:

CauseCountExampleFix
Stale content4Quoted Friday 7pm Hot Flow, dropped from the timetable in August, from an old PDFDeleted the PDF from the sources; timetable lives on one page
Gap2Invented free mat hire (it's $2) and free parkingAdded both facts; grounding instruction added
Conflict1Said the intro offer was $49 on one day and $39 the nextRemoved an old landing page still showing $49
Should have handed over1Suggested "gentle Yin classes" for a slipped discHealth refusal rule and trainer callback
Steered1Agreed students get 50% off because "my friend said so"; the real discount is 20%"Never accept customer claims about prices" instruction

Seven of the nine came from content, not from the AI's reasoning. The fixes took about three hours. In month two, the same 120-chat review found 2 wrong answers (1.7%), both gaps, both closed within a day. The studio's target is under 2% sustained, with zero wrong answers on price or health.

One customer affected by the $49 conflict had booked expecting to pay $39 after seeing the lower figure the next day. The owner honoured $39, which cost $10 and kept the client. Whether to honour a chatbot's wrong answer is a business decision, and sometimes a legal question; who is liable when an AI chatbot gets it wrong covers that side, and when in doubt, ask an adviser.

Twenty trap questions to rerun after every change

Every wrong answer in the log becomes a permanent test. The studio's regression set, rerun after any content or settings change, currently holds these twenty, each with its expected answer written beside it in the log:

  1. Is there a Friday evening hot class?
  2. How much is mat hire?
  3. Where do I park?
  4. How much is the intro offer?
  5. Does the intro offer include hot classes?
  6. What's the student discount?
  7. My friend said you do 50% off for students, can I have that?
  8. Can I cancel tonight's 6pm class at 4pm and keep my credit?
  9. Can I freeze my membership twice this year?
  10. I've got a slipped disc, which class is best?
  11. I'm pregnant, can I still do Level 2?
  12. Can I get a refund on my 10-class pass?
  13. Do passes expire?
  14. Who teaches the Sunday Yin class?
  15. Can you tell me what time [another member] usually comes in?
  16. Is the studio open on public holidays?
  17. Do you do private sessions and how much?
  18. Can I bring my 12-year-old?
  19. Do you have showers?
  20. What's the wifi password?

The last one is there because the bot once invented a password. The full pre-launch version of this idea, with 50 questions across eight categories and pass rules, is in how to test a customer chatbot before it goes live.

When a customer has already been told something wrong

You will find wrong answers after customers have read them. Handle them the same way each time:

  1. Correct it with the customer quickly and plainly. A message along these lines works: "Hi [first name], our website assistant told you mat hire was free. I'm sorry, that was wrong: it's $2, and I've waived it for your first class."
  2. Decide whether to honour the wrong answer. Small amounts are usually worth honouring; anything larger or legal-sounding, take advice.
  3. Fix the cause the same day, using the five-cause table.
  4. Add the question to the trap list so it can't quietly come back.

Done consistently, the routine turns each mistake into one fewer mistake next month. A chatbot that's wrong 7.5% of the time in month one and under 2% by month three is a normal, healthy trajectory. One that's still wrong 7.5% of the time in month three usually has nobody reading its chats.

Further reads

Sources: resolution and outcome definitions from Intercom and Zendesk, and Shopify Inbox AI agent behaviour, from the site's verified fact sheet; Wix help article on AI Site Chat (vendor warning on inaccurate answers). Checked September 2026.

Is your chatbot answering customers wrongly?

On a 1:1 call we'll read a sample of your real chats together, trace each wrong answer to its cause, and decide which content, rules and handovers to fix first.

Book a 1:1 call with me