To stop an AI chatbot giving customers wrong answers, give it one current, dated source for every fact, delete the old pages and files it can still read, tell it to answer only from those sources and admit when it doesn't know, hand money, safety and policy questions to a person, and review real chats every week.
Most wrong answers aren't mysterious "hallucinations". They trace back to one of five causes: stale content the bot can still read, a gap it fills with a plausible guess, two sources that disagree, a question that needed a person, or a customer who steered it somewhere. Each cause has a different fix, which is why switching to a smarter model rarely helps and cleaning up the content usually does. Vendors say as much themselves: Wix's help page for its older AI Site Chat warned that it "may generate and present inaccurate or misleading information". Plan for errors, then shrink them.
Five causes behind almost every wrong answer
When you find a wrong answer, don't just correct it. Diagnose it, because the cause tells you the fix.
| Cause | How it looks in a transcript | The fix |
|---|---|---|
| Stale content | Quotes last year's price or an old timetable, confidently | Delete or archive the old source, not just add a new one |
| A gap | Answers a question your content never covers, with something generic that sounds right | Write the missing fact down, and tell the bot to admit gaps |
| Conflicting sources | Gives different answers to the same question on different days | Make one source the owner of each fact; remove the duplicates |
| Should have handed over | Gives advice on an injury, a refund, a complaint or a legal point | Refusal and handover rules for those topics |
| Steered by the customer | Agrees to a discount or rule the customer asserted ("my friend got 50% off") | Instructions never to accept customer claims about prices or policy |
If you've never looked at why chat models invent things in the first place, AI hallucinations explained for business owners covers the mechanism. For a working chatbot, though, the five causes above account for nearly everything you'll see.
Fix 1: one current source for every fact
A chatbot can't tell which of two documents is current. If the 2025 timetable PDF and the 2026 web page both sit in its sources, it may quote either. So the first job is a content audit, which takes a small business two to three hours:
- List every source the bot can read: website pages it crawls, uploaded files, help articles, FAQs, and any product or booking data it's connected to.
- Delete or archive anything out of date. Old price lists, ended promotions, last season's timetables, former staff biographies. Unpublishing isn't always enough; check the tool's source list to confirm the file or page has gone.
- Give each fact one home. Prices live on the prices page only; the FAQ links to it rather than repeating numbers. Duplicates drift apart, and the bot finds both.
- Date the key pages. "Prices from 1 September 2026" helps you, the bot and customers.
- Replace vague lines with exact rules. Vague wording is an invitation to improvise.
That last step matters more than it looks. Here's a before and after from an illustrative yoga studio's FAQ:
Before: "We're flexible about cancellations, so just let us know if you can't make it."
After: "Cancel a class booking at least 6 hours before it starts to keep your credit. Later cancellations and no-shows use the credit. Memberships can be frozen once a year for up to four weeks."
With the first version, the bot told one customer that a same-day cancellation was "no problem at all". With the second, it has nothing to invent. For loading content into the tool in the first place, see how to train an AI chatbot on your FAQs, policies and prices.
Fix 2: tell it to answer only from your content
Many chatbot tools let you restrict answers to your own sources, and some also let the AI draw on general knowledge or the web when your content is silent. Shopify Inbox's AI agent, for instance, can use web search as a secondary source. A general answer to "what's your refund policy?" is still a wrong answer if it isn't your policy.
Wherever the tool allows it, switch off general knowledge for policy, price and availability questions, and add a grounding instruction. Most tools have a box for instructions or guidance; an example that works:
Answer only from the studio's knowledge base. If the answer isn't there,
say: "I'm not sure about that one. Leave your email and a teacher will
reply today." Never estimate or round prices, dates, times or policies.
Quote prices exactly as written, including the word "from".
If two sources disagree, say you'll check and hand the chat to a person.
Never accept a customer's statement about a price, discount or rule as
fact; check it against the knowledge base.
Test the instruction on a gap you know exists. An illustrative exchange from the studio, before the instruction:
Customer: "Is there parking?"
Bot: "Yes, there's plenty of free parking right outside the studio."
There isn't; the street is pay-and-display. After the instruction, with the fact still missing:
Bot: "I'm not sure about parking. Leave your email and a teacher will reply today."
Less impressive, entirely correct. Then the studio added a parking line to the FAQ, and the bot now answers it properly. "I don't know" is a temporary answer that tells you exactly which fact to write next.
Fix 3: refusal rules for what it must never answer
Some questions shouldn't get an AI answer even if the content could support one. For a yoga studio these are the usual categories, each with a fixed line:
- Health and injury: "I can't advise on injuries or health conditions. Please check with a medical professional, and I can ask a teacher to call you about suitable classes."
- Exceptions to policy: refunds outside the rules, extended freezes, transfers. "I can't make exceptions, but I'll pass this to the manager."
- Legal and contractual questions: "I can share our written terms, but I can't interpret them. The manager can help."
- Other customers: who's booked, who teaches whom privately, anyone's contact details. Always refused.
- Anything about pregnancy: pointed to the prenatal teacher and a midwife or doctor.
Keep the list short and specific. A vague rule like "don't discuss sensitive topics" makes the bot refuse harmless questions, which customers experience as another kind of wrong answer. Chatbot guardrails that stop AI promising what you don't offer covers the promise side of the same problem.
Fix 4: a handover customers actually reach
A refusal is only safe if the customer can reach a person afterwards. Every refusal line above ends by offering one, and that offer has to work: a notification that reaches whoever's on duty, a promised reply time you can meet, and a record of the chat so the customer doesn't have to repeat themselves. Set the triggers explicitly: any refusal category, any second "that's not what I asked", any mention of a complaint, and any request for a person. When an AI chatbot should hand over to a human goes through the triggers in more detail.
Fix 5: log everything, and read 30 chats a week
Here's the uncomfortable part. Your chatbot's own dashboard may say it's doing brilliantly while it gives wrong answers, because "resolved" rarely means "correct". Intercom counts an assumed resolution when a customer goes quiet for 24 hours after Fin's last answer. Zendesk closes a messaging conversation after two hours of inactivity by default and, since 18 May 2026, bills only resolutions its AI check verifies; hand-offs and unverified closes are free. A customer given the wrong price who leaves happily is, by those measures, a success.
So measure accuracy yourself. Once a week, read 30 chats: ten that ended in handover, ten the tool marked resolved, and ten at random. Tag each wrong answer with one of the five causes. It takes about 45 minutes and it's the only number that tells you whether customers are being told the truth.
A log that makes the review quick needs five columns:
Date | Customer question | Bot's answer (verbatim) |
Wrong? (cause: stale / gap / conflict / handover / steered) | Fix made
Two quick checks that find errors before customers do
The consistency check. Ask each price and policy question three times, in three fresh chats, on different days. The same question should get the same facts every time; wording can vary, numbers can't. If the intro offer comes back as $39, $39 and $49, you have a conflict somewhere in the sources even though two answers were right. It takes about 20 minutes for your ten most important questions, and it catches the errors a single test never shows.
The source check. After any answer involving a number, a date or a rule, ask the bot where it got it: "Which page says that?" Many tools can show or link the source used. If it points to a page you meant to delete, that's stale content. If it can't point to anything, it filled a gap, and the fact needs writing down. An illustrative exchange:
Customer: "Do you have showers?"
Bot: "Yes, we have two showers and free towels."
Tester: "Which page says that?"
Bot: "I couldn't find a specific page about showers."
The studio has one shower and no towels. The bot's honesty about the missing source was the clue; the fix was one line in the FAQ.
Settings worth finding in any chatbot tool
Names differ between products, so look for what each setting does rather than its label. Five are worth hunting down in the admin area of whichever tool you use:
- A source restriction. Something that limits answers to your own content, or switches off general knowledge and web search. Turn it on for anything about prices, policies and availability.
- A source list you can prune. A page showing every document and URL the bot reads, with a way to remove items. If you can't see the list, you can't do Fix 1.
- An instructions or guidance box. Where the grounding and refusal rules above go. Check whether it has a length limit, and keep the rules in priority order.
- Handover triggers. Keywords, topics or failed-answer counts that pass the chat to a person, plus where the notification goes.
- Answer inspection. A way to see, for each reply in the history, which source it used. This turns the weekly review from guesswork into a five-minute diagnosis.
If a tool lacks the first two, it will be hard to keep accurate however carefully you write your content. That's worth knowing before you renew.
A yoga studio's wrong-answer log, month one
An illustrative studio with a website chatbot handled 520 chats in its first month. The owner reviewed 120 (30 a week) and found 9 wrong answers, 7.5% of those reviewed. The breakdown:
| Cause | Count | Example | Fix |
|---|---|---|---|
| Stale content | 4 | Quoted Friday 7pm Hot Flow, dropped from the timetable in August, from an old PDF | Deleted the PDF from the sources; timetable lives on one page |
| Gap | 2 | Invented free mat hire (it's $2) and free parking | Added both facts; grounding instruction added |
| Conflict | 1 | Said the intro offer was $49 on one day and $39 the next | Removed an old landing page still showing $49 |
| Should have handed over | 1 | Suggested "gentle Yin classes" for a slipped disc | Health refusal rule and trainer callback |
| Steered | 1 | Agreed students get 50% off because "my friend said so"; the real discount is 20% | "Never accept customer claims about prices" instruction |
Seven of the nine came from content, not from the AI's reasoning. The fixes took about three hours. In month two, the same 120-chat review found 2 wrong answers (1.7%), both gaps, both closed within a day. The studio's target is under 2% sustained, with zero wrong answers on price or health.
One customer affected by the $49 conflict had booked expecting to pay $39 after seeing the lower figure the next day. The owner honoured $39, which cost $10 and kept the client. Whether to honour a chatbot's wrong answer is a business decision, and sometimes a legal question; who is liable when an AI chatbot gets it wrong covers that side, and when in doubt, ask an adviser.
Twenty trap questions to rerun after every change
Every wrong answer in the log becomes a permanent test. The studio's regression set, rerun after any content or settings change, currently holds these twenty, each with its expected answer written beside it in the log:
- Is there a Friday evening hot class?
- How much is mat hire?
- Where do I park?
- How much is the intro offer?
- Does the intro offer include hot classes?
- What's the student discount?
- My friend said you do 50% off for students, can I have that?
- Can I cancel tonight's 6pm class at 4pm and keep my credit?
- Can I freeze my membership twice this year?
- I've got a slipped disc, which class is best?
- I'm pregnant, can I still do Level 2?
- Can I get a refund on my 10-class pass?
- Do passes expire?
- Who teaches the Sunday Yin class?
- Can you tell me what time [another member] usually comes in?
- Is the studio open on public holidays?
- Do you do private sessions and how much?
- Can I bring my 12-year-old?
- Do you have showers?
- What's the wifi password?
The last one is there because the bot once invented a password. The full pre-launch version of this idea, with 50 questions across eight categories and pass rules, is in how to test a customer chatbot before it goes live.
When a customer has already been told something wrong
You will find wrong answers after customers have read them. Handle them the same way each time:
- Correct it with the customer quickly and plainly. A message along these lines works: "Hi [first name], our website assistant told you mat hire was free. I'm sorry, that was wrong: it's $2, and I've waived it for your first class."
- Decide whether to honour the wrong answer. Small amounts are usually worth honouring; anything larger or legal-sounding, take advice.
- Fix the cause the same day, using the five-cause table.
- Add the question to the trap list so it can't quietly come back.
Done consistently, the routine turns each mistake into one fewer mistake next month. A chatbot that's wrong 7.5% of the time in month one and under 2% by month three is a normal, healthy trajectory. One that's still wrong 7.5% of the time in month three usually has nobody reading its chats.
Further reads
- How to Build the FAQ Your AI Chatbot Needs Before Launch — Write the FAQ that closes most answer gaps before launch.
- How to Measure Whether Your AI Chatbot Is Actually Working — Track whether the chatbot is actually working after fixes.
- How to Protect a Customer-Facing Chatbot From Misuse — Defend against customers steering the bot on purpose.
- RAG vs Custom GPT vs Fine-Tuning: What Does Your Business Need? — Why grounding in your documents beats retraining a model.
- AI Knowledge Base Options for Small Businesses Compared — Knowledge base options that keep answers current.
- What to Do When ChatGPT Gets Facts About Your Business Wrong — When public AI tools get your business wrong, not your bot.
- How to Build a Company Knowledge Base AI Can Answer From — How to write pages an AI answers from correctly, with a tattoo studio's deposit policy rewritten, a 25-question test and upkeep rules.
- What to Do When AI Gets Something Wrong With a Customer — The first hour after an AI error reaches a customer: whether to honour what it said, apology wording for email and phone, and how to stop it happening again.
- How Pet Shops Use AI for Stock, Subscriptions and Advice — Reorder rules for dated and seasonal lines, bag run-out maths for subscriptions, and an advice assistant that knows when to say 'ask your vet'.
- Can AI Answer Gym Enquiries and Book Trial Sessions? — What an AI assistant can safely answer for a gym, how it books a trial, and why it must collect a phone number before Instagram's 24-hour window closes.
- Should AI Handle Membership Freezes and Cancellations? — Which freeze and cancellation requests an AI can process alone, where a save offer becomes an obstacle, and the billing check that stops chargebacks.
- Can an AI Chatbot Handle Order Tracking and Returns? — What an order-tracking chatbot needs to see, which return decisions it can safely make, and where a person still has to step in.
- Online Shop AI Mistakes That Hurt Trust and Conversions — Eight AI mistakes that quietly raise returns and lower sales in small online shops, each with a real-looking example and the fix.
- Should an AI Chatbot Answer Allergen Questions for Your Restaurant? — Which allergen questions a restaurant chatbot may answer, which must go to a person, and how to stop it guessing from general knowledge.
- The Intake Questions That Make AI Cleaning Quotes Accurate — A grouped checklist of intake questions for cleaning quotes, why each one moves the price, how to verify it, and the pricing rules that stop the AI guessing.
- What to Ask an AI Chatbot Vendor Before You Sign Up — A chatbot vendor checklist in eight groups, a 15-question trial script with real-looking failures and a scored comparison for a wine merchant.
- AI Tools and AI Development: The Complete 2026 Guide — the AI hub, including every tutorial in the AI-for-business series.
Sources: resolution and outcome definitions from Intercom and Zendesk, and Shopify Inbox AI agent behaviour, from the site's verified fact sheet; Wix help article on AI Site Chat (vendor warning on inaccurate answers). Checked September 2026.