Monitor it on three clocks: live alerts for anything urgent the AI hands over, a ten-minute daily check that every hand-off reached a person and every promise was kept, and a weekly read of 20 to 30 full conversations scored for errors. Keep hand-offs and errors in one log, and fix causes rather than single replies.
Don't rely on the vendor's dashboard for this. Most tools report a resolution rate, and several count a customer going quiet as success: Intercom counts an "assumed resolution" after 24 hours of silence following Fin's last answer, and HubSpot's Customer Agent counts a resolution when there's no hand-off within 72 hours. The two failures that hurt small businesses most, the hand-off that reached nobody and the wrong answer the customer believed, look like success on both.
The running example is a three-surgery dental practice (an illustration) with an AI chat assistant on its website and an AI voice agent answering the phone out of hours. Between them they handle about 800 conversations a month. The same routine works for a veterinary practice, an online shop or a trades business, with different alert words.
If your AI isn't live yet, set up this routine before launch, alongside the rules for when a chatbot should hand over to a human. Monitoring checks that those rules work in practice.
Two failures that dashboards count as successes
Here is what "resolved" means in four common tools. Each definition decides what you're billed for, and each can hide a customer who gave up.
| Tool | When a conversation counts as resolved | What it can hide |
|---|---|---|
| Intercom Fin ($0.99 per outcome) | Customer goes quiet for 24 hours after Fin's last answer ("assumed resolution") | Customers who left and phoned instead |
| Zendesk | After 2 hours of messaging inactivity by default (72 hours on email and web forms), plus an AI relevance check | Short, polite exits after an unhelpful answer |
| HubSpot Customer Agent (about $0.50 each) | No hand-off to a person within 72 hours | Wrong answers that never triggered a hand-off |
| Help Scout AI Answers ($0.75 each) | Not billed if the customer escalates, searches the docs, asks again or says they need more help | Customers who didn't do any of those, but still weren't helped |
The orphaned hand-off. The AI says "I've passed this to the team, someone will call you back", and the message lands in an inbox nobody checks at the weekend, or in a queue with no owner. The customer waits. The dashboard records a successful hand-off.
The believed error. The AI gives a wrong price, a wrong opening time or a wrong policy, confidently. The customer says thanks and leaves. The dashboard records a resolution, and you find out when the customer turns up expecting what they were told. For measuring overall performance, the tutorial on measuring whether your chatbot is working has a scorecard. Monitoring is narrower: catching these two failures fast enough to fix them before they repeat.
Map the hand-off chain before you monitor it
A hand-off isn't one event. It's a chain of five links, and monitoring means checking each one:
- The AI recognises it should hand over.
- It tells the customer what happens next, accurately.
- The hand-off lands somewhere: an inbox, a queue, a phone, a task list.
- A named person picks it up within a target time.
- The customer gets an outcome, and the AI's promise matches what happened.
The dental practice filled in the chain for its two channels. The weak links are in the last column.
| Link | Website chat | Out-of-hours phone agent | Weak point found |
|---|---|---|---|
| Recognise | Pain, swelling, bleeding, broken tooth, complaint, "speak to someone" | Same words, plus any caller who asks twice for a person | Callers who describe symptoms without the keywords |
| Tell the customer | "The team will reply by 10am on the next working day" | Urgent: emergency dentist number read out. Routine: callback next working day | Chat said "within the hour" in one early version |
| Land | Shared reception inbox, tagged "AI hand-off" | Voicemail transcript emailed to the same inbox | Transcripts went to the practice manager's personal address for two weeks |
| Pick up | Reception lead, first thing each morning | Reception lead, first thing each morning | No cover when the reception lead is on holiday |
| Outcome | Reply sent, tag closed | Callback made, note added | Nobody checked that tags were closed |
Two of those weak points existed before anyone looked. The transcripts going to a personal address, and the early "within the hour" wording, would both have produced orphaned hand-offs that no dashboard showed.
Live alerts: the short list that interrupts someone
Keep live alerts few, or people stop reading them. Five or six conditions are enough for most small businesses; everything else waits for the daily check. The dental practice uses these:
- A patient mentions swelling, bleeding that won't stop, a knocked-out tooth or difficulty swallowing. The AI gives the urgent-care instructions and an alert goes to the on-call phone.
- A message contains complaint or legal words: "complaint", "solicitor", "negligence", "report you".
- A customer asks for a person twice in one conversation.
- The AI says it has booked, moved or cancelled an appointment, but no matching change appears in the booking system within five minutes.
- The chat or phone agent fails to connect to the booking system, or stops answering altogether.
An illustrative alert, sent by text to whoever holds the on-call phone:
AI HAND-OFF (urgent): chat at 19:42. Patient reports swelling on left side of face since this morning, some difficulty opening mouth. AI gave emergency instructions and number. Patient's phone: [number]. Please call back or confirm they reached urgent care.
Each alert should say what the AI already told the customer. Otherwise the person calling back repeats the triage or, worse, contradicts it.
The ten-minute daily check
Every working morning, one person runs through the same five items. Here is the reception lead's check at the dental practice on a Monday, filled in:
| Check | What she looked at | Monday's result |
|---|---|---|
| Open hand-offs older than 12 hours | Inbox filtered by the "AI hand-off" tag | 7 from the weekend; all answered by 9:40 |
| Promises made | Conversations where the AI said "will", "we'll" or "someone will" | 4 callbacks promised; 3 in the diary, 1 missing, made at 9:15 |
| Bookings the AI made | AI booking log against the day's diary | 11 made, 11 in the diary, 1 in the wrong clinician's column |
| Customer-flagged replies | Thumbs-down ratings and "that's wrong" messages | 2; one about opening hours (see error log) |
| Integration errors | The tool's error or activity log | None |
The missing callback is exactly what this check exists to find. The AI had promised a patient a call about a replacement crown, and the hand-off had gone into the inbox without the tag, because the patient had written "crown" rather than using any of the tagging words. The fix wasn't to call more carefully; it was to make every conversation containing "we'll call" create a task, whatever the topic.
A realistic mistake from before the check existed: on a Friday evening, the chat assistant told a patient "someone will call you back first thing tomorrow". The practice is closed on Saturdays. The patient phoned on Monday, annoyed, having kept Saturday morning free. Nobody had noticed the phrase because the conversation was marked resolved. The chatbot's wording now names the next working day, and the promise search in the daily check would have caught it on Saturday's first run.
The weekly read: which 25 conversations to open
Random sampling finds typical problems. Targeted sampling finds the expensive ones. Read about 25 full conversations a week, chosen like this:
- 10 at random from all conversations.
- 5 that ended in a hand-off.
- 5 marked resolved that lasted under a minute, where people often leave after an unhelpful first answer.
- 5 that contain negative words such as "wrong", "useless", "not what I asked" or "ridiculous".
- Every complaint, whatever the count.
A quick sum on the effort: at about 90 seconds per short conversation and 3 minutes per long one, 25 conversations take 45 to 60 minutes. At 800 conversations a month, that's roughly one in eight read every month, enough to see patterns within two or three weeks.
Score each conversation on five lines: facts correct, action correct, hand-off correct (made when needed, not made when it wasn't), tone acceptable, and no personal data shared that shouldn't have been. A single "no" on any line goes into the error log.
You can use AI to pre-screen transcripts, as long as a person makes the call. This prompt works in ChatGPT, Claude or Copilot Chat on a business plan with transcripts that have had names and numbers removed:
Below are 40 chat transcripts from a dental practice's website assistant,
names and numbers removed. The practice's current facts are in FACTS.
For each transcript, return one line:
ID | problem type (wrong fact, wrong action, missed hand-off, unnecessary
hand-off, tone, personal data, none) | the exact sentence that shows it
Only report a problem if you can quote the sentence. If unsure, write "check".
FACTS: [opening hours, prices list, policies]
TRANSCRIPTS: [pasted]
An illustrative slice of the output:
T07 | wrong fact | "We're open until 7pm on Thursdays." (FACTS: until 6pm)
T12 | none |
T19 | missed hand-off | "I understand the pain is getting worse. Here are our opening hours."
T23 | check | "The hygienist appointment is $85." (FACTS lists two hygiene prices)
The pre-screen is a filter, not a verdict. In this run a person confirmed T07 and T19 as real errors, and T23 turned out correct for a standard appointment. For voice agents, the same scoring works on call transcripts; the tutorial on reviewing AI call transcripts covers the audio-specific checks.
An error log that points at causes
Log every error with its type, the cause you found, the fix and who owns it. The first two weeks at the dental practice produced this, lightly abridged:
| Date | Type | What happened | Cause | Fix and owner |
|---|---|---|---|---|
| Mon 3rd | Wrong fact | Told a patient "open until 7pm Thursdays" | Old late-opening page still in the knowledge base | Page removed; practice manager |
| Tue 4th | Missed hand-off | Worsening pain answered with opening hours | "Worse" and "getting worse" not in trigger list | Triggers widened; reception lead |
| Thu 6th | Wrong action | Check-up booked with a clinician on leave | Leave not blocked in the booking system | Leave now blocked in the diary; practice manager |
| Fri 7th | Tone | Replied "Unfortunately that isn't possible" to a bereaved patient cancelling | No instruction for bereavement | Hand-off rule added for bereavement; practice manager |
| Wed 12th | Unnecessary hand-off | Asked for a person over a parking question | Parking info missing | Parking page added; reception lead |
Notice that only one of the five causes was the AI misbehaving on its own. The rest were stale content, missing triggers and a diary that didn't match reality. That's typical, and it's why the log records causes, not only symptoms. The tutorial on stopping a chatbot giving wrong answers goes deeper into fixing content problems once the log has found them.
Thresholds for pulling the AI off a topic
Monitoring only helps if it can change what the AI is allowed to do. Agree in advance, in writing, what happens when the log shows a pattern, so nobody has to argue about it on a busy Monday. The dental practice uses four rules:
| If the log shows | Then | Until |
|---|---|---|
| Any missed urgent hand-off (pain, swelling, bleeding) | Out-of-hours agent switches to "take a message and give the emergency number" only | The trigger fix passes the 20-question retest twice |
| Two wrong prices or fees in one week | All price questions hand over to reception | The price list in the knowledge base is rebuilt from the current fee sheet |
| A booking made in the wrong diary twice in a week | AI takes booking requests but a person confirms them | A week of correct bookings in shadow mode |
| Personal data shared with the wrong person, once | AI stops discussing existing appointments at all | The practice manager has found the cause and agreed the fix |
Narrowing the AI's job is almost always better than switching it off altogether. Customers still get instant answers on the topics that work, and the risky topic goes back to people while you fix it. Most chat and voice tools let you do this by editing the instructions or the hand-off triggers, so it takes minutes rather than a support ticket.
Phone agents need two extra signals that chat doesn't. Watch for calls where the caller hangs up within 20 seconds of the AI answering, which often means they didn't want to talk to a machine, and for calls where the caller says "operator", "person" or "human" more than once. A rising count of either is an early sign that callers are giving up, weeks before it shows up as fewer bookings.
The practice's numbers after six weeks
These figures are illustrative but realistic for a practice of this size. In the first week of monitoring, the daily check found six orphaned hand-offs, mostly from the weekend and the misrouted transcripts, and the weekly read found nine errors in 25 conversations. By week six:
- Orphaned hand-offs: zero for three weeks running, because every "we'll call" creates a task.
- Errors in the weekly sample: two in 25, both tone rather than fact.
- Hand-off rate: from 14% of conversations to 11%, mostly because missing information (parking, payment plans) was added.
- Time spent: about 10 minutes a day and 50 minutes a week, roughly five hours a month shared between two people.
The practice judged five hours a month a fair price for knowing, rather than hoping, that 800 conversations were going well.
Retest after every change, including the vendor's
Every change can break something that worked. Keep a set of 20 test questions, including your nastiest real ones, and rerun them after you edit the knowledge base, change the instructions, or the vendor announces a model or product update. The dental practice's set includes "My face is swollen, what do I do?", "Can I bring my child to my appointment?", "How much is a filling?", "I want to make a complaint" and "Is the car park free?".
Vendors change things too, and not only models. Pricing definitions, retention periods and privacy defaults move. SimplePractice, for example, changed its Note Taker so that from 16 June 2026 new users are opted in by default to keeping de-identified transcripts. Put a monthly reminder in the diary to read your AI vendor's release notes and terms page, and check that the disclosure telling customers they're talking to AI still appears. If you serve customers in the EU, the EU AI Act's Article 50 duty to tell people they're dealing with a chatbot has applied since 2 August 2026.
When a wrong answer has already reached the customer
Contact the customer as soon as the log finds it, correct the information plainly, and decide whether to honour what the AI said. The before and after below shows the difference tone makes.
Before (the first draft a receptionist wrote): "Our chatbot made an error regarding our Thursday opening hours. The correct closing time is 6pm. Apologies for any inconvenience."
After: "Hi [first name], you asked our website assistant about Thursday evenings and it told you we're open until 7pm. That was wrong, and I'm sorry: we close at 6pm. Your 6:30pm request hasn't been booked. I can offer 5:15pm this Thursday or 6:15pm next Tuesday, when we do open late. Just reply with the one that suits."
The second version says exactly what went wrong, doesn't hide behind "the chatbot", and fixes the customer's actual problem in the same message. The tutorial on what to do when AI gets something wrong with a customer covers when to honour a price the AI quoted.
Who does what in a team of five to fifteen
Monitoring fails when it belongs to everyone. Name people for four jobs, and a backup for each:
- Daily checker: usually the reception or customer-service lead. Ten minutes, first thing.
- Weekly reviewer: the practice or office manager. About an hour, same day each week.
- Fix owner: whoever maintains the knowledge base and instructions. Fixes within a week, logs what changed.
- Escalation owner: for a clinic, a clinician on call for anything clinical; for other businesses, the owner for complaints and refunds.
Write the holiday cover into the rota. The dental practice's worst week before monitoring was the reception lead's holiday, when hand-offs arrived as normal and nobody was looking for them.
Further reads
- How to Write Call Scripts and Escalation Rules for an AI Receptionist — Write the escalation rules your monitoring will check.
- How to Test a Customer Chatbot Before It Goes Live — Build the question set you'll rerun after changes.
- AI Chatbot Disclosure: What to Tell Customers at the Start of a Chat — Wording for telling customers they're talking to AI.
- AI Incident Response Plan for Small Businesses (With Template) — A template for the rare serious failure.
- Chatbot Guardrails: Stop AI Promising What You Don't Offer — Stop the promises before you have to catch them.
- How a Small Clinic Can Use AI at the Front Desk Safely — Where monitoring fits in a clinic's front-desk setup.
- How to Set Up Human Review for AI Work Without Slowing Down — Four levels of human review matched to risk, how to make each check take under a minute, how many to sample, and when to relax or tighten.
- Customer-Facing or Back-Office: Where Should AI Go First? — A side-by-side comparison of customer-facing and back-office AI as a first move, with a scoring sheet, the middle route, and two worked decisions.
- How Insurance Brokers Use AI to Handle Claims Enquiries — A claims-enquiry workflow for small brokers: AI structures notifications, flags urgency and drafts updates, while coverage answers stay with the broker.
- How to Stop Zapier and Make Automations Breaking Silently — Four layers of protection against quiet failures: notifications, error handlers that alert someone, heartbeat and volume checks, and validation of AI outputs.
- Should a Custom GPT Answer Your Customers? Limits and Safer Options — Customers must sign in to use a GPT, you can't read the chats, and there's no handover. Here are the chatbot options built for customers, with costs.
- How to Pilot Your First AI Agent Without Risking Customers — Run your first AI agent in shadow mode, then behind approvals, then live in a narrow window. A café's catering agent shows each stage, cost and stop rule.
- Done-for-You vs Done-With-You AI Implementation: Which Suits You? — What done-for-you and done-with-you AI implementation leave you holding, how upkeep decides it, and a four-shop butcher that uses both.
- What Is an AI Audit? What It Covers and What It Costs — The four jobs sold as an AI audit, what a usage and risk audit checks, a food truck audited in an afternoon, and how audit fees are built up.
- How to Hire Your First Customer Service Person Alongside AI — Split enquiries between AI and a person, work out the hours left, write the job description, and test candidates on escalations and wrong bot answers.
- AI Consultant Handover Checklist: What You Need Before They Leave — Everything an AI consultant should hand over before they leave, grouped into a checklist with how to verify each item, a filled runbook and a scored example.
- AI Tools and AI Development: The Complete 2026 Guide — the AI hub, including every tutorial in the AI-for-business series.
Sources: Intercom Fin pricing and resolution definition; Zendesk automated resolution documentation; HubSpot Customer Agent knowledge base; Help Scout AI Answers pricing; SimplePractice Note Taker notice; EU AI Act Article 50 guidance.