How to Monitor AI That Talks to Customers: Hand-Offs and Errors

Coding Liquids tutorial cover featuring Sagnik Bhattacharya for How to Monitor AI That Talks to Customers: Hand-Offs and Errors.
Coding Liquids tutorial cover featuring Sagnik Bhattacharya for How to Monitor AI That Talks to Customers: Hand-Offs and Errors.

Monitor it on three clocks: live alerts for anything urgent the AI hands over, a ten-minute daily check that every hand-off reached a person and every promise was kept, and a weekly read of 20 to 30 full conversations scored for errors. Keep hand-offs and errors in one log, and fix causes rather than single replies.

Don't rely on the vendor's dashboard for this. Most tools report a resolution rate, and several count a customer going quiet as success: Intercom counts an "assumed resolution" after 24 hours of silence following Fin's last answer, and HubSpot's Customer Agent counts a resolution when there's no hand-off within 72 hours. The two failures that hurt small businesses most, the hand-off that reached nobody and the wrong answer the customer believed, look like success on both.

Follow me on Instagram@sagnikteaches

The running example is a three-surgery dental practice (an illustration) with an AI chat assistant on its website and an AI voice agent answering the phone out of hours. Between them they handle about 800 conversations a month. The same routine works for a veterinary practice, an online shop or a trades business, with different alert words.

Connect on LinkedInSagnik Bhattacharya

If your AI isn't live yet, set up this routine before launch, alongside the rules for when a chatbot should hand over to a human. Monitoring checks that those rules work in practice.

Subscribe on YouTube@codingliquids

Two failures that dashboards count as successes

Here is what "resolved" means in four common tools. Each definition decides what you're billed for, and each can hide a customer who gave up.

ToolWhen a conversation counts as resolvedWhat it can hide
Intercom Fin ($0.99 per outcome)Customer goes quiet for 24 hours after Fin's last answer ("assumed resolution")Customers who left and phoned instead
ZendeskAfter 2 hours of messaging inactivity by default (72 hours on email and web forms), plus an AI relevance checkShort, polite exits after an unhelpful answer
HubSpot Customer Agent (about $0.50 each)No hand-off to a person within 72 hoursWrong answers that never triggered a hand-off
Help Scout AI Answers ($0.75 each)Not billed if the customer escalates, searches the docs, asks again or says they need more helpCustomers who didn't do any of those, but still weren't helped

The orphaned hand-off. The AI says "I've passed this to the team, someone will call you back", and the message lands in an inbox nobody checks at the weekend, or in a queue with no owner. The customer waits. The dashboard records a successful hand-off.

The believed error. The AI gives a wrong price, a wrong opening time or a wrong policy, confidently. The customer says thanks and leaves. The dashboard records a resolution, and you find out when the customer turns up expecting what they were told. For measuring overall performance, the tutorial on measuring whether your chatbot is working has a scorecard. Monitoring is narrower: catching these two failures fast enough to fix them before they repeat.

Map the hand-off chain before you monitor it

A hand-off isn't one event. It's a chain of five links, and monitoring means checking each one:

  1. The AI recognises it should hand over.
  2. It tells the customer what happens next, accurately.
  3. The hand-off lands somewhere: an inbox, a queue, a phone, a task list.
  4. A named person picks it up within a target time.
  5. The customer gets an outcome, and the AI's promise matches what happened.

The dental practice filled in the chain for its two channels. The weak links are in the last column.

LinkWebsite chatOut-of-hours phone agentWeak point found
RecognisePain, swelling, bleeding, broken tooth, complaint, "speak to someone"Same words, plus any caller who asks twice for a personCallers who describe symptoms without the keywords
Tell the customer"The team will reply by 10am on the next working day"Urgent: emergency dentist number read out. Routine: callback next working dayChat said "within the hour" in one early version
LandShared reception inbox, tagged "AI hand-off"Voicemail transcript emailed to the same inboxTranscripts went to the practice manager's personal address for two weeks
Pick upReception lead, first thing each morningReception lead, first thing each morningNo cover when the reception lead is on holiday
OutcomeReply sent, tag closedCallback made, note addedNobody checked that tags were closed

Two of those weak points existed before anyone looked. The transcripts going to a personal address, and the early "within the hour" wording, would both have produced orphaned hand-offs that no dashboard showed.

Live alerts: the short list that interrupts someone

Keep live alerts few, or people stop reading them. Five or six conditions are enough for most small businesses; everything else waits for the daily check. The dental practice uses these:

  • A patient mentions swelling, bleeding that won't stop, a knocked-out tooth or difficulty swallowing. The AI gives the urgent-care instructions and an alert goes to the on-call phone.
  • A message contains complaint or legal words: "complaint", "solicitor", "negligence", "report you".
  • A customer asks for a person twice in one conversation.
  • The AI says it has booked, moved or cancelled an appointment, but no matching change appears in the booking system within five minutes.
  • The chat or phone agent fails to connect to the booking system, or stops answering altogether.

An illustrative alert, sent by text to whoever holds the on-call phone:

AI HAND-OFF (urgent): chat at 19:42. Patient reports swelling on left side of face since this morning, some difficulty opening mouth. AI gave emergency instructions and number. Patient's phone: [number]. Please call back or confirm they reached urgent care.

Each alert should say what the AI already told the customer. Otherwise the person calling back repeats the triage or, worse, contradicts it.

The ten-minute daily check

Every working morning, one person runs through the same five items. Here is the reception lead's check at the dental practice on a Monday, filled in:

CheckWhat she looked atMonday's result
Open hand-offs older than 12 hoursInbox filtered by the "AI hand-off" tag7 from the weekend; all answered by 9:40
Promises madeConversations where the AI said "will", "we'll" or "someone will"4 callbacks promised; 3 in the diary, 1 missing, made at 9:15
Bookings the AI madeAI booking log against the day's diary11 made, 11 in the diary, 1 in the wrong clinician's column
Customer-flagged repliesThumbs-down ratings and "that's wrong" messages2; one about opening hours (see error log)
Integration errorsThe tool's error or activity logNone

The missing callback is exactly what this check exists to find. The AI had promised a patient a call about a replacement crown, and the hand-off had gone into the inbox without the tag, because the patient had written "crown" rather than using any of the tagging words. The fix wasn't to call more carefully; it was to make every conversation containing "we'll call" create a task, whatever the topic.

A realistic mistake from before the check existed: on a Friday evening, the chat assistant told a patient "someone will call you back first thing tomorrow". The practice is closed on Saturdays. The patient phoned on Monday, annoyed, having kept Saturday morning free. Nobody had noticed the phrase because the conversation was marked resolved. The chatbot's wording now names the next working day, and the promise search in the daily check would have caught it on Saturday's first run.

The weekly read: which 25 conversations to open

Random sampling finds typical problems. Targeted sampling finds the expensive ones. Read about 25 full conversations a week, chosen like this:

  • 10 at random from all conversations.
  • 5 that ended in a hand-off.
  • 5 marked resolved that lasted under a minute, where people often leave after an unhelpful first answer.
  • 5 that contain negative words such as "wrong", "useless", "not what I asked" or "ridiculous".
  • Every complaint, whatever the count.

A quick sum on the effort: at about 90 seconds per short conversation and 3 minutes per long one, 25 conversations take 45 to 60 minutes. At 800 conversations a month, that's roughly one in eight read every month, enough to see patterns within two or three weeks.

Score each conversation on five lines: facts correct, action correct, hand-off correct (made when needed, not made when it wasn't), tone acceptable, and no personal data shared that shouldn't have been. A single "no" on any line goes into the error log.

You can use AI to pre-screen transcripts, as long as a person makes the call. This prompt works in ChatGPT, Claude or Copilot Chat on a business plan with transcripts that have had names and numbers removed:

Below are 40 chat transcripts from a dental practice's website assistant,
names and numbers removed. The practice's current facts are in FACTS.
For each transcript, return one line:
ID | problem type (wrong fact, wrong action, missed hand-off, unnecessary
hand-off, tone, personal data, none) | the exact sentence that shows it
Only report a problem if you can quote the sentence. If unsure, write "check".
FACTS: [opening hours, prices list, policies]
TRANSCRIPTS: [pasted]

An illustrative slice of the output:

T07 | wrong fact | "We're open until 7pm on Thursdays." (FACTS: until 6pm)
T12 | none |
T19 | missed hand-off | "I understand the pain is getting worse. Here are our opening hours."
T23 | check | "The hygienist appointment is $85." (FACTS lists two hygiene prices)

The pre-screen is a filter, not a verdict. In this run a person confirmed T07 and T19 as real errors, and T23 turned out correct for a standard appointment. For voice agents, the same scoring works on call transcripts; the tutorial on reviewing AI call transcripts covers the audio-specific checks.

An error log that points at causes

Log every error with its type, the cause you found, the fix and who owns it. The first two weeks at the dental practice produced this, lightly abridged:

DateTypeWhat happenedCauseFix and owner
Mon 3rdWrong factTold a patient "open until 7pm Thursdays"Old late-opening page still in the knowledge basePage removed; practice manager
Tue 4thMissed hand-offWorsening pain answered with opening hours"Worse" and "getting worse" not in trigger listTriggers widened; reception lead
Thu 6thWrong actionCheck-up booked with a clinician on leaveLeave not blocked in the booking systemLeave now blocked in the diary; practice manager
Fri 7thToneReplied "Unfortunately that isn't possible" to a bereaved patient cancellingNo instruction for bereavementHand-off rule added for bereavement; practice manager
Wed 12thUnnecessary hand-offAsked for a person over a parking questionParking info missingParking page added; reception lead

Notice that only one of the five causes was the AI misbehaving on its own. The rest were stale content, missing triggers and a diary that didn't match reality. That's typical, and it's why the log records causes, not only symptoms. The tutorial on stopping a chatbot giving wrong answers goes deeper into fixing content problems once the log has found them.

Thresholds for pulling the AI off a topic

Monitoring only helps if it can change what the AI is allowed to do. Agree in advance, in writing, what happens when the log shows a pattern, so nobody has to argue about it on a busy Monday. The dental practice uses four rules:

If the log showsThenUntil
Any missed urgent hand-off (pain, swelling, bleeding)Out-of-hours agent switches to "take a message and give the emergency number" onlyThe trigger fix passes the 20-question retest twice
Two wrong prices or fees in one weekAll price questions hand over to receptionThe price list in the knowledge base is rebuilt from the current fee sheet
A booking made in the wrong diary twice in a weekAI takes booking requests but a person confirms themA week of correct bookings in shadow mode
Personal data shared with the wrong person, onceAI stops discussing existing appointments at allThe practice manager has found the cause and agreed the fix

Narrowing the AI's job is almost always better than switching it off altogether. Customers still get instant answers on the topics that work, and the risky topic goes back to people while you fix it. Most chat and voice tools let you do this by editing the instructions or the hand-off triggers, so it takes minutes rather than a support ticket.

Phone agents need two extra signals that chat doesn't. Watch for calls where the caller hangs up within 20 seconds of the AI answering, which often means they didn't want to talk to a machine, and for calls where the caller says "operator", "person" or "human" more than once. A rising count of either is an early sign that callers are giving up, weeks before it shows up as fewer bookings.

The practice's numbers after six weeks

These figures are illustrative but realistic for a practice of this size. In the first week of monitoring, the daily check found six orphaned hand-offs, mostly from the weekend and the misrouted transcripts, and the weekly read found nine errors in 25 conversations. By week six:

  • Orphaned hand-offs: zero for three weeks running, because every "we'll call" creates a task.
  • Errors in the weekly sample: two in 25, both tone rather than fact.
  • Hand-off rate: from 14% of conversations to 11%, mostly because missing information (parking, payment plans) was added.
  • Time spent: about 10 minutes a day and 50 minutes a week, roughly five hours a month shared between two people.

The practice judged five hours a month a fair price for knowing, rather than hoping, that 800 conversations were going well.

Retest after every change, including the vendor's

Every change can break something that worked. Keep a set of 20 test questions, including your nastiest real ones, and rerun them after you edit the knowledge base, change the instructions, or the vendor announces a model or product update. The dental practice's set includes "My face is swollen, what do I do?", "Can I bring my child to my appointment?", "How much is a filling?", "I want to make a complaint" and "Is the car park free?".

Vendors change things too, and not only models. Pricing definitions, retention periods and privacy defaults move. SimplePractice, for example, changed its Note Taker so that from 16 June 2026 new users are opted in by default to keeping de-identified transcripts. Put a monthly reminder in the diary to read your AI vendor's release notes and terms page, and check that the disclosure telling customers they're talking to AI still appears. If you serve customers in the EU, the EU AI Act's Article 50 duty to tell people they're dealing with a chatbot has applied since 2 August 2026.

When a wrong answer has already reached the customer

Contact the customer as soon as the log finds it, correct the information plainly, and decide whether to honour what the AI said. The before and after below shows the difference tone makes.

Before (the first draft a receptionist wrote): "Our chatbot made an error regarding our Thursday opening hours. The correct closing time is 6pm. Apologies for any inconvenience."

After: "Hi [first name], you asked our website assistant about Thursday evenings and it told you we're open until 7pm. That was wrong, and I'm sorry: we close at 6pm. Your 6:30pm request hasn't been booked. I can offer 5:15pm this Thursday or 6:15pm next Tuesday, when we do open late. Just reply with the one that suits."

The second version says exactly what went wrong, doesn't hide behind "the chatbot", and fixes the customer's actual problem in the same message. The tutorial on what to do when AI gets something wrong with a customer covers when to honour a price the AI quoted.

Who does what in a team of five to fifteen

Monitoring fails when it belongs to everyone. Name people for four jobs, and a backup for each:

  • Daily checker: usually the reception or customer-service lead. Ten minutes, first thing.
  • Weekly reviewer: the practice or office manager. About an hour, same day each week.
  • Fix owner: whoever maintains the knowledge base and instructions. Fixes within a week, logs what changed.
  • Escalation owner: for a clinic, a clinician on call for anything clinical; for other businesses, the owner for complaints and refunds.

Write the holiday cover into the rota. The dental practice's worst week before monitoring was the reception lead's holiday, when hand-offs arrived as normal and nobody was looking for them.

Further reads

Sources: Intercom Fin pricing and resolution definition; Zendesk automated resolution documentation; HubSpot Customer Agent knowledge base; Help Scout AI Answers pricing; SimplePractice Note Taker notice; EU AI Act Article 50 guidance.

Want a monitoring routine for your customer-facing AI?

On a 1:1 call we'll map where your chatbot or AI phone agent hands customers over, find the links in that chain that can fail silently, and set up checks your team can run in minutes a day.

Book a 1:1 call with me