How to Measure Customer Reaction After Introducing AI

Coding Liquids tutorial cover featuring Sagnik Bhattacharya for How to Measure Customer Reaction After Introducing AI.
Coding Liquids tutorial cover featuring Sagnik Bhattacharya for How to Measure Customer Reaction After Introducing AI.

Compare four signals from before and after the AI went live: a one-question satisfaction check sent after interactions the AI touched, complaints and requests for a human, customer behaviour such as rebookings and cancellations, and a weekly read of 20 real conversations. Give it four to six weeks, and trust behaviour over survey scores.

What follows is a measurement plan sized for a business with a few hundred customers rather than a few million: which signals to collect, survey wording that doesn't lead people, a prompt that sorts conversations for you, and a rule for deciding whether to keep, adjust or pull back the AI. The main caveat is small numbers. If 30 people answer your survey, two unhappy customers can move the score noticeably, so read the comments and the behaviour alongside it.

Follow me on Instagram@sagnikteaches

Pin down which customer moments the AI now touches

"We've introduced AI" is too vague to measure. Customers react to specific moments, so list every place where they now meet something an AI produced or decided. Typical touchpoints in a small business:

Connect on LinkedInSagnik Bhattacharya
  • An assistant on your website or messaging channel that answers enquiries.
  • Emails or messages drafted by AI that a person checks and sends.
  • Reminders, follow-ups or updates written or timed by AI.
  • Summaries or reports sent to customers, such as meeting notes or progress notes.
  • Phone answering or call routing.

For each one, write down three facts: roughly how many customers pass through it each week, whether they can tell AI was involved, and whether a person checks the output before it reaches them. That last point changes what you're measuring. If a colleague rewrites every AI draft, customers are mostly reacting to your colleague. If the assistant replies on its own at 11pm, they're reacting to the AI directly, and that's where to look hardest.

Subscribe on YouTube@codingliquids

Filled in for an illustrative twelve-room guesthouse, the list looked like this:

TouchpointCustomers a weekCan they tell?Checked by a person?
Website chat assistant (availability, parking, breakfast times)About 90Yes, it says soNo, replies on its own day and night
Pre-arrival emails with directions and check-in detailsAbout 40Not reallyYes, the owner sends each one
Replies to online reviewsAbout 15PossiblyYes

The chat assistant was the obvious place to measure: the most contact, visible to guests, and nobody checking it. The pre-arrival emails were left to the "didn't notice" test described further down.

Pick the one or two touchpoints with the most customer contact and measure those properly. Measuring everything thinly produces numbers too small to read.

Four signals, ranked by how far to trust them

SignalWhat it tells youWhere the data usually livesHow often to check
1. Behaviour: rebookings, renewals, repeat orders, cancellations, unsubscribesWhat customers actually did. Hardest to fake and slowest to move.Booking system, invoicing or accounts software, email platformMonthly
2. Escalations and complaints: "can I speak to someone", repeated questions, complaints that mention the new processWhere the AI creates frictionInbox, chat logs, phone notes, complaints logWeekly
3. A weekly read of 20 conversationsTone, confusion, praise and the reasons behind the numbersChat transcripts, email threads, message historyWeekly
4. A one-question surveyWhat customers say about the experienceA form tool or a reply-with-a-number messageCollected continuously, reviewed monthly

Behaviour comes first because it's what pays the bills. A customer who grumbles in a survey but rebooks is a different problem from one who says nothing and quietly leaves. Surveys come last because the people who answer them tend to be the very pleased and the very annoyed; the middle mostly stays silent.

Response time is missing from this list on purpose. Faster replies are an operational result, not a customer reaction. Customers may love the speed, or they may get a quick answer that's wrong. Track speed separately and check the quality of the work, not just how fast it arrives.

Every signal needs a "before" to compare with. If you didn't record one, rebuild it from the previous two or three months of records: most booking systems, inboxes and invoicing tools let you count past cancellations, repeat customers and complaints. The tutorial on setting a baseline before you introduce AI covers how to do this properly next time.

Survey wording that doesn't steer the answer

Ask about the customer's outcome, not about your AI. "How much did you enjoy our new AI assistant?" invites politeness and tells you nothing about whether they got what they needed. Two standard question types suit a small business:

  • Customer effort: "How easy was it to get what you needed?" scored 1 to 7. Effort is the thing AI most often makes better or worse, so this is usually the best single question.
  • Satisfaction (CSAT): "How satisfied were you with [the reply / the booking / the update]?" scored 1 to 5, reported as the share of 4s and 5s.

Net Promoter Score (the 0-10 "how likely are you to recommend us" question) measures overall loyalty and moves slowly, so it won't show you the effect of one change within six weeks. Keep it if you already run it, but don't rely on it here.

Use the same question before and after, send it within a day of the interaction, and never ask the same customer more than once a month. Wording you can copy:

After an enquiry or booking (email or text):
"Quick question about your enquiry this week: how easy was it to
get what you needed? Reply with a number from 1 (very hard) to
7 (very easy). Anything we could have done better? (optional)"

Monthly check-in for regular customers:
"Over the last month, have our messages about [appointments /
lessons / orders] been more useful than before, about the same,
or less useful? What's one thing you'd change?"

If you've told customers about the AI:
"Some of our replies are now drafted with AI and checked by
[our team / your teacher]. Has that made any difference for you?
Better / No difference / Worse / Didn't notice"

Keep "Didn't notice" as an option. For most back-office uses of AI it's the answer you're hoping for, and without it people pick "no difference" or "better" to be kind.

Sorting 20 conversations a week with an AI helper

Numbers tell you something moved; conversations tell you why. Each week, pull 20 conversations at random from the touchpoint you're measuring. Random matters: if you only read the ones a colleague flagged, you'll see a worse picture than reality, and if you only read the ones you happened to open, a better one.

Reading 20 threads takes about 30 minutes. An AI tool can do the first sort in two, provided you remove names, email addresses and phone numbers first, or use a business plan that doesn't train on your content. Then check its tags on five of the 20 yourself, because AI misreads sarcasm and polite frustration. A starting prompt:

You are helping a small business review customer conversations.
Below are [N] conversations separated by ###. Personal details
have been removed.

For each conversation, give one row with:
1. Conversation number
2. Outcome: resolved / resolved after a person stepped in / unresolved
3. Customer signals (any that apply): confused, repeated a question,
   asked for a person, frustrated, pleased, neutral
4. The customer's exact words that support each signal
   (quote, 20 words maximum)
5. What our reply got wrong, if anything (one line, or
   "nothing obvious")

Only tag a signal you can support with a direct quote. If unsure,
write "unclear". Don't infer feelings the customer didn't express.

Then total each signal and list the three most common problems
in our replies.
###
[paste conversations]

A few rows of what comes back for the guesthouse's chats (illustrative):

3 | resolved | pleased | "Great, thanks, that's all I needed" | nothing obvious
4 | unresolved | repeated a question, frustrated | "I asked about
    parking for a van, not a car" | answered for cars only
5 | resolved | pleased | "Great, thanks for nothing" | nothing obvious
6 | resolved after a person stepped in | asked for a person |
    "Can I just ring someone?" | didn't know the late check-in fee

Rows 3, 4 and 6 are right, and row 4 is the useful one: a specific gap you can fix. Row 5 is the sarcasm problem in action. "Thanks for nothing" was tagged pleased because it contains "thanks", and the "nothing obvious" in the last column is wrong too, since the guest had asked about dogs and got a reply about the garden. That's why you check five rows yourself every week. If the same misreading recurs, add a line to the prompt: "Sarcastic thanks counts as frustrated."

Log the weekly totals in a simple sheet: one row per week, one column per signal. After four weeks you'll see whether "asked for a person" or "confused" is rising, flat or falling. If you already run a sampling routine for quality, fold this into it; the weekly spot-check routine for AI support replies uses the same sample.

Worked example: a music teaching studio's AI-drafted practice notes

This one is illustrative. Consider a music teaching studio with three teachers and about 140 pupils, most of them children, so the "customer" who reacts is usually a parent. Before AI, teachers sent a practice note after some lessons, when they had time. The new process: after each lesson, the teacher records 30 seconds of spoken notes, an AI tool turns them into a short practice note, and the teacher checks it before it's emailed to the parent.

The baseline, taken from the previous term's records and one survey sent in its final week:

  • Re-enrolment for the next term: 127 of 140 pupils (91%).
  • Complaints: 2 in the term.
  • Survey question: "How clear is it what your child should practise this week?" on a 1-7 scale. Average 5.6 from 41 replies.

Six weeks in:

  • Survey average 5.9 from 37 replies. Up, but with fewer than 40 replies each time, a 0.3 shift could easily be noise.
  • The weekly read of 20 parent replies showed "confused" falling from 3 or 4 a week to 1 or 2. Good.
  • In weeks two and three, four parents used the words "generic" or "same as last week", and one asked whether the teacher had actually written the note.
  • Mid-term cancellations: 3, against 4 at the same point last term. Too few to read.

The survey on its own said "fine". The conversation read found a real problem: the AI was smoothing every note into the same shape. The studio made two changes. Teachers now say one specific thing in their recording ("bars 12 to 16, left hand only, slowly"), and the prompt tells the AI to keep that sentence word for word at the top. The studio also added a line to each note: "Drafted with AI from your teacher's comments and checked by your teacher." By week six the "generic" comments had stopped, and nobody asked again who wrote the notes.

End of term: 126 of 140 re-enrolled (90%). Essentially unchanged, which for a change aimed at clarity rather than retention is the right result: parents understood practice better and nobody left because of it.

When the signals disagree

They often will. Read the combination, not any single number:

What you seeLikely meaningWhat to do
Survey steady, requests for a person risingThe AI handles easy cases; harder ones hit a wallReview how and when the AI hands over to a person
Survey down, behaviour steadyCustomers notice and dislike something, but not enough to leave yetTreat it as an early warning; fix tone or accuracy before it shows up in cancellations
Survey up, behaviour downOnly happy customers are answeringLook at who cancelled or stopped buying and read their last conversations
Praise for speed, complaints about wrong detailsSpeed is hiding errorsTighten checking before a wrong price or date causes real damage
Everything flat, nobody mentions a changeCustomers haven't noticedFor behind-the-scenes AI, that's usually success

The third row is the one owners find hardest to believe, so here's how it looked for an illustrative meal-prep delivery business. Its one-question survey average rose slightly after it added an AI assistant to its customer messages. Subscription cancellations, meanwhile, went from about 6 a month to 11. Reading the last conversation of each customer who cancelled showed a pattern the survey never would: four of the eleven had asked the assistant how to pause deliveries for a holiday, and been given the steps for cancelling instead. The customers who answered surveys were the ones for whom it worked. The fix was a pause instruction in the assistant's reference material and a rule that any message containing "pause", "skip" or "holiday" goes to a person.

One pattern needs faster action than any table can give: a single conversation where the AI promised something you don't offer, quoted a wrong price or gave unsafe advice. Don't wait for the monthly review. The steps in what to do when AI gets something wrong with a customer apply straight away.

Keeping the before-and-after comparison honest

Small businesses are seasonal and changes pile on top of each other, so a raw before-and-after can mislead in either direction. Five habits keep it fair:

  1. Compare like with like. Set this term against last term, or this December against last December. A September baseline won't tell you much about a January result.
  2. Change one thing at a time. If you raised prices or changed your booking system in the same month, you can't credit or blame the AI for what happened.
  3. Respect small numbers. With fewer than about 30 survey replies per period, treat moves of under half a point on a 1-7 scale as noise, and look for the same direction over three periods before believing it.
  4. Don't let the vendor's dashboard mark its own homework. Chat tools often report "deflection" or "containment", meaning chats that ended without a person. A customer who gave up and phoned instead counts as contained. At the guesthouse, the chat tool reported 82% of conversations contained. Checking 50 of those "contained" chats against the phone log and inbox found 9 guests who had rung or emailed about the same thing within 48 hours, mostly about parking. Knock those out and the real figure is nearer 67%, still useful, and now pointing at one clear gap in the assistant's knowledge. The tutorial on measuring whether an AI chatbot is actually working explains how to cross-check this.
  5. Record what else changed. Keep a one-line log of anything unusual: a staff absence, a price change, a busy fortnight. When a number jumps, the log usually explains it.

Deciding what to do with what you find

At the four-to-six-week mark, apply a simple rule so the decision doesn't come down to whoever feels most strongly:

  • Keep as it is if behaviour is steady or better, requests for a person aren't rising, and no single complaint theme appears in two consecutive weekly reads.
  • Adjust if a theme shows up in three or more of the 20 conversations for two weeks running. Change one thing (the prompt, the handover rule, the disclosure line), then measure for another two weeks.
  • Pull back to person-first if rebookings or renewals fall outside their normal range, or if any conversation involved a wrong promise about price, dates or safety. Put a person in front of the AI while you fix it.

Most of the damage AI does to customer relationships comes from a small set of repeat problems: sounding canned, getting facts wrong, and making it hard to reach a person. The tutorial on AI mistakes that damage customer trust is a useful list to read your conversation tags against.

Once the numbers settle, move to a monthly check: the survey average, the behaviour figures and one read of 20 conversations. It takes under an hour and it's the difference between knowing how customers feel about your AI and assuming.

Questions owners ask about measuring customer reaction

Should I tell customers about the AI before I measure their reaction?

Measure first if you can, but don't hide it for the sake of a cleaner test. If customers talk to a chat assistant, tell them at the start of the chat. For AI drafts that a person checks and sends, a short line in your footer or on a page about how you work is usually enough. Changing disclosure mid-test changes the reaction, so note the date you added it.

What if almost nobody answers my survey?

Under about 15 replies a month, stop treating the score as a number and treat each reply as a comment. Lean harder on behaviour and on the weekly conversation read. You can raise replies by asking at the moment of contact, for example one question at the end of a chat or a reply-with-a-number text, rather than a link to a separate form.

Can I rely on the thumbs-up ratings inside my chat tool?

Use them as a supporting signal only. They're usually answered by a small, self-selected group, and customers who gave up and left never rate at all. Compare the ratings with how many chats end with a request for a person or a repeat enquiry by email or phone within a day or two.

How do I measure reaction to AI that customers never see directly?

Back-office AI, such as drafting invoices or sorting enquiries, shows up in customers' experience indirectly: faster replies, fewer errors, fewer chasers. Track complaints about mistakes and the number of times customers chase you. If those are flat or better and nobody mentions a change, customers are reacting the way you want: not noticing.

Further reads

Sources: general survey-method definitions (CSAT, Customer Effort Score, Net Promoter Score); no vendor figures are quoted in this tutorial.

Not sure whether customers like your new AI?

On a 1:1 call we'll look at where AI now touches your customers, choose the signals worth tracking for your volumes, and set up a simple weekly check your team can run.

Book a 1:1 call with me