Track five things monthly: the real resolution rate (vendor-reported resolutions minus customers who came back by phone or email within seven days), accuracy from a weekly sample of 20-30 conversations, how fast humans pick up handovers, cost per genuinely solved chat, and one business result such as bookings. Compare them with your pre-launch baseline.
Vendor dashboards flatter chatbots because "resolved" usually means "nobody asked for a human". Intercom treats a conversation as an assumed resolution if the customer goes quiet for 24 hours after Fin's last answer; Zendesk closes a conversation once it has gone quiet (two hours by default on messaging, 72 hours on email) and, since 18 May 2026, bills only the resolutions its AI check verifies; hand-offs and unverified closes are free. A customer who gave up and rang a competitor looks, on those dashboards, like a success.
Why the resolution rate on your dashboard runs high
The vendors aren't hiding anything; their definitions are published, and some are careful. Zendesk runs a second AI review before billing a resolution, and Intercom deducts a resolution if the customer comes back to the same conversation. But none of them can see what happened outside the chat window: the phone call ten minutes later, the email to the owner's address, the booking that went elsewhere. And they bill on that same definition, so a generous count also means a bigger invoice on per-resolution pricing.
Three kinds of conversation inflate the number:
- Gave up. The answer was vague or wrong, the customer closed the tab, and silence was scored as success.
- Went elsewhere. The customer got an answer but didn't trust it, and phoned to check.
- Wrong but accepted. The bot gave an incorrect price or policy, the customer believed it, and the problem surfaces later as a dispute.
The measures below exist to catch each of those. None needs special software: a conversation export, your phone log and an hour a week.
The five-line scorecard
| Measure | How to calculate it | Where the data comes from | Warning sign |
|---|---|---|---|
| Real resolution rate | (Vendor resolutions minus came-back-anyway minus sampled wrong answers) ÷ conversations with a genuine question | Chatbot export, phone log, inbox, your weekly sample | Falling for two months, or more than 15 points below the vendor figure |
| Accuracy | Correct answers ÷ answers graded in the weekly sample | 20-30 random conversations a week | Any wrong price, policy or safety answer; under 90% overall |
| Handover pick-up time | Median time from the bot handing over to a human's first reply, split in hours and out of hours | Help desk or inbox timestamps | Longer than the time the bot told the customer to expect |
| Cost per genuine resolution | Monthly chatbot cost ÷ real resolutions | Invoice plus the calculation above | Higher than the cost of a person answering the same chats |
| Business result | Bookings, qualified leads or orders that started in chat | Booking system or CRM, tagged by source | Flat or down compared with the pre-launch baseline |
If you didn't record numbers before launch, reconstruct what you can from the month before: enquiry count, how many were repeat questions, how long replies took and how many bookings came from the website. The tutorial on setting a baseline before introducing AI covers the method.
Set the bot up so it can be measured at all
Half the measures above are impossible if the chatbot was switched on with default settings. Four changes, each taking a few minutes, make the numbers collectable:
- Ask for a phone number or email before booking or handing over. Without contact details, the came-back-anyway check has nothing to match on. Ask only at the point it's genuinely needed ("so the team can reply"), not as a gate at the start, which drives people away.
- Tag conversations by topic. Most chatbot tools and help desks let you tag automatically or with a rule: price, booking, existing job, complaint, not offered. Five to eight tags is plenty. Topic tags turn "accuracy was 91%" into "accuracy on prices was 78%", which tells you what to fix.
- Mark where bookings came from. If a chat leads to a booking, the booking record needs a source field that says "chat". In a booking system or CRM that's a dropdown; in a diary it's a letter in the entry. Without it, the business-result line is a guess.
- Record handover times. The handover should create an item in your help desk or shared inbox with a timestamp, so the gap between the bot passing it on and a person replying can be measured, rather than remembered.
A surveying firm that skipped step 3 spent its first review meeting arguing about whether the bot had produced any bookings at all. Adding one dropdown settled it the following month.
Running the came-back-anyway check
This is the single most revealing check, and the one almost nobody runs. It answers one question: of the chats the vendor counts as resolved, how many customers contacted you again about the same thing within seven days?
- Export the resolved conversations for the month, with any phone number or email the customer gave and the date. Most tools export to CSV from their reporting or conversations screen.
- Export your other contact logs for the same period plus seven days: the call log from your phone system, and received emails from the enquiries inbox.
- Match them. In a spreadsheet, put the chat contacts in one column and the call and email contacts in another, and flag any chat contact that appears in the other log within seven days. A
COUNTIFSon phone number and date range does it; so does pasting both lists into an AI assistant with a clear instruction. - Read the matches. Not every match is a failure: someone who chatted about prices and then rang to book is a success. Mark each as "booked", "same question again" or "complaint".
Only "same question again" and "complaint" come off the resolution count. Chats where the customer gave no contact details can't be matched, so the check undercounts; treat the result as a floor, not a precise figure.
If you use an AI assistant for step 3, keep the instruction narrow and check its matches by hand:
Below are two lists. List A: chatbot conversations marked resolved, with date,
phone or email, and the customer's first message. List B: phone calls and emails
received, with date, number or address, and a one-line note where available.
For every entry in List A, find entries in List B with the same phone number or
email address dated 0-7 days later. Output a table: List A date, contact, first
message, List B date, List B note. Do not guess matches on name alone.
If there is no match, leave the row out.
Grading a weekly sample of conversations
Pick 20-30 conversations at random each week, plus every one with a thumbs-down or a complaint. Grade each on one of four levels:
- Correct and complete: right answer, right conditions, sensible next step.
- Correct but incomplete: nothing false, but a condition or next step was missing.
- Wrong: stated something untrue about price, policy, availability or safety.
- Should have handed over: answered something it should have passed to a person.
Here's an illustrative conversation from a two-van locksmith's website chat, received at 10.40pm:
Customer: how much to get into my flat, locked out, keys inside
Bot: Our call-out charge is $85 and most lockouts are opened without
damage. Would you like us to book a visit?
Customer: yes asap
Bot: Thanks. Please leave your name, number and address and we'll
confirm a time.
The vendor counted that as resolved: the customer didn't ask for a human. Graded properly, it's wrong twice. $85 is the daytime call-out figure; after 8pm the firm charges more, and that condition was in the price entry but the bot dropped it. And a night-time lockout should never end with "we'll confirm a time": the rule was to hand straight to the on-call phone and tell the customer to expect a call within 15 minutes. The fix was a tighter price entry and a routing rule on the words "locked out" after hours. Without the sample, the firm would have learned about it from an angry customer disputing the bill at the door. The steps for fixing answers like this are in stopping a chatbot giving customers wrong answers.
Keep a running tally in a simple sheet, one row per graded conversation, with the grade, the topic and a note on the cause. After a month, sort by topic. Errors cluster, and the cluster tells you which content to rewrite.
A locksmith's month, measured properly
The two-van locksmith's first full month, with illustrative figures, shows how far the dashboard and reality can drift apart.
- Conversations with a genuine question: 410. The dashboard showed 279 resolved, a 68% resolution rate.
- Came-back-anyway check: 52 of those 279 customers rang or emailed within seven days with the same question or a complaint. That leaves 227.
- Weekly samples: 100 resolved conversations graded across the month, 9 of them wrong. Applying that 9% to the remaining 227 removes about 20 more. That leaves roughly 207.
- Real resolution rate: 207 ÷ 410, about 50%. Eighteen points below the dashboard.
Cost next. On a per-resolution plan at $0.99, the firm paid for the vendor's 279 resolutions: $276. Divided by 207 genuine resolutions, that's about $1.33 each. The owner's estimate of answering a chat himself was four minutes at around $22 an hour including overheads, so about $1.47. On cost alone, a near tie.
The business result settled it. Forty-six jobs were booked from chat in the month, 31 of them between 8pm and 8am, when nobody would otherwise have answered the website. And the handover numbers showed the weak spot: in working hours a person picked up handovers in a median of 25 minutes, but out-of-hours non-emergency handovers waited a median of 11 hours, while the bot was telling customers "someone will be in touch shortly". The change that month was one sentence in the handover message, giving the honest time ("we'll reply by 9am"), and a routing rule sending lockouts straight to the on-call phone.
Reading what the chatbot couldn't answer
Most tools keep a list of questions the bot couldn't answer or handed over. It's the best source of new content, but a raw list of 130 questions is hard to act on. An AI assistant can cluster it:
Here are the customer questions our website chatbot couldn't answer last month.
Group them into topics. For each topic give: a short name, the number of
questions, two example questions in the customer's words, and whether it looks
like (a) missing content we could write, (b) something needing a human, or
(c) something we don't offer. Sort by number of questions.
An illustrative result for the locksmith:
1. Car key replacement (23) - "do you do car keys", "lost my only car key"
-> (c) not offered. Add an entry saying so and suggest an auto locksmith.
2. Lock upgrade for insurance (17) - "what lock do I need for my insurance",
"is a 5 lever lock ok" -> (a) missing content, needs a careful answer.
3. Price for a specific job (15) - "how much to change 3 locks" -> (b) human,
but could give a per-lock range.
4. Are you open now (9) - "are you open", "can someone come now"
-> (a) missing: hours and night cover aren't in the content.
What you'd check before acting on it: that the counts add up to roughly the list's length, and that the example questions really belong to their groups. Clustering tools sometimes fold two different topics together. Here, "are you open now" deserved priority despite its lower count, because every one of those customers wanted work that night.
What to measure when the bot's job is bookings or leads
Not every chatbot exists to answer questions. For many small firms the real job is capturing enquiries, so the business-result line of the scorecard matters most.
A six-person surveying firm used its bot mainly to explain survey types and take booking requests. The measure that mattered was booking requests per 100 conversations, and how many became paid surveys. Before the bot, the website contact form produced about 30 requests a month; after two months the bot produced 44, of which 29 booked. The owner also tracked something dashboards don't: how many bot-captured requests lacked the property value or the buyer's deadline, which the surveyor then had to chase by phone. It started at one in three and fell to one in eight after the bot's questions were reordered.
A kitchen fitter measured lead quality rather than volume. Design visits booked from chat rose, but so did visits where the customer's budget turned out far below any kitchen the firm sold. Adding a budget-band question to the bot, with the firm's typical price ranges stated plainly, cut the wasted visits. The volume number looked worse; the business result improved.
Things that quietly distort the numbers
Before trusting any month's figures, rule out the usual distortions. Each has fooled a small business owner into a wrong conclusion.
- Your own team's tests. Staff checking the bot after a content change can add dozens of conversations. Exclude your own devices or tag test chats, or the accuracy and resolution figures include your experiments.
- Spam and automated traffic. Chat widgets attract bots selling marketing services. They usually leave after one message and can be scored as resolved. Filter out conversations with no genuine question before calculating anything, which is why the scorecard's denominator is "conversations with a genuine question".
- Existing customers checking on jobs. "When's my engineer coming?" isn't something most chatbots can answer, so it drags the resolution rate down. Track it as its own topic; it may point to a job-status message you could send automatically instead.
- Seasonal swings. A heating installer's chat volume and topics in the first cold week look nothing like midsummer. Compare against the same month's baseline where you can, or at least note the season beside each monthly report.
- A content change mid-month. If you rewrote the price entries on the 14th, the month's accuracy is an average of two different bots. Note the date of every change in the report so the next month's figures can be read properly.
The measuring itself needs a budget of time. For the locksmith it settled at about an hour a week: 30 minutes grading the sample, 15 on the unanswered list and 15 on the came-back check, done monthly in one sitting. That hour is what keeps a bot answering customers correctly; without it, content drifts and errors surface as complaints instead.
A one-page monthly report
Put the five numbers in the same format every month, so trends are visible and the conversation about the bot is about evidence. A filled-in example for the locksmith's second month:
CHATBOT REPORT - Month 2 (Month 1)
Conversations with a genuine question 438 (410)
Vendor-reported resolution rate 66% (68%)
Came back anyway 31 (52)
Sampled accuracy 95% (91%)
Real resolution rate 56% (50%)
Handover pick-up, in hours (median) 20 min (25 min)
Handover pick-up, out of hours 9h 10m (11h)
Cost per genuine resolution $1.21 ($1.33)
Jobs booked from chat 52 (46)
Top content fix this month: night call-out price and routing
Notice that the vendor's figure fell slightly while the real rate rose. That's typical after tightening handover rules: more conversations correctly go to a person, which lowers the vendor's resolution count and raises the number of customers genuinely helped.
Fix it, change it or switch it off
After two or three months the numbers tell you which of three moves to make. These thresholds are my working rules, not industry standards.
- Fix the content when accuracy is under 90% or errors cluster on a few topics. That's most bots in their first quarter, and rewriting FAQs, policies and prices as single-topic entries is usually the answer.
- Fix the handovers when pick-up times exceed what the bot promises, or complaints mention being left waiting. Change the promise, the routing or the staffing; the rules on when a chatbot should hand over to a human help here.
- Switch it off when, after three months of fixes, the real resolution rate is under about a third, the business result is flat against the baseline, and cost per genuine resolution exceeds the cost of a person answering. A good contact form with an honest reply time beats a bot that doesn't help.
For a wider view of spot-checking every kind of AI support reply, not just the chatbot, a weekly sampling routine for AI support replies extends the grading method above.
Measuring a chatbot: questions owners ask
What is a good resolution rate for a small business chatbot?
There's no universal figure, and vendor-quoted benchmarks use their own generous definitions. Measure your own real resolution rate after removing customers who came back through another channel and answers your sampling found wrong. Many small firms find a real rate between a third and two thirds is realistic, depending on how repetitive their questions are. The trend over three months matters more than the starting number.
How many conversations should I read each week?
Twenty to thirty is enough for most small businesses, picked at random rather than chosen, plus every conversation that ended in a complaint or a thumbs-down. At lower volumes, read them all for the first month. The aim is to catch patterns of wrong answers early, not to produce a statistically precise error rate.
Should I add a satisfaction survey at the end of the chat?
A one-click thumbs up or down is worth having because it costs the customer nothing, but treat the results with caution. People who leave annoyed rarely click anything, so satisfaction scores from chat surveys run high. Pair the survey with the came-back-anyway check and your own sampling, which catch the unhappy customers the survey misses.
How long before I can judge whether the chatbot works?
Give it six to eight weeks. The first two weeks mostly reveal content gaps, which you fix; weeks three to six show whether those fixes hold; by week eight you have enough conversations to compare with the baseline you recorded before launch. Judging after one week usually means judging the content you forgot to write, not the chatbot.
Further reads
- How to Calculate AI ROI for Your Business (Worked Example) — Turn the scorecard into a return-on-investment figure.
- 9 AI Chatbot Mistakes That Lose Small Businesses Customers — The failures your measurements are most likely to reveal.
- How to Measure Customer Reaction After Introducing AI — Wider ways to read how customers feel about AI contact.
- How to Measure an AI Receptionist's Return in the First 90 Days — The same discipline applied to an AI phone receptionist.
- How to Test a Customer Chatbot Before It Goes Live — Catch the worst errors before the numbers ever start.
- How to Build a KPI Dashboard With AI When You Have No Data Team — Put these chatbot numbers on a simple monthly dashboard.
- AI KPIs for Small Businesses: 12 Metrics Worth Tracking — Twelve AI metrics explained with how to measure each, an example and how it misleads, plus a guide to choosing yours and a one-screen monthly sheet.
- Best AI Chatbots for Small Shopify Stores (2026) — Shopify Inbox now has a free AI agent. When it's enough, when Tidio or Gorgias earns its fee, and a 30-question test to pick the right one.
- How Wedding Venues Use AI for Viewings, Enquiries and Follow-Ups — Use AI at three points in a venue's sales pipeline: the first reply, the notes after each viewing, and follow-ups timed around provisional holds.
- How Day Spas Can Take Bookings From Instagram DMs With AI — Three ways to put AI in a spa's Instagram inbox, the knowledge sheet it needs, when to hand over to a therapist, and a 20-message test before launch.
- Your First 30 Days of AI Customer Support for an Online Shop — Week by week: collect real questions, let AI draft while people send, automate tracking, delivery and returns, then review 30 conversations.
- How to Build the FAQ Your AI Chatbot Needs Before Launch — How to source, write and test the FAQ a small business chatbot answers from, with a copyable entry format and a handover list.
- How Campsites and Glamping Sites Use AI to Answer Guests — A site facts sheet, the questions guests ask at each stage, where the AI should answer, and the rules it must never bend on a campsite.
- Can a Takeaway Take Orders on WhatsApp With AI? — How an AI ordering bot on WhatsApp handles a real takeaway order, what Meta now charges per message, and the Friday-night failures to plan for.
- How Small Hotels Use AI to Answer Guest Messages Around the Clock — The mechanics of overnight AI replies in a small hotel: how messages flow, what the AI may say at 2am, and how to check its work each morning.
- Can an AI Chatbot Handle Order Tracking and Returns? — What an order-tracking chatbot needs to see, which return decisions it can safely make, and where a person still has to step in.
- AI Tools and AI Development: The Complete 2026 Guide — the AI hub, including every tutorial in the AI-for-business series.
Sources: Intercom Fin help centre on confirmed and assumed resolutions; Zendesk help on automated resolutions for AI agents; Help Scout and Intercom pricing pages.