Pull a fixed random sample of the AI's conversations every week, about 30 for most small businesses, weighted towards high-risk topics, handovers and anything the dashboard marked resolved. Score each against a six-question sheet (correct, complete, safe, handed over properly, genuinely resolved, on-brand), fix critical errors the same day, and track the pass rate week by week.
Reading a handful of chats when you have a spare moment feels like quality assurance but isn't. You read the memorable ones, and the memorable ones are rarely the dangerous ones: the confident wrong answer that the customer accepted without complaint never looks memorable. A routine fixes the sample size, the selection and the scoring, so this week's result means something next to last week's.
What a weekly sample can and can't tell you
A little probability tells you what to expect from a sample, and it explains why 30 is a sensible default rather than 5. If a certain kind of error happens in a fraction p of conversations, the chance that a random sample of n contains at least one is 1 − (1 − p)n. In practice:
| True error rate | Chance a 10-chat sample catches one | Chance a 30-chat sample catches one | Chance a 50-chat sample catches one |
|---|---|---|---|
| 10% (1 in 10) | 65% | 96% | Over 99% |
| 5% (1 in 20) | 40% | 79% | 92% |
| 2% (1 in 50) | 18% | 46% | 64% |
Two conclusions follow. First, ten chats a week misses a 1-in-20 problem more often than it finds it. Second, a clean week proves less than it feels like. A useful rule of thumb (statisticians call it the rule of three): if you find zero errors in n conversations, you can be about 95% confident the true rate is below 3 ÷ n. Zero errors in 30 means "probably under 10%", which is not reassuring for errors that cost money. Zero in four weeks of 30, 120 in all, means "probably under 2.5%", which is better.
So a weekly sample does two jobs well: it catches frequent problems fast, and over a month it builds evidence about rare ones. It will not catch a 1-in-500 error, which is why the selection below deliberately over-samples the conversations where rare, serious errors live.
Choosing which conversations to pull each week
A purely random sample spends most of its time on easy questions, because most questions are easy. Split the week's conversations into groups (strata) and draw a set number from each, so the risky groups get more attention than their share of volume.
A filled-in sampling plan for an illustrative six-person insurance broker whose AI assistant answers website chat and email:
| Group | Share of weekly volume | Chats sampled | Why |
|---|---|---|---|
| Cover questions ("am I covered for...") | 12% | 8 | A wrong answer here has direct financial consequences |
| Claims in progress or new claims | 8% | 5 | Distressed customers; handover must be fast and correct |
| Marked resolved with no reply from the customer | 20% | 6 | Silence often means the customer gave up |
| Handed over to a person | 15% | 4 | Checks the handover happened and carried the context |
| Negative rating or complaint words | 3% | 3 | Direct evidence of harm |
| Everything else (renewal dates, documents, payments) | 42% | 4 | Keeps a baseline on routine answers |
| Total | 100% (about 420 chats) | 30 |
The mix changes as you learn. If three weeks pass without a finding in routine answers, move two of those four chats to whichever group produced last week's critical finding.
Drawing a random sample in a spreadsheet
"Random" has to mean random, not "the first eight in the list" (which are all Monday morning) or "ones that look interesting". Most help desks and chat tools let you filter conversations handled by the AI and export them. From the export:
- Add a column that labels each conversation's group, using the topic or tag your tool assigns plus a filter for ratings and handovers.
- Add a column with
=RAND()in every row. - Copy that column and paste it back as values only. RAND recalculates every time the sheet changes, so without this step your sample reshuffles while you work.
- Sort by group, then by the random column, and take the top rows from each group as your plan specifies.
- Save the sample list with the week's date. Next month you will want to know exactly which conversations were reviewed.
The whole draw takes five minutes once the export is saved as a routine report.
The scoring sheet: six questions per conversation
Every conversation gets the same six questions, answered yes or no, with a severity for any no. Keeping the questions fixed is what makes weeks comparable.
- Correct: do the facts match your policies, prices and documents?
- Complete: did it answer the question the customer actually asked, not a nearby one?
- Safe: did it stay within scope: no advice it isn't allowed to give, no promises, no personal data shown to the wrong person?
- Handover: did it pass the conversation to a person when it should (claims, complaints, distress, anything it couldn't answer), and only then?
- Genuinely resolved: if marked resolved, was the customer's problem actually settled?
- On-brand: did it sound like your business, at the right level of warmth for the situation?
Illustrative rows from the broker's sheet:
| Chat | Customer asked | Finding | Question failed | Severity |
|---|---|---|---|---|
| 0917-044 | "Is my van covered if I use it for deliveries?" | AI said "Yes, business use is included on your policy." It can't see the class of use on the certificate and should have handed over | Correct, Safe, Handover | Critical |
| 0917-061 | "When does my home policy renew?" | Correct date from the documents, polite close | None | |
| 0917-102 | "My shop flooded last night, what do I do?" | Gave the claims line and handed over within one exchange | None | |
| 0917-130 | "How much to cancel my policy?" | Quoted the fee correctly; customer went silent; marked resolved. The customer phoned the next day, upset about the fee | Genuinely resolved | Major |
| 0917-188 | "My car was stolen from the drive" | Correct next steps, but opened with "Great question!" | On-brand | Minor |
Severity levels and what each one triggers
A finding without a consequence is just a note. Tie each severity to an action and a deadline:
- Critical: wrong information about cover, money or legal position; a promise the business can't keep; personal data exposed; a distressed or vulnerable customer not handed over. Action: fix the instruction or knowledge source the same day, contact the affected customer, and consider switching the AI off for that topic until the fix is tested.
- Major: incomplete or misleading but low-harm answers; wrong handover routing; false "resolved" status. Action: fix within the week and add the case to your test questions.
- Minor: tone, length, formatting. Action: batch into a monthly instructions update.
Write the "contact the customer" step into the routine. In the van example, someone at the broker called the customer to explain that cover depends on the class of use shown on their certificate and to check it with them. That call is uncomfortable, and it is far cheaper than a declined claim the customer believed was covered. Then widen the search: filter the whole week's export for the same topic (here, every chat mentioning "covered" or "cover") and read all of them, not just the sampled one. In the broker's case that turned up two more customers who had been given the same reassurance, and both got the same call. A sample finds the problem; only a full search of that topic finds everyone it affected. Why an assistant gives confident wrong answers in the first place, and how to reduce them at source, is covered in how to stop an AI chatbot giving customers wrong answers.
One Monday session at a six-person insurance broker
Putting the pieces together, here is how the illustrative broker's first full session ran, with times:
- 9:00 to 9:05. The account executive running the review opens last week's saved export: 418 AI-handled conversations.
- 9:05 to 9:10. Groups labelled, random column added and pasted as values, 30 conversations drawn per the plan.
- 9:10 to 9:50. Scoring. Most chats take about a minute; claims and cover questions take two or three.
- 9:50 to 10:00. Findings logged, the critical one raised with the owner, fixes assigned.
Results: 24 of 30 passed every question, an 80% pass rate. The six failures were one critical (the van cover answer), two major (a false "resolved" and a handover that dropped the policy number, forcing the customer to repeat everything) and three minor tone issues. By Friday the instructions said, in plain words, "Never confirm whether something is covered. Explain that cover depends on the policy documents and offer to pass the question to a broker", the handover template carried the policy number, and the cover-question group went up from 8 to 10 chats for the next month, with the extra two taken from the routine group so the total stayed at 30.
The first month's log, one row per Monday:
| Week | Sampled | Passed | Critical | Major | Minor | Main fix that week |
|---|---|---|---|---|---|---|
| 1 | 30 | 24 (80%) | 1 | 2 | 3 | "Never confirm cover" instruction |
| 2 | 30 | 26 (87%) | 1 | 1 | 2 | Claim numbers routed straight to the claims handler |
| 3 | 30 | 28 (93%) | 0 | 1 | 1 | Old fee schedule removed from the sources |
| 4 | 30 | 27 (90%) | 0 | 1 | 2 | Scoring guide: no praise openers |
By week four the pass rate had settled at 90% to 93%, with no critical findings in weeks three and four. Two clean weeks is 60 conversations, which by the rule of three supports "critical errors are probably under 5%". Two more clean weeks, 120 in all, would support "probably under 2.5%". That is the kind of statement you can take to a compliance review, and it comes from an hour a week. Note the week-two critical finding: a different topic from week one. Fixing one error type doesn't mean the assistant is fixed, which is why the sample keeps running after the first good week.
The same routine adapts to other sectors by changing the groups. For an illustrative mortgage adviser, the high-risk group is any conversation containing words like "borrow", "afford" or "how much can I", because the assistant must hand those to an adviser rather than answer; sampling that group heavily checks the handover rule that matters most. The broader question of when a handover should happen is covered in when an AI chatbot should hand over to a human.
Why "resolved" in the dashboard needs its own sample
Vendors count resolutions in ways that suit billing and dashboards, and those definitions don't always mean the customer was helped. Intercom's Fin, charged at $0.99 per resolved outcome, counts an "assumed resolution" when the customer goes quiet for 24 hours after Fin's last answer. Zendesk waits two quiet hours on messaging (72 on email and web forms) before an AI check grades the conversation, and since 18 May 2026 you pay only for the resolutions it verifies, not for hand-offs or unverified closes. HubSpot's Customer Agent counts a resolution when there's no hand-off within 72 hours. Help Scout's AI Answers ($0.75 per resolution) is stricter, and doesn't bill a resolution if the customer escalates, searches the help docs, asks again or says they need more help. Even the strict definitions measure silence, though, and silence isn't satisfaction: a customer who gave up and rang instead looks exactly like a customer who got their answer.
That is why the broker's plan samples six "resolved with no reply" conversations every week, and why "genuinely resolved" is one of the six questions. The cancellation-fee chat in the scoring table was one of these: resolved on the dashboard, an angry phone call in reality. If your plan is billed per resolution, false resolutions also cost money directly. The tutorial on measuring whether your AI chatbot is actually working covers the dashboard metrics; the weekly sample is what tells you whether to believe them.
Letting AI pre-screen transcripts without trusting it
At higher volumes, you can ask a chat assistant to pre-screen a larger batch and point you at likely problems, then read those plus your normal random sample. Remove names, policy numbers and contact details first, and use a business plan that doesn't train on your data. A pre-screen prompt along these lines:
You are reviewing customer support chats for an insurance
broker. Our rules are attached. For each chat below, answer:
1. Did the assistant state or imply that something is
covered? (quote the sentence)
2. Did the customer mention a claim, complaint, distress or
money difficulty? If so, was the chat handed over?
3. Does the chat end without the customer confirming the
answer helped?
Output one line per chat: ID, flags raised, quoted evidence.
Do not score tone.
[paste 60 anonymised transcripts]
A few lines of what it returned (illustrative):
0918-012 - Q1: "your policy should cover that" - FLAG
0918-027 - Q2: customer mentions "can't afford the renewal";
not handed over - FLAG
0918-033 - no flags
0918-041 - Q3: ends after fee explanation - FLAG
Read in full, the chats flagged as 012 and 027 were genuine problems. But when the reviewer checked 033, it contained "you're fine to drive it abroad for a week", a cover statement the pre-screen missed because it didn't match the idea of "covered". The pre-screen widens the net; it doesn't replace the random sample, because the random sample is the only part of the routine that finds problems nobody thought to prompt for.
Keeping reviewers consistent
If two people take turns reviewing, their scores drift apart within a month: one marks "Great question!" as a minor finding, the other lets it pass; one calls a slow handover major, the other minor. Once a month, have both reviewers score the same five conversations independently, then compare. Where they disagree, write the decision into a one-page scoring guide with the example attached ("Opening with praise for the question: minor, every time"). Twenty minutes of calibration keeps the pass rate measuring the assistant rather than the reviewer.
From findings to fixes by Friday
The routine only pays off if findings change something. Every critical or major finding should end in one of three fixes, recorded next to the finding:
- An instruction change. Before: "Answer customer questions about their policies helpfully." After: "Never confirm whether a loss or activity is covered. Say cover depends on the policy documents and offer to pass the question to a broker."
- A knowledge fix. The assistant quoted last year's cancellation fee because the old fee schedule was still in its sources. Remove or replace the outdated document, not just correct the answer.
- A routing change. Conversations mentioning a claim number go straight to the claims handler, skipping the assistant entirely.
Add every critical finding to a list of test questions and re-run them after each fix, so a repaired answer stays repaired after the next update. Once a month, look at the four weekly pass rates together, the groups where findings cluster and the fixes made; the tutorial on running a monthly AI quality review in 30 minutes gives that meeting a structure. And if the weekly hour starts to feel like too much at your volume, dedicated QA software can score every conversation automatically; what to compare in customer service QA software covers when that becomes worth paying for.
Spot-checking AI replies: follow-up questions
Should the person who set up the AI assistant do the spot-checks?
Not alone. The person who wrote the instructions knows what the assistant was meant to say and reads replies charitably. Rotate the weekly review between two or three people who know your policies, and have the set-up owner fix what the reviewers find. That split also keeps the review going when one person is away.
Do I need customers' permission to review their chat transcripts?
Reviewing conversations for quality is a normal business use, but it should be covered by your privacy notice, and reviewers should see only what they need. If you paste transcripts into a separate AI tool, remove names and account details first and use a business plan that doesn't train on your data. Ask your data-protection adviser if you're unsure.
How many conversations should I sample if volume is very low?
If the assistant handles fewer than about 60 conversations a week, read all of them for the first month rather than sampling. Once you know where it goes wrong, drop to 20 a week weighted towards the risky topics. Below roughly 20 conversations a week, reading everything remains practical and tells you more.
Further reads
- Chatbot Guardrails: Stop AI Promising What You Don't Offer — Stop the promises your sample keeps catching.
- AI Error Log: Track Mistakes and Stop Them Happening Again — A log format for findings that stops repeats.
- How to Test a Customer Chatbot Before It Goes Live — The pre-launch tests that make weekly findings rarer.
- How to Write Reply Templates That Keep AI Replies On-Script — Reply templates that keep answers to sensitive questions on-script.
- Who Is Liable When Your AI Chatbot Gets It Wrong? — What's at stake when a critical finding reached a customer.
- How to Handle Sensitive Customer Conversations Without AI — Which conversations to route away from AI altogether.
- How to Set Up Human Review for AI Work Without Slowing Down — Four levels of human review matched to risk, how to make each check take under a minute, how many to sample, and when to relax or tighten.
- How to Pilot AI in Shadow Mode Before Customers See It — How to run AI in parallel with your team, log and grade what it would have done, and decide from real numbers when it's safe to let customers see it.
- How to Check AI Is Doing Good Work, Not Just Fast Work — A rubric template, a 30-minute weekly sampling routine, a re-runnable test set and the warning signs, shown through a garden centre's plant-care replies.
- How to Measure Customer Reaction After Introducing AI — Four signals, survey wording that doesn't lead, a conversation-sorting prompt and a decision rule for reading small-business numbers honestly.
- Best AI Chatbots for Small Shopify Stores (2026) — Shopify Inbox now has a free AI agent. When it's enough, when Tidio or Gorgias earns its fee, and a 30-question test to pick the right one.
- Your First 30 Days of AI Customer Support for an Online Shop — Week by week: collect real questions, let AI draft while people send, automate tracking, delivery and returns, then review 30 conversations.
- Should You Let AI Auto-Post Replies to Google Reviews? — A sorting rule for which Google reviews an AI may answer on its own, why star ratings make a poor trigger, and the guardrails for the auto-post lane.
- How Small IT Support Businesses Use AI to Resolve Tickets Faster — Use AI to prepare the evidence and next check for each ticket, while technicians approve changes and confirm that service is restored.
- Photo Checklists and Quality Control for Cleaning Teams Using AI — Set checkpoints a camera can prove, a fixed photo set per clean, an AI pass that flags which jobs need a supervisor, and clear privacy rules for clients' homes.
- Should a Small Business Let AI Answer Customer Messages? — Sort your last 100 messages, pick the right level of AI involvement, set red lines by business type, and know what each channel costs per reply.
- AI Tools and AI Development: The Complete 2026 Guide — the AI hub, including every tutorial in the AI-for-business series.
Sources: Intercom and Zendesk documentation on how AI resolutions are counted (via the verified fact sheet, September 2026). Sampling figures are standard probability calculations.