How to Spot-Check AI Support Replies: A Weekly Sampling Routine

Coding Liquids tutorial cover featuring Sagnik Bhattacharya for How to Spot-Check AI Support Replies: A Weekly Sampling Routine.
Coding Liquids tutorial cover featuring Sagnik Bhattacharya for How to Spot-Check AI Support Replies: A Weekly Sampling Routine.

Pull a fixed random sample of the AI's conversations every week, about 30 for most small businesses, weighted towards high-risk topics, handovers and anything the dashboard marked resolved. Score each against a six-question sheet (correct, complete, safe, handed over properly, genuinely resolved, on-brand), fix critical errors the same day, and track the pass rate week by week.

Reading a handful of chats when you have a spare moment feels like quality assurance but isn't. You read the memorable ones, and the memorable ones are rarely the dangerous ones: the confident wrong answer that the customer accepted without complaint never looks memorable. A routine fixes the sample size, the selection and the scoring, so this week's result means something next to last week's.

Follow me on Instagram@sagnikteaches

What a weekly sample can and can't tell you

A little probability tells you what to expect from a sample, and it explains why 30 is a sensible default rather than 5. If a certain kind of error happens in a fraction p of conversations, the chance that a random sample of n contains at least one is 1 − (1 − p)n. In practice:

Connect on LinkedInSagnik Bhattacharya
True error rateChance a 10-chat sample catches oneChance a 30-chat sample catches oneChance a 50-chat sample catches one
10% (1 in 10)65%96%Over 99%
5% (1 in 20)40%79%92%
2% (1 in 50)18%46%64%

Two conclusions follow. First, ten chats a week misses a 1-in-20 problem more often than it finds it. Second, a clean week proves less than it feels like. A useful rule of thumb (statisticians call it the rule of three): if you find zero errors in n conversations, you can be about 95% confident the true rate is below 3 ÷ n. Zero errors in 30 means "probably under 10%", which is not reassuring for errors that cost money. Zero in four weeks of 30, 120 in all, means "probably under 2.5%", which is better.

Subscribe on YouTube@codingliquids

So a weekly sample does two jobs well: it catches frequent problems fast, and over a month it builds evidence about rare ones. It will not catch a 1-in-500 error, which is why the selection below deliberately over-samples the conversations where rare, serious errors live.

Choosing which conversations to pull each week

A purely random sample spends most of its time on easy questions, because most questions are easy. Split the week's conversations into groups (strata) and draw a set number from each, so the risky groups get more attention than their share of volume.

A filled-in sampling plan for an illustrative six-person insurance broker whose AI assistant answers website chat and email:

GroupShare of weekly volumeChats sampledWhy
Cover questions ("am I covered for...")12%8A wrong answer here has direct financial consequences
Claims in progress or new claims8%5Distressed customers; handover must be fast and correct
Marked resolved with no reply from the customer20%6Silence often means the customer gave up
Handed over to a person15%4Checks the handover happened and carried the context
Negative rating or complaint words3%3Direct evidence of harm
Everything else (renewal dates, documents, payments)42%4Keeps a baseline on routine answers
Total100% (about 420 chats)30

The mix changes as you learn. If three weeks pass without a finding in routine answers, move two of those four chats to whichever group produced last week's critical finding.

Drawing a random sample in a spreadsheet

"Random" has to mean random, not "the first eight in the list" (which are all Monday morning) or "ones that look interesting". Most help desks and chat tools let you filter conversations handled by the AI and export them. From the export:

  1. Add a column that labels each conversation's group, using the topic or tag your tool assigns plus a filter for ratings and handovers.
  2. Add a column with =RAND() in every row.
  3. Copy that column and paste it back as values only. RAND recalculates every time the sheet changes, so without this step your sample reshuffles while you work.
  4. Sort by group, then by the random column, and take the top rows from each group as your plan specifies.
  5. Save the sample list with the week's date. Next month you will want to know exactly which conversations were reviewed.

The whole draw takes five minutes once the export is saved as a routine report.

The scoring sheet: six questions per conversation

Every conversation gets the same six questions, answered yes or no, with a severity for any no. Keeping the questions fixed is what makes weeks comparable.

  1. Correct: do the facts match your policies, prices and documents?
  2. Complete: did it answer the question the customer actually asked, not a nearby one?
  3. Safe: did it stay within scope: no advice it isn't allowed to give, no promises, no personal data shown to the wrong person?
  4. Handover: did it pass the conversation to a person when it should (claims, complaints, distress, anything it couldn't answer), and only then?
  5. Genuinely resolved: if marked resolved, was the customer's problem actually settled?
  6. On-brand: did it sound like your business, at the right level of warmth for the situation?

Illustrative rows from the broker's sheet:

ChatCustomer askedFindingQuestion failedSeverity
0917-044"Is my van covered if I use it for deliveries?"AI said "Yes, business use is included on your policy." It can't see the class of use on the certificate and should have handed overCorrect, Safe, HandoverCritical
0917-061"When does my home policy renew?"Correct date from the documents, polite closeNone
0917-102"My shop flooded last night, what do I do?"Gave the claims line and handed over within one exchangeNone
0917-130"How much to cancel my policy?"Quoted the fee correctly; customer went silent; marked resolved. The customer phoned the next day, upset about the feeGenuinely resolvedMajor
0917-188"My car was stolen from the drive"Correct next steps, but opened with "Great question!"On-brandMinor

Severity levels and what each one triggers

A finding without a consequence is just a note. Tie each severity to an action and a deadline:

  • Critical: wrong information about cover, money or legal position; a promise the business can't keep; personal data exposed; a distressed or vulnerable customer not handed over. Action: fix the instruction or knowledge source the same day, contact the affected customer, and consider switching the AI off for that topic until the fix is tested.
  • Major: incomplete or misleading but low-harm answers; wrong handover routing; false "resolved" status. Action: fix within the week and add the case to your test questions.
  • Minor: tone, length, formatting. Action: batch into a monthly instructions update.

Write the "contact the customer" step into the routine. In the van example, someone at the broker called the customer to explain that cover depends on the class of use shown on their certificate and to check it with them. That call is uncomfortable, and it is far cheaper than a declined claim the customer believed was covered. Then widen the search: filter the whole week's export for the same topic (here, every chat mentioning "covered" or "cover") and read all of them, not just the sampled one. In the broker's case that turned up two more customers who had been given the same reassurance, and both got the same call. A sample finds the problem; only a full search of that topic finds everyone it affected. Why an assistant gives confident wrong answers in the first place, and how to reduce them at source, is covered in how to stop an AI chatbot giving customers wrong answers.

One Monday session at a six-person insurance broker

Putting the pieces together, here is how the illustrative broker's first full session ran, with times:

  • 9:00 to 9:05. The account executive running the review opens last week's saved export: 418 AI-handled conversations.
  • 9:05 to 9:10. Groups labelled, random column added and pasted as values, 30 conversations drawn per the plan.
  • 9:10 to 9:50. Scoring. Most chats take about a minute; claims and cover questions take two or three.
  • 9:50 to 10:00. Findings logged, the critical one raised with the owner, fixes assigned.

Results: 24 of 30 passed every question, an 80% pass rate. The six failures were one critical (the van cover answer), two major (a false "resolved" and a handover that dropped the policy number, forcing the customer to repeat everything) and three minor tone issues. By Friday the instructions said, in plain words, "Never confirm whether something is covered. Explain that cover depends on the policy documents and offer to pass the question to a broker", the handover template carried the policy number, and the cover-question group went up from 8 to 10 chats for the next month, with the extra two taken from the routine group so the total stayed at 30.

The first month's log, one row per Monday:

WeekSampledPassedCriticalMajorMinorMain fix that week
13024 (80%)123"Never confirm cover" instruction
23026 (87%)112Claim numbers routed straight to the claims handler
33028 (93%)011Old fee schedule removed from the sources
43027 (90%)012Scoring guide: no praise openers

By week four the pass rate had settled at 90% to 93%, with no critical findings in weeks three and four. Two clean weeks is 60 conversations, which by the rule of three supports "critical errors are probably under 5%". Two more clean weeks, 120 in all, would support "probably under 2.5%". That is the kind of statement you can take to a compliance review, and it comes from an hour a week. Note the week-two critical finding: a different topic from week one. Fixing one error type doesn't mean the assistant is fixed, which is why the sample keeps running after the first good week.

The same routine adapts to other sectors by changing the groups. For an illustrative mortgage adviser, the high-risk group is any conversation containing words like "borrow", "afford" or "how much can I", because the assistant must hand those to an adviser rather than answer; sampling that group heavily checks the handover rule that matters most. The broader question of when a handover should happen is covered in when an AI chatbot should hand over to a human.

Why "resolved" in the dashboard needs its own sample

Vendors count resolutions in ways that suit billing and dashboards, and those definitions don't always mean the customer was helped. Intercom's Fin, charged at $0.99 per resolved outcome, counts an "assumed resolution" when the customer goes quiet for 24 hours after Fin's last answer. Zendesk waits two quiet hours on messaging (72 on email and web forms) before an AI check grades the conversation, and since 18 May 2026 you pay only for the resolutions it verifies, not for hand-offs or unverified closes. HubSpot's Customer Agent counts a resolution when there's no hand-off within 72 hours. Help Scout's AI Answers ($0.75 per resolution) is stricter, and doesn't bill a resolution if the customer escalates, searches the help docs, asks again or says they need more help. Even the strict definitions measure silence, though, and silence isn't satisfaction: a customer who gave up and rang instead looks exactly like a customer who got their answer.

That is why the broker's plan samples six "resolved with no reply" conversations every week, and why "genuinely resolved" is one of the six questions. The cancellation-fee chat in the scoring table was one of these: resolved on the dashboard, an angry phone call in reality. If your plan is billed per resolution, false resolutions also cost money directly. The tutorial on measuring whether your AI chatbot is actually working covers the dashboard metrics; the weekly sample is what tells you whether to believe them.

Letting AI pre-screen transcripts without trusting it

At higher volumes, you can ask a chat assistant to pre-screen a larger batch and point you at likely problems, then read those plus your normal random sample. Remove names, policy numbers and contact details first, and use a business plan that doesn't train on your data. A pre-screen prompt along these lines:

You are reviewing customer support chats for an insurance
broker. Our rules are attached. For each chat below, answer:
1. Did the assistant state or imply that something is
   covered? (quote the sentence)
2. Did the customer mention a claim, complaint, distress or
   money difficulty? If so, was the chat handed over?
3. Does the chat end without the customer confirming the
   answer helped?
Output one line per chat: ID, flags raised, quoted evidence.
Do not score tone.

[paste 60 anonymised transcripts]

A few lines of what it returned (illustrative):

0918-012 - Q1: "your policy should cover that" - FLAG
0918-027 - Q2: customer mentions "can't afford the renewal";
           not handed over - FLAG
0918-033 - no flags
0918-041 - Q3: ends after fee explanation - FLAG

Read in full, the chats flagged as 012 and 027 were genuine problems. But when the reviewer checked 033, it contained "you're fine to drive it abroad for a week", a cover statement the pre-screen missed because it didn't match the idea of "covered". The pre-screen widens the net; it doesn't replace the random sample, because the random sample is the only part of the routine that finds problems nobody thought to prompt for.

Keeping reviewers consistent

If two people take turns reviewing, their scores drift apart within a month: one marks "Great question!" as a minor finding, the other lets it pass; one calls a slow handover major, the other minor. Once a month, have both reviewers score the same five conversations independently, then compare. Where they disagree, write the decision into a one-page scoring guide with the example attached ("Opening with praise for the question: minor, every time"). Twenty minutes of calibration keeps the pass rate measuring the assistant rather than the reviewer.

From findings to fixes by Friday

The routine only pays off if findings change something. Every critical or major finding should end in one of three fixes, recorded next to the finding:

  1. An instruction change. Before: "Answer customer questions about their policies helpfully." After: "Never confirm whether a loss or activity is covered. Say cover depends on the policy documents and offer to pass the question to a broker."
  2. A knowledge fix. The assistant quoted last year's cancellation fee because the old fee schedule was still in its sources. Remove or replace the outdated document, not just correct the answer.
  3. A routing change. Conversations mentioning a claim number go straight to the claims handler, skipping the assistant entirely.

Add every critical finding to a list of test questions and re-run them after each fix, so a repaired answer stays repaired after the next update. Once a month, look at the four weekly pass rates together, the groups where findings cluster and the fixes made; the tutorial on running a monthly AI quality review in 30 minutes gives that meeting a structure. And if the weekly hour starts to feel like too much at your volume, dedicated QA software can score every conversation automatically; what to compare in customer service QA software covers when that becomes worth paying for.

Spot-checking AI replies: follow-up questions

Should the person who set up the AI assistant do the spot-checks?

Not alone. The person who wrote the instructions knows what the assistant was meant to say and reads replies charitably. Rotate the weekly review between two or three people who know your policies, and have the set-up owner fix what the reviewers find. That split also keeps the review going when one person is away.

Do I need customers' permission to review their chat transcripts?

Reviewing conversations for quality is a normal business use, but it should be covered by your privacy notice, and reviewers should see only what they need. If you paste transcripts into a separate AI tool, remove names and account details first and use a business plan that doesn't train on your data. Ask your data-protection adviser if you're unsure.

How many conversations should I sample if volume is very low?

If the assistant handles fewer than about 60 conversations a week, read all of them for the first month rather than sampling. Once you know where it goes wrong, drop to 20 a week weighted towards the risky topics. Below roughly 20 conversations a week, reading everything remains practical and tells you more.

Further reads

Sources: Intercom and Zendesk documentation on how AI resolutions are counted (via the verified fact sheet, September 2026). Sampling figures are standard probability calculations.

Want a QA routine sized to your support volume?

On a 1:1 call we'll look at how your AI assistant is set up, decide which conversations to sample and how many, and build a scoring sheet and fix-list your team can run in an hour a week.

Book a 1:1 call with me