Score a weekly sample of transcripts against a written scorecard that separates quality (was the problem understood, solved and logged, with the right tone) from compliance (identity checked, recording and AI disclosed, no card details or passwords captured, no promises outside your terms). Let AI flag risky calls first, but confirm every flagged moment against the audio before acting.
Transcripts are evidence with a margin of error. Speech-to-text mishears names, numbers and technical terms, and speaker labels can swap, so a line that seems to show a technician reading out a password may be a misheard "password reset". Treat the transcript as the index and the recording as the record.
Quality and compliance are two separate reviews
Call quality monitoring usually blends two questions that deserve different handling. Quality asks whether the call went well: the caller's problem understood, the right fix, a clear explanation, an accurate ticket. It's scored, trended and used for coaching. Compliance asks whether anything happened that shouldn't have: an account reset for someone who wasn't verified, a card number spoken on a recorded line, a promise the contract doesn't support. It isn't averaged. One failure is an incident.
Keeping them apart stops a common distortion, where a friendly technician scores 9 out of 10 overall despite resetting multi-factor authentication for a stranger. It also means different people can own each: in an illustrative 12-person IT support firm, the helpdesk lead reviews quality and the service manager owns compliance.
The scorecard, filled in for one helpdesk call
Here's the firm's scorecard applied to a real-looking call: a client's employee locked out of Microsoft 365 after changing phones. The call lasted 8 minutes 20 seconds.
| Item | Result | Evidence (timestamp) |
|---|---|---|
| Q1 Problem understood and restated | 2 of 2 | 00:48 "So the new phone doesn't have your authenticator app yet" |
| Q2 Correct fix | 2 of 2 | MFA re-registration walked through, 04:30 to 06:10 |
| Q3 Next steps explained | 1 of 2 | Didn't say the old phone's app would stop working |
| Q4 Ticket accurate | 1 of 2 | Ticket says "password reset"; it was an MFA reset |
| Q5 Clear, no unexplained jargon | 2 of 2 | Explained "authenticator app" at 01:15 |
| Quality total | 8 of 10 | |
| C1 Recording notice at start | Pass | 00:03 "Calls are recorded for quality and training" |
| C2 Identity verified before any reset | Fail | 03:55 caller: "It's [first name] from accounts, you know me." No callback or passphrase before the reset at 04:12 |
| C3 No passwords or card data spoken | Pass | AI flag at 06:40 was a transcription error; audio checked |
| C4 Promises within the contract | Fail | 07:10 "I'll get someone out to you today"; the client's contract says next working day for on-site |
| C5 Only necessary personal data collected | Pass |
The outcome: C2 goes to the service manager the same day, because an unverified MFA reset is exactly how account takeovers start. C4 becomes a coaching point and a call to the client to reset expectations. Q4 gets the ticket corrected. The quality score of 8 is decent, and it's recorded, but it's the least important number on the sheet.
Evidence with timestamps is what makes the card usable. A reviewer who writes "didn't verify identity" invites an argument; one who writes "03:55, no callback before the 04:12 reset" invites a fix.
How many calls to read, and which ones
Reading every call is rarely realistic, and reading a handful chosen by instinct misses patterns. A mixed sample works better. The firm's helpdesk takes about 620 calls a month handled by five technicians, plus about 90 after-hours calls answered by an AI voice agent. Its plan:
- Random sample of technician calls: 8 a week, about 35 a month, spread so every technician gets at least 6 a month. New starters get double for their first two months.
- Every call the AI pre-screen flags: around 20 a month, many of them false alarms that take two minutes to dismiss.
- Every voice-agent call that did more than take a message: about 25 a month, because the agent's scope is where the surprises are.
- Every call linked to a complaint or an incident, whenever one comes in.
To keep the random sample honest, don't let reviewers pick. Export the week's call list from the phone system into a spreadsheet, add a column of random numbers, sort by it, and take the top eight, skipping any technician who already has two that week. It takes three minutes and removes the temptation to review only the calls you already suspect.
That's roughly 80 calls a month out of 710, about 11%, but weighted heavily towards the calls most likely to contain a problem. For a deeper look at sampling maths and rotating reviewers, the weekly sampling routine for AI support replies applies the same thinking to written replies.
An AI pre-screen that quotes the transcript
An AI model is good at reading 700 transcripts for specific moments and poor at judging them, so give it a narrow job and make it show its evidence. The firm's prompt:
You are checking an IT helpdesk call transcript for compliance.
Report only the items below. For each, give the timestamp and an
exact quote. If an item isn't in the transcript, write "not found".
Don't guess and don't summarise the call.
1 IDENTITY: was the caller verified (callback to the number on file,
or the client's passphrase) BEFORE any password reset, MFA reset
or access change?
2 SENSITIVE DATA: any password, card number, card security code or
full bank details spoken by anyone.
3 NOTICES: the recording notice; for AI agent calls, the statement
that the caller is speaking to an AI assistant.
4 PROMISES: any commitment on response time, on-site visits, cost
or refunds.
5 OTHER: at most two further moments a compliance reviewer should hear.
Its illustrative output for the call scored above:
1 IDENTITY: not found before the MFA reset at 04:12. At 03:55 the
caller says "It's [first name] from accounts, you know me"; the
technician replies "No problem" and continues.
2 SENSITIVE DATA: 06:40 "my password is reset one two three".
Possible password spoken.
3 NOTICES: 00:03 "Calls are recorded for quality and training."
4 PROMISES: 07:10 "I'll get someone out to you today."
5 OTHER: none.
Items 1 and 4 were real. Item 2 was not: the audio at 06:40 says "my password reset was at one twenty-three", meaning the time of a failed attempt, and the transcription engine mangled it. That's the pattern to expect from a pre-screen: it finds the right moments and gets some of them wrong, which is fine as long as a person listens before anything happens. Requiring exact quotes and timestamps is what makes the listening take seconds rather than minutes.
If you'd rather buy this than build it, dedicated quality-assurance tools for support teams do scoring, sampling and AI flagging in one place; what to compare in customer service QA software with AI covers the choice.
Check the transcript against the audio before you act
Every flagged moment gets a short listen before it becomes a finding, and no one is ever disciplined on a transcript alone. Technical helpdesk calls are hard work for transcription engines. Errors the firm's reviewers see regularly, with what the audio actually said:
- "Reset your entry ID" for "reset your Entra ID" (Microsoft's current name for what was Azure Active Directory).
- "Open in tune" for "open Intune".
- "Ticket fifty-two eighty" for "ticket fifteen-two-eighty".
- Speaker labels swapped for a stretch of the call, so the technician's question appears as the caller's.
- Names of client staff spelt three different ways in one transcript.
Two minutes of audio around each flag is usually enough. Note in the scorecard that the audio was checked, as the firm did for C3, so nobody later wonders whether a finding rests on a transcription glitch.
Compliance points that matter on an IT support line
Most compliance lists are written for regulated call centres. An IT support helpdesk has its own short list, and the first item is the most important.
- Identity before access. Password resets, MFA resets and new-device approvals are what attackers phone the helpdesk for. A caller who "sounds like" the finance director is not verified; a callback to the number on file, or the client's agreed passphrase, is. Voice cloning makes "I recognised the voice" worth even less than it used to be, which spotting deepfake voice and video scams covers in detail.
- The recording notice. Rules on recording calls differ between jurisdictions: some need only one party's consent, others need everyone's, and many expect a clear notice. Check the rules that apply to you and your callers, and make sure the notice plays on every call, including those the AI agent answers.
- AI disclosure. If you sell to customers in the EU, the AI Act's duty to tell people they're interacting with an AI system has applied since 2 August 2026. A voice agent should say so in its first sentence.
- Card details. The card industry's security standard (PCI DSS) says card security codes must not be stored after authorisation, and its guidance treats audio recordings that can be searched as storage. A transcript is searchable text. If a client reads a card number to pay an invoice, you now hold it in a recording, a transcript and possibly an AI summary. Send payment links instead, and redact and delete any transcript where card details slip through.
- Passwords spoken aloud. Never ask for one. Every transcript that contains a password is a password store nobody secured.
- Promises outside the contract. Response times, on-site visits and "we won't charge for this" all bind the firm in the client's mind.
- Retention. Transcripts are personal data under data-protection law such as the GDPR. Decide how long you keep recordings and transcripts, set the deletion in the tool, and make sure the AI summaries attached to tickets follow the same rule.
Here's how the card-details problem tends to surface. A client's office manager calls to pay an overdue invoice and, before the technician can stop her, reads out the card number, expiry and security code. The call is recorded, transcribed, summarised by AI and the summary pasted into the ticket. The weekly review finds it three days later, in three places. The firm deletes the recording and transcript, edits the ticket, tells the office manager what happened, and adds a line to the technician script: "I can't take card details on this line, but I'll email you a secure payment link now."
Reviewing the AI voice agent's own calls
When an AI agent answers calls, its transcripts get the same compliance check plus two extra questions: did it say it's an AI, and did it stay inside its job? The firm's after-hours agent is meant to take details, judge priority from a short list, and page the on-call engineer only for a priority-one outage. Its transcripts showed a different habit in the first month:
Caller: Our whole office can't get online and we open at eight.
Agent: I'm sorry to hear that. An engineer will be on site within
the hour to get you back up and running.
The paging was correct; the promise was not. The contract offers a one-hour response for priority-one issues, meaning a call back from an engineer, not a site visit. Four calls that month contained a version of the same promise. The fix was in the agent's script: "The on-call engineer has been alerted and will call you back within one hour." How to write those scripts and escalation rules properly is covered in call scripts and escalation rules for an AI receptionist.
The tool you use decides what you get to review. Quo (formerly OpenPhone), for example, includes AI call summaries and transcripts for all calls on its Business plan ($23 per user a month billed annually, $33 monthly) and Scale plan ($35 annually, $47 monthly); on Starter ($15 annually, $19 monthly) they're only produced for calls its AI agent, Sona, handles. Meeting-focused transcription tools such as Otter.ai run from $8.33 per user a month billed annually on Pro. Whatever you use, check three settings: how long transcripts are kept, who can see them, and whether they can be exported for review.
The same day a compliance failure turns up
A quality finding can wait for the monthly coaching note. A compliance failure can't, because the risk it created may still be live. For the unverified MFA reset in the scorecard above, the service manager worked through this the same afternoon:
- Confirm with the real person. Call the employee back on the number held on file, not the number that called in, and ask whether they requested the reset.
- Check for misuse. Review the account's recent sign-ins and any new devices or mailbox rules since the reset. If anything looks wrong, or the employee didn't make the request, revoke active sessions and reset again with proper verification.
- Tell the client if there was any risk. In this case the caller was genuine, so the client got a short note explaining the process gap and the fix. Had it not been, the client would have heard within the hour.
- Record it. What happened, what was checked, what changed. The same record answers the client's questions later and feeds the monthly review.
The technician involved hears about it privately, with the timestamp, and the conversation is about the process rather than blame. Most unverified resets happen because a familiar voice and a busy queue make the check feel rude; the fix that works is making the check automatic, not making people more anxious.
From findings to coaching and fixes
Reviews only pay off if something changes afterwards. The firm runs three short routines:
- Calibration, 20 minutes a week. The helpdesk lead and the service manager score the same two calls separately, then compare. Where they disagree, the scorecard wording gets clearer. Without this, "next steps explained" means something different to every reviewer.
- Coaching notes, per technician, monthly. Two things done well, one to work on, each tied to a timestamp. AI can draft these from the scorecards; turning call recordings into coaching notes with AI shows the approach, which carries over from sales calls to support calls.
- Process fixes, as found. A compliance failure that recurs is a process problem. Three unverified resets in a month led the firm to add a hard stop to its ticketing tool: the MFA reset form won't submit without a verification method selected.
A month of reviews at a 12-person IT support firm
Putting the illustration together, here's the monthly effort and what it found:
- Random technician calls: 35 at about 7 minutes each, 245 minutes.
- AI-flagged calls: 20 at about 5 minutes each, 100 minutes. Twelve were false alarms.
- Voice-agent calls: 25 at about 4 minutes each, 100 minutes.
- Calibration: 4 sessions of 20 minutes, 80 minutes.
That's 525 minutes, about 8.75 hours a month. At an internal cost of $55 an hour, roughly $480. In the first month it found three MFA resets without verification, one card number read aloud, four voice-agent promises of on-site visits, six tickets that didn't match the call, and two calls where a technician's frustration showed. None of these would have appeared in a customer satisfaction score. The three unverified resets alone justify the time: a single account takeover at a client costs far more than a year of reviews, and it lands on the firm's reputation as well as the client's.
By the third month, with the ticketing hard stop and the revised agent script in place, unverified resets and on-site promises had both dropped to zero in the sample, and the firm cut the random sample from 8 calls a week to 6, keeping the flagged and voice-agent reviews as they were.
Further reads
- How to Set Up an AI Receptionist Without Losing Callers — Set up an AI receptionist so its calls start out compliant.
- How Small IT Support Businesses Use AI to Resolve Tickets Faster — Use AI on the ticket side of an IT helpdesk too.
- AI vs Human Transcription: Which Is Worth Paying For? — Decide when a human transcript is worth paying for.
- AI Error Log: Track Mistakes and Stop Them Happening Again — Log what reviews find so fixes stick.
- How to Run a Monthly AI Quality Review in 30 Minutes — Roll call findings into a monthly quality review.
- AI Chatbot Disclosure: What to Tell Customers at the Start of a Chat — Word the AI disclosure your voice agent should open with.
- How to Set Up an AI Phone Line for Takeaway Orders — Seven steps from counting your calls to a tested AI phone line that takes takeaway orders into your till, with the three routes compared.
- What to Do When an AI Receptionist Gets a Booking Wrong — A step-by-step response for AI receptionist booking errors: the callback script, a cause-finding table, who absorbs the cost and an incident log.
- How Advisers Use AI to Prepare for Annual Client Reviews — A four-week countdown for review prep with AI: what it drafts, what the adviser checks, and the prompts that keep figures and advice in human hands.
- AI vs Human Answering Service for HVAC Firms — When an HVAC firm should let AI answer the phone, when a person should, and why most end up with a hybrid that escalates emergencies.
- Setting Up an AI Answering Service That Never Misses an Emergency — Emergency definitions, a word-for-word screening script, a transfer chain with backups, and monthly test calls that prove your AI line catches urgent calls.
- Can AI Triage Emergency Plumbing Calls Out of Hours? — How an AI phone agent can sort night-time plumbing calls into wake-me-now, first-thing and advice, with scripted safety lines and a test plan.
- How to Write Case Studies From Customer Interviews Using AI — Interview questions, transcript-to-brief and draft prompts, and an approval routine for case studies where every result comes from the customer, not the model.
- How to Write Sales Call Scripts With AI That Don't Sound Scripted — Swap word-for-word scripts for an AI-built call map: openers, questions and objection cards drawn from your own best calls, then rehearsed aloud with AI.
- How to Set Up an AI Phone Agent for After-Hours Calls — Nine stages from logging your night calls to the morning handover, with tool costs, a copyable script and an electrician's setup worked through.
- AI Call Summaries: Log Every Phone Enquiry in Your CRM — Turn every inbound phone enquiry into a CRM record with named fields and a callback date, using a phone system that summarises calls for you.
- AI Tools and AI Development: The Complete 2026 Guide — the AI hub, including every tutorial in the AI-for-business series.
Sources: PCI Security Standards Council guidance and FAQs on card verification codes and audio recordings; Quo pricing page (AI call summaries and transcripts by plan, Sona AI agent); vendor facts summarised in our verified fact sheet (Otter.ai pricing, EU AI Act Article 50 transparency duties).