Compare six things: how many conversations the AI actually scores (every one, or a monthly allowance), whether it works with your helpdesk, how easily you can build your own scorecard, whether it checks AI-agent replies as well as staff, how it supports calibration and coaching, and the real per-agent price, since several vendors only quote on request.
For a small support team the first decision is usually not which dedicated tool, but whether the quality scoring built into your helpdesk is enough. Zendesk and Intercom both now score conversations with AI inside their own products, while dedicated tools such as EvaluAgent, MaestroQA and ScorebuddyCX go further on scorecards, coaching and analytics. Whichever you pick, AI scores are only useful once they've been checked against human judgement on your own conversations, so budget time for calibration as well as the licence.
What AI adds to quality assurance, and what it can't judge
Traditional QA means a team lead reads a handful of each agent's conversations a week and scores them against a scorecard. It's slow and it samples a tiny slice. AI changes three things:
- Coverage. Tools can score every conversation. Zendesk describes its AutoQA as giving "100% coverage", analysing every interaction including AI agents and voice.
- Triage. Instead of random samples, the AI surfaces outliers: angry customers, long threads, conversations that broke a rule.
- AI-agent QA. If a chatbot answers customers, its replies need checking too. EvaluAgent, for example, describes automated scoring and hallucination detection for bot conversations.
What AI scoring can't do reliably is judge whether the answer was right for your business. It can tell that an agent apologised and offered a solution; it can't know that the solution breaks your returns policy unless you've given it the policy and tested that it applies it. That's why scorecards and calibration matter more than any feature list.
The comparison criteria that matter
| Criterion | Why it matters | Ask the vendor | Answer to be wary of |
|---|---|---|---|
| AI scoring coverage | Scoring everything is the main gain over manual QA | "Is every conversation scored, or is there a monthly AI-score allowance?" | An allowance far below your monthly volume |
| Helpdesk integration | QA must see conversations where they happen | "Is there a native integration with [your helpdesk], including voice and chat?" | "Via export" or "coming soon" |
| Custom scorecards | Generic categories score what's easy, not what matters to you | "Can we write our own criteria, weights and auto-fail rules, and can AI score custom criteria?" | AI only scores the vendor's fixed categories |
| AI-agent QA | Bot replies reach customers unreviewed | "Can you score our chatbot's conversations, and flag invented answers?" | Human conversations only |
| Calibration | Keeps AI and human scores aligned | "How do we compare AI scores with our reviewers' and correct drift?" | No calibration workflow at all |
| Coaching and disputes | Scores should lead to improvement | "Can agents see reasons, and dispute a score?" | Scores visible only to managers |
| Pricing basis | Seats, reviewers, conversations or credits change the bill | "Who counts as a seat, and what's the total for our team on annual billing?" | No written quote, or every agent charged when only leads use it |
| Data handling | Customer conversations are personal data | "Where is conversation data processed, how long is it kept, and is it used to train models?" | Vague answers or no data processing agreement |
Five options a small support team is likely to shortlist
Names and prices checked on each vendor's own pages in September 2026. Several quote on request, so the published figures are only a starting point.
| Product | Published pricing | AI scoring | Worth knowing |
|---|---|---|---|
| Zendesk (built-in QA and the Workforce Engagement bundle) | Suite Professional $115 an agent a month billed annually lists automatic scoring of human and AI agents; the Workforce Engagement bundle is $50 an agent a month billed annually | AutoQA across every interaction, including AI agents and voice | Add-ons are bought for all agents on the account; only for Zendesk users |
| Intercom CX Score | Needs Fin and Intercom's Pro add-on; check current prices with Intercom | Scores every meaningful conversation 1-5 with reasons, for Fin and human replies | Closer to a customer-experience score than a full QA scorecard |
| EvaluAgent | AutoQM and Improvement from $35 a user a month; with Conversation Intelligence from $65; pay-per-conversation pricing for AI agents | Automated scoring on custom scorecards, plus bot QA and hallucination detection | Works across helpdesks and voice; volume discounts available |
| MaestroQA | Tailored pricing by quote only | Automated QA, AI coaching and conversation analytics | Get a written quote for your exact seat count before comparing |
| ScorebuddyCX (the Scorebuddy QA platform) | Foundation, Accelerate and Elite tiers, all by quote; 14-day free trial | AI auto scoring on Accelerate (500 AI scores a month) and Elite (1,000), with extra credits as an add-on | Check the monthly AI-score allowance against your conversation volume |
Two patterns stand out. Built-in options are simplest if you're already on that helpdesk, but they lock QA to it. Dedicated tools cost extra and take setup, but give you richer scorecards and work if you change helpdesk later; Freshdesk versus Zendesk for a small support team covers that upstream choice.
An AI quality score, with the errors marked
Here's why calibration matters. A kitchen fitting company's customer care team handles delivery and snagging queries. A customer writes that a cabinet door arrived cracked, three weeks after installation. The agent replies with a warm apology, promises a free replacement door "and a discount on your next order", and closes the ticket. The customer hasn't replied since.
An illustrative AI auto-score:
Conversation #48213 Overall: 92% (PASS)
Greeting and tone .............. 5/5 "Warm, empathetic apology"
Understanding the issue ........ 5/5 "Identified cracked door"
Policy adherence ............... 5/5 "Offered replacement"
Resolution ..................... 5/5 "Resolved: customer did not reply"
Grammar and clarity ............ 4/5 "Minor punctuation"
What the team lead's review found:
- Policy adherence should fail. The company's policy covers damage reported within 7 days of installation; after that, a surveyor visits first. The agent skipped the visit and promised a discount nobody can authorise. The AI scored "offered replacement" as good because it had never been given the policy.
- "Resolved" means nothing here. The customer simply hadn't replied. The same trap exists in helpdesk resolution counting: Intercom counts an "assumed resolution" after 24 hours of silence following Fin's last answer, and Zendesk closes a messaging conversation after 2 hours of inactivity by default before its AI check decides whether the resolution is verified and billable (since 18 May 2026 only verified resolutions are charged). Silence isn't satisfaction.
- Tone 5/5 is fair but beside the point. A lovely reply that commits the business to the wrong thing is worse than a plain one that follows the process.
The corrected score: policy adherence 1/5 with an auto-fail flag, resolution "pending: awaiting surveyor visit", overall FAIL, and a coaching note. The fix on the software side was to add the damage policy to the scoring instructions as a specific rule ("damage reported more than 7 days after installation: must book a surveyor visit before offering replacement") and to change the resolution criterion to require confirmation from the customer or a completed action.
Scoring the chatbot's conversations too
If an AI agent answers customers first, its conversations are where QA software earns its keep, because nobody reads them otherwise. A realistic case from the same company, whose website chatbot answers delivery questions:
Customer: When will my kitchen be delivered? Order ends 4471.
Bot: Your kitchen is scheduled for delivery on Thursday between
8am and 12pm. Our driver will call 30 minutes before arrival.
An illustrative AI-agent QA flag:
Conversation #51902 (AI agent) FLAG: possible invented detail
- "Thursday between 8am and 12pm": no delivery date found in the
order record the bot retrieved (status: "awaiting cabinet stock").
- "driver will call 30 minutes before": not in the knowledge base;
current policy is a text message the evening before.
Suggested action: human follow-up with the customer today.
Both details were invented: the order had no delivery slot, and the call-ahead promise didn't match the company's practice. Left unchecked, the customer would have taken Thursday morning off work for nothing. The fixes were a human follow-up that day, a knowledge-base line stating the real delivery-contact policy, and a bot rule that it must say "we'll confirm your date as soon as stock arrives" whenever the order status isn't "scheduled". This is the case for insisting on AI-agent QA in any shortlist: the bot makes the same mistake for every customer who asks, until someone notices.
Data questions specific to QA tools
QA software reads every customer conversation, which makes it one of the most data-hungry tools a support team buys. Before signing, get written answers to four questions:
- Which AI models score the conversations, and whose are they? Many QA tools send conversation text to a model provider to score it. Ask who that provider is and whether conversation data is used to train anything.
- How long are conversations and scores kept? Retention should match your helpdesk's, not quietly exceed it.
- Can personal details be masked? Some tools offer anonymisation of reviewed conversations; ask whether it applies before the AI sees the text or only on screen.
- Do you record calls? Voice QA depends on call recordings and transcripts, so customers need to be told calls are recorded, and the rules on recording vary by jurisdiction; check yours with your data-protection adviser.
The answers go into the same file as your data processing agreement. If a vendor can't say which model provider it uses, treat that as a failed criterion, whatever the demo looked like.
Write the scorecard before you choose the software
The scorecard is the thing you're really buying software to apply, so draft it first. A filled-in example for that six-person customer care team:
CUSTOMER CARE SCORECARD v1, [month, year]
Criterion Weight Scored by Auto-fail?
1. Correct policy applied 30% AI + human Yes, if wrong
(warranty, damage window, deposits, delivery changes)
2. Issue fully understood 15% AI
3. Clear next step and date 20% AI Yes, if missing
4. No unauthorised promises 15% AI + human Yes (discounts,
(discounts, free extras, dates we can't control) free items)
5. Tone matches house style 10% AI
6. Resolution confirmed 10% AI -
(customer confirmed OR action completed; silence doesn't count)
Calibration: team lead and one senior agent both score the same
10 conversations on the first Monday of each month; review any
criterion where AI and human scores differ on 3 or more.
Notice what carries the weight: correct policy and promises, not greetings. Once this exists, every demo becomes a simple test: can the tool score criteria 1 and 4 accurately on your past conversations? Weekly human sampling still matters alongside it; a weekly sampling routine for AI support replies shows how to choose which conversations to read.
A worked choice for a six-agent team
The kitchen fitting company's six agents handle about 2,400 conversations a month on Zendesk Suite Team ($55 an agent a month billed annually). Today the team lead manually scores 10 conversations per agent a week, 60 in all, at about 6 minutes each: 6 hours a week, roughly 10% of monthly volume. The team lead's time is costed at an illustrative $35 an hour.
| Option | Licence cost a year | Coverage | Team lead's QA time |
|---|---|---|---|
| A: Stay manual | $0 | About 10% of conversations | 6 hours a week: about $10,900 a year |
| B: Zendesk Workforce Engagement bundle | 6 × $50 × 12 = $3,600 | Every conversation | About 2 hours a week reviewing flags and calibrating: about $3,640 |
| C: Upgrade to Suite Professional | 6 × $60 extra × 12 = $4,320 | Automatic scoring, plus the other Professional features | About 2 hours a week: about $3,640 |
| D: EvaluAgent AutoQM, from $35 a user | 6 × $35 × 12 = $2,520 (starting price; get a quote) | Every conversation on custom scorecards | About 2 hours a week, plus setup: about $3,640 |
On these illustrative numbers, any of B, C or D costs less in total than staying manual, mostly because the team lead gets four hours a week back while coverage goes from a tenth of conversations to all of them. The time estimates are the soft part; the trial below is how you'd firm them up.
Between the automated options, the company leans towards B for simplicity, since it's already on Zendesk and the bundle includes workforce management it would like for rota planning. It would choose C instead if it wanted other Suite Professional features anyway, and D if it expected to leave Zendesk or needed EvaluAgent's scorecards and bot QA across several channels. The deciding test, though, is whether each tool scores the policy criterion correctly in a trial on the company's own conversations. What customer service software with AI costs per agent puts these add-on prices in the context of the whole helpdesk bill.
Where QA software purchases go wrong in small teams
- Buying before the scorecard exists. The vendor's default categories get adopted, and the team ends up optimising greetings.
- Trusting AI scores without calibration. As the cracked-door example shows, a confident 92% can hide a policy failure. Check agreement between AI and human scores monthly, criterion by criterion.
- Counting seats wrongly. Zendesk's add-ons are bought for all agents on the account; other vendors count reviewers, agents or conversations. Get the total for your team in writing.
- Running out of AI scores. Allowance-based plans can quietly fall back to partial coverage. Compare the allowance with your monthly volume, including chatbot conversations.
- Using scores as a stick. If agents learn that scores drive discipline, they write for the scorer, not the customer. Use them to start coaching conversations and let agents dispute them.
- Forgetting the bot. If an AI agent answers customers, its conversations need the same scorecard. A policy error repeated by a bot a hundred times a day is the costliest kind.
A 30-day trial plan for two tools
- Week 1: prepare. Finalise the scorecard. Pick 50 past conversations and have the team lead score them by hand, including ten known problem cases like the cracked door.
- Week 2: score the same 50 in each trial tool. Configure the scorecard and policy rules in each, then compare AI and human scores criterion by criterion. Count how many of the ten problem cases each tool caught.
- Week 3: run live. Let each tool score a week of real conversations. Time how long the team lead spends reviewing flags.
- Week 4: decide. Choose on three numbers: problem cases caught, agreement with human scores on the weighted criteria, and weekly review time. Then check the written quote and the data processing terms before signing.
A trial built this way answers the question the demos never do: does this tool score your policies correctly on your conversations? Everything else on the comparison table is secondary to that. For the vendor-side checks that go with any purchase, a small-business scorecard for evaluating AI software vendors covers security, support and contract terms.
Choosing QA software: questions support leads ask
Can AI auto-scoring replace human QA reviews?
Not entirely. AI scoring covers every conversation for things it can judge reliably, such as structure, tone and whether required steps happened. A person still needs to check whether answers were actually correct, calibrate the AI against human scores regularly, and handle anything that goes to coaching or a formal conversation with an agent. Think of AI as the triage, and people as the judges.
How many conversations should a small team score by hand?
Enough to keep the AI honest and to coach properly. A workable starting point is a few conversations per agent each week, chosen partly at random and partly from AI-flagged outliers, plus a monthly calibration session where two people score the same ten conversations and compare. Adjust once you see how often human and AI scores disagree.
Should agents see their AI quality scores?
Yes, with context. Agents who can see scores and the reasons behind them improve faster and trust the system more. Make it clear that scores start coaching conversations rather than decide pay or discipline on their own, and give agents a way to dispute a score. A score nobody can question quickly becomes a score nobody believes.
Further reads
- How to Run a Two-Week AI Tool Trial Before You Commit — Run the trial described here with a proper two-week plan.
- How to Check AI Is Doing Good Work, Not Just Fast Work — Measure whether AI support is good, not just fast.
- 9 AI Chatbot Mistakes That Lose Small Businesses Customers — The chatbot failures your AI-agent QA should catch.
- How to Hire Your First Customer Service Person Alongside AI — Staffing a small team that works next to AI agents.
- How to Write a House Style Guide Your Team and AI Both Follow — The style rules your QA scorecard should check against.
- Tidio vs Intercom: Which AI Chat Tool Suits a Small Business? — If your helpdesk choice is still open.
- Best AI Customer Support Software for Small Teams in 2026 — Nine support tools compared on what small teams actually pay for AI: per-resolution fees, credits and bundles, each with a worked example.
- Can AI Handle Customer Complaints Without Making Them Worse? — Where AI speeds up complaint handling, the five ways its replies inflame customers, and one hotel complaint followed from angry email to resolved.
- How to Review AI Call Transcripts for Quality and Compliance — How a small helpdesk reviews AI call transcripts: a filled-in scorecard, sampling numbers, an AI pre-screen prompt and the compliance checks that matter.
- How to Run a Monthly AI Quality Review in 30 Minutes — A fixed 30-minute monthly review for AI emails, chatbots, posts and automations: random samples, a copyable scoring sheet and a rule for fixes that get done.
- AI Tools and AI Development: The Complete 2026 Guide — the AI hub, including every tutorial in the AI-for-business series.
Sources: Zendesk pricing page, Zendesk QA product page and help article on buying workforce engagement add-ons; EvaluAgent pricing page; MaestroQA pricing page; ScorebuddyCX pricing page; Intercom help articles on CX Score; Intercom, Zendesk and Help Scout resolution definitions and prices as of September 2026.