Track AI with a handful of numbers in four groups: time (hours saved, time per task, response time), quality (edit rate, error rate, escalations, customer satisfaction), adoption (active users, share of eligible work going through AI) and money (cost per task, spend against budget, revenue effects). Most small businesses need four to six, each compared with a pre-AI baseline.
A KPI, or key performance indicator, is simply a number you check regularly to see whether something is working. Below, each of the 12 gets a plain definition, a way to measure it without a data team, an illustrative example, and the way it can mislead you. After the list there's a guide to choosing your four to six, and a one-screen monthly sheet.
Every one of these needs a "before" figure to mean anything. If you haven't measured the current process yet, set a baseline first; two weeks is usually enough.
Time: is AI actually saving hours?
1. Net hours saved per week
What it is: the time a job takes now, including checking and correcting AI output, subtracted from the time it took before, multiplied by volume. How to measure: baseline minutes per item minus current minutes per item, times items per week. Example: a pet shop's product descriptions drop from 15 minutes to 5 minutes including checking; at 40 new lines a month, that's about 6.7 hours a month. Watch-out: quoting gross time ("the AI writes it in seconds") instead of net, and counting hours that vanish into a quieter afternoon rather than going on other work.
The gap between gross and net is usually bigger than people expect. An illustrative four-person cleaning company timed its quote follow-up emails: 9 minutes each by hand. The AI draft appears in about 20 seconds, which is the figure the keenest user quoted in the team meeting. Timed properly, with reading the draft, correcting the address and adding a line about the specific clean, it took 4 minutes. At 35 follow-ups a week, the gross claim was about 5 hours saved; the net figure was just under 3. Three hours a week is still well worth having, and it's the only number that belongs on the sheet.
2. Time per task
What it is: the median minutes of hands-on work for one item. How to measure: time ten items a day for one week each month. Example: median reply time falls from 6 to 3 minutes. Watch-out: if staff send only the easy items through AI, time per task looks great while the hard ones pile up elsewhere. Pair it with coverage (number 9).
Use the median, not the average, because one awkward item distorts a small sample. Say a bookkeeping practice times its first ten client queries on a Tuesday and gets 3, 4, 2, 5, 3, 25, 4, 3, 4 and 2 minutes. The average is 5.5 minutes, dragged up by one 25-minute query about a missing invoice. The median is 3.5 minutes, which is what a typical query actually takes. Keep the 25-minute one in your notes, though: if long queries keep turning up, they deserve their own look.
3. Response or turnaround time
What it is: how long a customer waits, from their message arriving to your first real reply. How to measure: email or message timestamps on a sample of 30 a month; use the median. Example: an optician's non-clinical enquiries go from a median of 20 working hours to 6. Watch-out: saving effort doesn't automatically cut waiting time. If staff still clear the inbox twice a day, turnaround won't move much, and that's worth knowing.
Split the figure if your week isn't uniform. An illustrative florist's overall median fell from 5 hours to about 2 after AI drafting started, which looked like a clear win. Split by day, weekday messages had dropped to about 40 minutes, while weekend orders still waited around 26 hours, because nobody opened the inbox between Saturday lunchtime and Monday. For a business where many orders are for weekend deliveries, the weekend number was the one that mattered, and AI drafting alone was never going to fix it.
Quality: is the work still good?
4. Edit rate
What it is: the share of AI drafts used with no or only minor edits. How to measure: the person checking marks each item as none, minor, heavy or discarded; count a sample of 30 to 50 a month. Example: 45% no or minor edits in month one, 70% by month three as instructions improve. Watch-out: a sudden jump to nearly 100% "no edits" often means people have stopped reading drafts properly, not that the AI got perfect.
Agree what "minor" means before anyone starts marking, or each person will draw the line differently. A workable rule: minor means no fact changed and no more than a sentence or two reworded; heavy means any fact corrected or most of the draft rewritten; discarded means the draft was abandoned. A few lines from a typical log:
Draft 14 minor changed "Hi there" to the customer's name
Draft 15 heavy wrong delivery date for a bulk order
Draft 16 none
Draft 17 discarded customer was asking about a different order
Count discards as failures, not as missing data. A draft nobody could use cost checking time and saved nothing.
If the log runs to a few hundred lines, a chat assistant can do the tallying, provided you check its arithmetic. Paste the month's log with any customer names removed and ask:
Below is a month of edit-level marks for AI drafts, one per line.
Count how many are none, minor, heavy and discarded, and give the
share of none plus minor as a percentage of ALL lines, including
discarded. Then list the three most common reasons given for heavy
edits, quoting the notes.
[paste log]
An illustrative reply for a 48-line log:
None: 19. Minor: 14. Heavy: 11. Discarded: 3. Share of none or minor: 33 of 44, 75%. Most common heavy-edit reasons: wrong delivery date (4), missing bulk discount (3), wrong branch address (2).
The counts add to 47, not 48, and the percentage is out of 44 rather than all 48 lines: the assistant has quietly left out the discards, and missed a line, despite being told. The missing line turns out to be another heavy edit, so the true figure is 33 of 48, about 69%, which falls just short of a 70% target rather than clearing it. The reasons list is the useful part and usually reliable; the arithmetic is what you check, by adding the four counts before you copy anything onto the sheet.
5. Error rate
What it is: errors found per 100 AI outputs in a checked sample, such as a wrong price, a missed question, an invented fact. How to measure: check a fixed random sample of 20 to 30 outputs each month against the source information. Example: 4 errors per 100 in month one, 1 per 100 in month three. Watch-out: it only counts what you look for. Check the same kinds of thing each month, or the trend means nothing. Checking AI is doing good work, not just fast work shows how to build the checklist.
A filled-in month from an illustrative dog-grooming salon shows what "check the same things" means in practice. Its checklist had four questions for every sampled reply draft: does the price match the price list, are dates and times right, is every question the customer asked answered, and is nothing promised that the salon doesn't offer? Of 25 drafts checked in March, 23 passed everything. One quoted the small-dog price for a large cockapoo cross, and one ignored a second question tucked into the last line of the email. That's 2 errors in 25, or 8 per 100, and both pointed at specific fixes: a line in the price list for crosses, and "answer every question in the message" added to the instructions.
6. Escalation or handover rate
What it is: for chat assistants and sorting tools, the share of conversations or items passed to a person. How to measure: most chat tools report it; otherwise count the "pass to staff" label in your log. Example: 20% of website chats handed to staff. Watch-out: it can mislead in both directions. Too high and the AI isn't saving much; too low and it may be answering questions it should pass on. If you pay per resolved conversation, as with HubSpot's Customer Agent at 50 credits (about $0.50) each, check exactly what counts as "resolved". For chat assistants in particular, measuring whether your AI chatbot is working goes further.
Here's the "too low" direction in practice. An illustrative garden centre's website chat assistant reported handing only 5% of conversations to staff, which looked excellent. Reading 20 transcripts told a different story: when customers asked whether a plant was safe for cats, it answered from general knowledge instead of passing the question on, sometimes with a confident answer staff wouldn't have given without checking. The fix was a short list of topics that must always go to a person (pet and child safety, pesticides, refunds), after which the handover rate rose to 11%. The higher figure was the healthier one.
7. Customer satisfaction on AI-touched work
What it is: how customers react to work the AI helped with. How to measure: complaints per 100 AI-assisted interactions, review mentions, or a one-question follow-up ("Did we answer your question?"). Example: complaints about wrong information stay at or below the baseline of 2 a month. Watch-out: small businesses have small numbers, so one bad week looks like a crisis. Look at the quarterly trend. Measuring customer reaction after introducing AI covers survey wording.
The one-question follow-up is the cheapest of these to run. An illustrative tutoring agency added "Did this reply answer your question? Yes / Partly / No" to one reply in five for a quarter and got 41 responses: 33 yes, 6 partly and 2 no. The numbers were fine; the value was in the "partly" answers. Five of the six were parents asking about discounts for siblings, which the agency's reference notes didn't cover, so the AI had answered around the question. One added paragraph fixed it.
Adoption: are people using it?
8. Active use rate
What it is: people who used the AI tool in a given week, divided by paid seats. How to measure: most business plans give admins some usage view. In Microsoft 365, the Microsoft Copilot usage report in the admin centre (under Reports, then Usage) shows active users and an active-users rate over the last 7, 28, 90 or 180 days. Example: 5 of 8 seats active weekly is 63%. Watch-out: logging in isn't the same as using it well. A high rate with no time saving means people are dabbling, not working differently.
Comparing two windows tells you more than either alone. Say a small architecture practice with 12 Copilot seats sees 9 active users over the last 28 days but only 4 over the last 7. People tried it in the first few weeks and most have since stopped, which a single monthly figure of 75% would have hidden. That's the moment to ask the five who drifted away what they tried and where it let them down, before deciding whether they need a seat at renewal.
9. Coverage
What it is: the share of eligible items that actually went through the AI-assisted route. How to measure: count eligible items (for example all non-clinical emails) and the number where staff used the AI draft. Example: an optician receives 60 eligible emails in a week and uses drafts for 45: 75% coverage. Watch-out: low coverage alongside glowing reports usually means staff skip the tool on awkward cases. Ask which items they skip and why.
At the optician, asking about the 15 skipped emails took five minutes and settled it. Eleven were contact-lens re-orders. The reference files had no lens prices or supply lengths, so the drafts for those were useless and staff had quietly stopped opening them. Adding the lens price list moved coverage to about 90% the following month, and it turned a vague "the AI isn't great at some things" into a specific, fixable gap.
Money: is it paying its way?
10. Cost per completed task
What it is: monthly AI costs for a workflow divided by tasks completed with AI. How to measure: subscriptions and platform charges for that workflow, divided by its volume. Example: $70 a month across 400 tasks is about 18 cents a task, against a staff cost of roughly $1.67 for five minutes at $20 an hour. Watch-out: it ignores checking time, which is why it belongs alongside net hours saved, not instead of it.
With a fixed-price plan, cost per task depends heavily on volume, so a low-volume month makes a workflow look expensive. Take an automation on Zapier Professional at $29.99 a month (billed monthly) for 750 tasks, where each run uses two ordinary action steps plus an AI step on the three-task tier: five tasks a run, so up to 150 runs a month. At 150 runs, the platform costs about 20 cents a run. At 60 runs in a quiet month, it's 50 cents. Neither figure is wrong, but compare like with like: the same month last year, or a rolling three-month figure.
11. Spend against budget and unused seats
What it is: total monthly AI spend across every tool, compared with your budget, plus the number of seats unused for 30 days. How to measure: card statements and invoices each month, matched against your tool list. Example: $240 spent against a $250 cap, with two seats unused since last month. Watch-out: AI charges hide inside other subscriptions and in usage-based credits, so check invoices, not just the list of tools. Auditing your AI subscriptions helps find them.
A realistic find from a first spend check: an illustrative bakery with three AI-related lines it knew about (two seats on a business AI plan and an automation plan) found a fourth on the email marketing tool's invoice, an AI writing add-on at about $15 a month that someone had ticked during a free trial months earlier. Nobody used it. It hadn't appeared on the tool list because it wasn't a separate tool, which is exactly why this KPI is measured from invoices.
12. A revenue-linked outcome
What it is: one sales number the AI should move, such as enquiry-to-booking conversion, rebooking rate, or sales from products listed faster. How to measure: the same calculation, from the same system, as the baseline. Example: enquiries converted to bookings rise from 30% to 34%. Watch-out: sales move for many reasons. Compare the same season, and treat any rise as partly caused by AI, not wholly.
The fairest version compares like with like twice over. Say an illustrative bike repair shop converts 38% of service enquiries into bookings in spring last year and 44% this spring, after introducing AI-drafted replies. It also cut its basic service price in April, so not all of that is AI. Within this spring, though, enquiries answered with an AI draft converted at 45% and those answered by hand (evenings, when the owner replied from a phone) at 41%. That smaller gap, about four points, is the more honest estimate of what the drafting added.
Picking your four to six
Start from what you wanted AI to achieve, then take the core KPIs for that goal plus one guardrail that must not get worse.
| Your main goal | Core KPIs | Guardrail |
|---|---|---|
| Save staff time on admin | 1 Net hours saved, 2 Time per task, 4 Edit rate | 5 Error rate |
| Reply to customers faster | 3 Turnaround, 9 Coverage | 7 Customer satisfaction |
| Handle more enquiries without hiring | 1 Net hours saved, 3 Turnaround, 12 Conversion | 7 Customer satisfaction |
| Keep AI spending under control | 8 Active use, 10 Cost per task, 11 Spend and unused seats | 1 Net hours saved |
| Run a customer-facing chat assistant | 6 Escalation rate, 7 Customer satisfaction | 5 Error rate |
A worked choice: the dog-grooming salon above wanted to handle more enquiries without taking on another receptionist. From the table, it took net hours saved, turnaround and conversion, plus customer satisfaction as the guardrail, and added error rate because its drafts quote prices. It dropped cost per task: it pays for one flat subscription, so the figure would only ever restate the bill. Five KPIs, all collected from the inbox, the booking system and a 25-draft sample, in about 20 minutes a month.
Very low volumes need a different approach. An illustrative surveying practice uses AI to draft cover letters for about eight reports a month. At that volume, one bad letter moves the edit rate by more than ten points and the error rate by about 12 per 100, so month-to-month percentages are noise. Track it quarterly instead, as plain counts ("2 of 24 letters needed a fact corrected this quarter"), and keep a single note of what each error was. The counts are honest about how little evidence there is, and the notes tell you what to fix.
If a KPI would take more than 15 minutes a month to collect, and it isn't central to your goal, drop it. A short list you actually keep up beats a long one abandoned in month two.
A one-screen monthly KPI sheet
One row per KPI, one sheet per month. Here's how an illustrative pet shop might fill it in for its product-description workflow:
WORKFLOW: Product descriptions for new lines MONTH: April
KPI | Baseline | Target | This month | Last month | Better?
Net hours saved/month | 0 | 6 | 6.5 | 5.0 | yes
Time per description | 15 min | 5 min | 5 min | 7 min | yes
Edit rate (none/minor) | n/a | 70% | 68% | 55% | yes
Errors per 100 checked | 3 | <= 3 | 2 | 4 | yes
Cost per description | n/a | < $1 | $0.63 | $0.83 | yes
Action this month: add feeding-guide table to reference files
Owner: web shop manager Reviewed with: AI lead, 2 May
The "action" line matters most. A KPI sheet that never leads to a change is just paperwork. Each month, pick one thing to adjust based on the numbers, and see whether it moved the next month's figures. If you'd like these on a live dashboard rather than a sheet, building a KPI dashboard without a data team shows how.
How often to look, and who looks
Monthly is right for most small businesses. Weekly numbers bounce around too much to act on, except in the first month of a new workflow, when a quick weekly glance catches set-up problems early. Give it 20 minutes: the AI lead and each workflow owner go through the sheet together, agree one action per workflow, and note anything the owner needs to decide.
Taking the pet shop's April sheet above, those 20 minutes might go like this. Five minutes on the numbers: everything moved the right way, and edit rate at 68% is two points under target but up 13 on March, so the direction matters more than the gap. Five minutes on the one figure that could mislead: cost per description fell from $0.83 to $0.63 partly because the shop listed more new lines in April, not only because the drafts improved, so nobody should quote it as an efficiency gain. Five minutes on the heavy edits, most of which were feeding amounts copied wrongly from supplier sheets, which is where the "feeding-guide table" action came from. The last five minutes go on anything for the owner: here, nothing, so the sheet goes in the folder and the quarterly meeting gets three months of them together.
Once a quarter, take the sheets to the owner. That's the meeting where you decide whether each workflow stays, changes or stops, and whether the seats you're paying for still match the people using them.
Further reads
- How to Calculate AI ROI for Your Business (Worked Example) — Turn several of these KPIs into a single return figure.
- How to Measure Time Saved After Rolling Out AI in a Small Firm — A closer look at getting the time KPIs right.
- How to Set Up Human Review for AI Work Without Slowing Down — The review step that produces your edit-rate data.
- How to Measure Whether Copilot Is Paying for Itself — Applying these metrics to Microsoft 365 Copilot seats.
- How to Review an AI Tool After 90 Days: Keep, Fix or Cancel — Where your monthly KPIs feed a keep or cancel call.
- Who in Your Team Actually Needs a Paid AI Licence? — Act on a poor active-use figure by trimming seats.
- How to Write a One-Page AI Strategy for Your Business — The seven boxes a one-page AI strategy needs, a filled-in garden centre example, and five tests that show whether your page will guide real decisions.
- Why AI Saves Time but Not Money, and How to Capture the Gain — Why hours saved by AI vanish from the accounts, the four routes that turn them into money, and a worked example from a small brewery.
- How to Measure Whether AI Is Improving Your Marketing Results — Six stages for telling whether AI is improving your marketing or just producing more of it, with a noise check, tracking sheet and worked quarter.
- How to Track Whether AI Search Is Sending You Customers — Build a practical AI search measurement sheet that separates visibility, visits, enquiries and customers without double counting.
- How to Get a Weekly Business Summary Emailed to You by AI — Get a Monday email with your five key numbers and a short AI commentary. Covers the numbers tab, Zapier and Make builds, assistant scheduled tasks and checks.
- Did Your AI Pilot Work? How to Set Success Criteria That Hold Up — How to write AI pilot success criteria that can actually fail: a baseline margin, a grading rubric, guardrails, sample sizes and pass, extend or stop rules.
- AI Tools and AI Development: The Complete 2026 Guide — the AI hub, including every tutorial in the AI-for-business series.
Sources: Microsoft Learn (Microsoft 365 admin centre Copilot usage report); HubSpot pricing and credits pages. Checked September 2026.