Run a monthly AI quality review as a fixed 30-minute meeting: pull five random examples from each AI job (emails, chatbot chats, posts, automations), score each for accuracy, tone, completeness and policy on a simple 0-2 scale, check automation error logs, then agree no more than three fixes with an owner and date. Prepare the sample beforehand.
Thirty minutes is enough only if the review deals in evidence rather than impressions. "The chatbot seems fine" is an impression. "Two of five chats promised same-day printing, which we don't offer" is evidence you can act on. Most monthly reviews fail by trying to cover everything; this one looks at a small random sample and a handful of numbers, and it ends with decisions.
The worked example is an illustrative print shop with six staff that uses AI in four places: drafting quote emails, a chat assistant on its website, social posts and product descriptions, and a Zapier automation with an AI step that turns online order forms into job tickets. The same structure works for any business with a few AI jobs running.
The review sits between two other habits. Day-to-day, people check AI output before it goes out; when something slips through, it goes in an error log. The monthly review is where the log, the numbers and a fresh sample meet, so patterns get fixed instead of patched one at a time. If you don't have an error log yet, the tutorial on keeping an AI error log sets one up in an afternoon.
What happens before the 30 minutes start
The meeting stays short because the preparation happens beforehand, ideally by a single person spending about 20 minutes on the last working day of the month. Prepare three things:
- The sample. Five random items from each AI job, pasted into one shared document. Random matters: if you pick examples yourself, you'll unconsciously pick the ones you remember, which are usually the best or the worst.
- The numbers. A few figures per job, the same ones every month, so trends show up: complaints mentioning AI output, chatbot handovers to a person, automation runs that errored, error-log entries.
- Vendor changes. Any notice from a tool you use about a price, feature, model or privacy change. These arrive by email and get ignored unless someone is looking for them.
For random picks without special software: export the month's items to a spreadsheet, add a column with =RAND(), sort by it and take the top five. For emails, the print shop exports the "AI draft" folder; for chats, its chat tool's transcript export; for the automation, the task history.
The agenda, minute by minute
| Minutes | What happens | Output |
|---|---|---|
| 0-5 | Read the numbers against last month. Anything up sharply gets a question mark. | One line per AI job: steady, better or worse |
| 5-20 | Score the sample: five items per job, about 45 seconds each, using the scoring sheet | A score per item and a note on every zero |
| 20-25 | Pick the fixes: what caused each zero, and which three problems matter most | No more than three actions |
| 25-30 | Record: owner, due date and how you'll know the fix worked; glance at vendor changes | The review log updated |
Keep the clock visible. The discipline of 45 seconds per item forces you to score what's in front of you rather than debate what the AI might do in other cases.
A scoring sheet you can copy
Four criteria, each scored 0, 1 or 2. Any item scoring 0 on accuracy or policy counts as a failure whatever its total, because a well-written wrong answer is still wrong.
| Criterion | 2 | 1 | 0 |
|---|---|---|---|
| Accuracy | Every fact, price, date and product detail correct | A minor slip a customer wouldn't act on | A wrong fact someone could act on |
| Tone | Sounds like your business | Generic but acceptable | Wrong for the situation (cheery about a complaint, stiff with a regular) |
| Completeness | Answers the actual question and says what happens next | Answers it but leaves an obvious follow-up | Misses the question or the key detail |
| Policy | Within what you offer and promise | Vague where it should be specific | Promises something you don't offer, or breaks a rule |
Filled in, three rows from the print shop's sheet look like this (illustrative):
| Item | Acc. | Tone | Comp. | Policy | Note |
|---|---|---|---|---|---|
| Quote email: 250 A5 flyers | 2 | 2 | 1 | 2 | Didn't mention artwork deadline for Friday delivery |
| Chat: "can you print today?" | 0 | 2 | 2 | 0 | Promised same-day printing; we need 48 hours |
| Instagram post: wedding stationery | 1 | 1 | 2 | 2 | Seven hashtags; Instagram caps posts at five |
Two scorers should agree within a point on most items. If they don't, the criteria need a sentence of clarification, not a longer discussion. If your AI jobs are mostly support replies and you want a deeper weekly version for those, the weekly sampling routine for AI support replies goes into more detail.
One month at the print shop, scored
Here is the illustrative September review in full: 20 items sampled across four jobs, plus the numbers.
| AI job | Items | Failures (a zero on accuracy or policy) | Average score (out of 8) | Numbers this month |
|---|---|---|---|---|
| Quote emails | 5 | 0 | 7.2 | 1 complaint (wrong paper weight quoted) |
| Website chat | 5 | 2 | 5.4 | 38% of chats handed to staff, up from 29% |
| Posts and product copy | 5 | 0 | 6.6 | 2 posts edited after publishing |
| Order-form automation | 5 runs | 1 | — | 23 errored runs, up from 2 |
Three findings came out of 15 minutes of scoring:
- The chat assistant was promising same-day printing. Two of five sampled chats said "we can usually print the same day". The source was an old FAQ page from a promotion, still in the assistant's knowledge base. The rise in handovers had the same cause: customers who were told "same day" then asked staff to confirm.
- The automation's errors had jumped from 2 to 23. Someone had renamed a field on the order form from "Paper stock" to "Paper type", so the AI step received nothing and guessed. One of the five sampled runs had created a job ticket for the wrong stock.
- Posts were fine but careless with platform rules. Seven hashtags on one post, when Instagram has capped posts at five since December 2025.
The three actions agreed, with owners: remove the old FAQ from the chat assistant's sources and add an explicit "orders take 48 hours" rule (the owner, by Friday); map the renamed form field and replay the 23 failed runs (the office manager, by Wednesday); add "maximum five hashtags" to the social post prompt (whoever writes posts, today). Each had a check for next month: no "same day" in sampled chats, errored runs back under five, no post over five hashtags.
The quick sum on whether 30 minutes a month is worth it: that's six hours a year. The single wrong-stock job ticket cost the shop a reprint of 500 wedding invitations, roughly $180 in materials and three hours of press and finishing time. One catch like that a year pays for the review several times over.
Five causes behind a zero, and the fix each one needs
The five minutes set aside for picking fixes go faster if you know the usual causes. Almost every failure the print shop has found traces back to one of these, and the cause decides the fix. Rewriting the prompt is the reflex, but it's the right answer less than half the time.
| Cause | How it shows up in the sample | The fix |
|---|---|---|
| Stale source material | The AI repeats an old price, policy or offer word for word | Update or remove the source document; the prompt is fine |
| Missing rule | Output is reasonable but breaks something nobody wrote down (hashtag limits, lead times) | Add one line to the prompt or saved instructions |
| Changed input | An automation's output goes wrong from a specific date | Find what changed upstream (a form field, a column name) and map it again |
| Tool or model change | Output gets longer, more formal or differently structured across the board | Re-run your standard test inputs and adjust instructions |
| Skipped human check | An error that the day-to-day check should have caught | Ask why the check was skipped; usually time pressure or unclear ownership |
In the September review above, the same-day promise was stale source material, the automation errors were a changed input and the hashtags were a missing rule. Three different fixes, none of which would have come from "improve the prompt".
Automations fail quietly: the checks that catch them
AI steps inside automations are the easiest to forget, because nobody reads their output directly. The monthly review should always look at run history, not just a sample of results, because the platforms behave in ways that hide problems:
- Zapier automatically pauses a Zap when 95% of its runs error over seven days, which stops the damage but also stops the work. And if you've added an error handler, Zapier doesn't send the usual error emails when the handler runs, so a failing step can go unnoticed.
- Power Automate switches flows off after 14 days of continuous failure, and flows that haven't triggered for 90 days are switched off too unless you have a Premium licence.
- Make lets you attach error handlers (now named Skip, Retry, Resume, Commit and Rollback). Skip in particular quietly drops the item that failed, which is exactly what you need to see in a monthly review.
So the prep for any automation is: count errored runs this month versus last, open two or three failures to see why, and sample five successful runs to check the AI step's output made sense. A broader look at automations nobody owns is in the automation audit for Zapier and Make.
Chatbot numbers that look good and mean little
If a chat assistant is one of your AI jobs, be careful which number you track. Vendors report "resolutions", and their definitions are generous. Intercom counts an "assumed resolution" for Fin when the customer goes quiet for 24 hours after Fin's last answer. Zendesk's default closes a messaging conversation after two hours without activity, or 72 hours for email and web forms, then an AI check decides whether it counts as a verified, billable resolution (since 18 May 2026 hand-offs and unverified closes are free). A customer who gave up and phoned you instead counts as resolved in both.
For the review, track handovers to a person and read the sample. A rising handover rate, like the print shop's jump from 29% to 38%, often points to one specific wrong answer rather than a general decline. If you run a Shopify store with the free AI agent in Shopify Inbox, remember it can use web search as a secondary source, so include two or three of your own policy questions in the monthly sample and check its answers match your policies. The tutorial on measuring whether your AI chatbot is working covers the fuller set of metrics.
Vendor changes to scan for each month
The last five minutes include a quick look at what changed in your tools, because a vendor change can break a working setup overnight. Examples from the last year show how varied these are:
- Retirements. Any custom GPT your team relies on will stop working on 11 December 2026, when OpenAI retires them. Microsoft retired Excel's
=COPILOT()worksheet function on 14 September 2026, leaving cached results in place but allowing no new formulas. Anything built on either needs a replacement plan. - Renamed products. ChatGPT agent became ChatGPT Work; NotebookLM became Gemini Notebook. Renames rarely break things but confuse staff instructions.
- Privacy defaults. Vendors sometimes change what they keep by default. When one does, your data handling rules may need updating.
- Price and credit changes. A credit-based AI feature that gets more expensive per use shows up first as a larger bill.
If you use Microsoft 365 Copilot, the admin centre's Copilot usage report (active users over 7, 28, 90 or 180 days) is worth a glance at the same time: licences nobody uses are a cost issue, and a sudden drop in use sometimes means people have lost trust in the output.
Adapting the review for other businesses
The agenda stays the same; the AI jobs and the numbers change.
- An online clothing shop samples chat answers about returns and sizing, AI product descriptions and the abandoned-cart emails. Its key number is returns with the reason "not as described", which rises when product copy drifts.
- A software reseller samples AI-drafted renewal notices and the monthly licence summaries an AI step writes from its distributor reports. Accuracy dominates the scoring: one wrong licence count is a billing dispute.
- An events company samples AI-drafted proposals and supplier emails. Its failures are usually completeness: dates, guest numbers or dietary notes missing from a proposal.
- A handmade jewellery seller working alone does a 15-minute version: five product descriptions, five customer replies, one number (messages asking something the listing should have answered).
Turning findings into fixes without a backlog
The rule that keeps this to 30 minutes is the limit of three actions a month. Everything else found goes into the error log to be looked at next time. Three fixes that actually happen beat twelve that sit on a list.
Each action needs three things written down: who, by when, and how next month's review will tell whether it worked. The print shop's review log is a single spreadsheet tab with one row per action. Two rows from October, the month after the review above (illustrative):
| Month | AI job | Problem | Fix | Owner, due | Check | Result next month |
|---|---|---|---|---|---|---|
| Sep | Website chat | Promised same-day printing | Old FAQ removed; 48-hour rule added | Owner, Fri 4th | No "same day" in sampled chats | 0 of 5; handovers down to 30% |
| Sep | Order automation | Renamed form field broke AI step | Field remapped; 23 runs replayed | Office manager, Wed 2nd | Errored runs under 5 | 1 errored run |
The last column is what makes the log useful: it shows whether fixes worked, not just that they were made. After a few months, that log is also the best evidence you have of what AI is doing well and badly in your business, which is useful when deciding whether to keep, change or cancel a tool. The tutorial on reviewing an AI tool after 90 days uses exactly that kind of record, and the one on checking that AI is doing good work, not just fast work covers the thinking behind the scoring.
Running the review: common questions
Who should attend a monthly AI quality review?
Two people is ideal in a small business: whoever owns the AI tools and one person who deals with customers directly. The customer-facing person spots wrong tone and missing information quickly. More than three people turns a 30-minute review into an hour of discussion.
Is five examples per AI job enough to find problems?
Five random examples won't catch rare errors, but they reliably catch problems that affect a meaningful share of output, and those are the ones worth fixing first. Rare but serious errors are what the error log and customer complaints are for. If one job keeps producing low scores, sample ten of it the following month.
What if we find nothing wrong for several months?
Keep the review, but shorten it or move one AI job to every other month. A clean run usually means the prompts and guardrails are working, and the review is what will tell you when a tool update or a new product breaks them.
Should the review include how much time AI is saving?
Once a quarter, yes: time a few tasks with and without AI, or compare output volumes. Monthly, stay focused on quality, because a tool that saves time while producing wrong answers is costing you more than it looks.
Further reads
- How to Set Up Human Review for AI Work Without Slowing Down — Set up human review for day-to-day AI work without slowing down.
- AI Incident Response Plan for Small Businesses (With Template) — What to do when the review finds a mistake that reached customers.
- How to Audit Your AI Subscriptions and Cut Wasted Spend — Add a quick subscription check to the quarterly version of the review.
- Customer Service QA Software With AI: What to Compare — When QA software is worth it for reviewing support replies.
- What Is an AI Audit? What It Covers and What It Costs — How a fuller AI audit differs from a monthly review.
- Chatbot Guardrails: Stop AI Promising What You Don't Offer — Fix chatbot answers that promise what you don't offer.
- What Are Diners Really Saying? Using AI to Analyse Your Reviews — A method for turning months of restaurant reviews into a few counted, checked themes, with a codebook, a tagging prompt and a worked nine-month example.
- How to Reply to Hotel Reviews With AI Without Sounding Canned — Why AI review replies read as templated, and the voice sheet, reply shapes and prompt that make each one sound written by the person who runs the place.
- Gym Automation Mistakes That Increase Member Churn — Ten ways automated messages quietly drive gym members away, each with a real-looking example, the fix, and a check you can run this week.
- Photo Checklists and Quality Control for Cleaning Teams Using AI — Set checkpoints a camera can prove, a fixed photo set per clean, an AI pass that flags which jobs need a supervisor, and clear privacy rules for clients' homes.
- AI Content Approval Workflow: Draft, Check, Sign Off — A three-stage approval workflow for AI-drafted content, with risk lanes, a filled-in check sheet, sign-off records and the numbers for a small agency.
- How to Catch Outdated Information in AI Answers — How to catch stale facts in AI answers: the topics that date fastest, a prompt that exposes dates, a five-minute check and real 2026 examples.
- How to Review AI Call Transcripts for Quality and Compliance — How a small helpdesk reviews AI call transcripts: a filled-in scorecard, sampling numbers, an AI pre-screen prompt and the compliance checks that matter.
- How to Audit Your Website for AI Content That Needs Fixing — A content audit for sites with AI-drafted pages: inventory, leftover searches, a five-check score, keep-fix-merge decisions and a translation agency's numbers.
- How Accurate Is ChatGPT? What Owners Should Expect by Task — Where ChatGPT is dependable, where it guesses, and a 20-case test that measures its accuracy on your own work before you trust it with customers.
- AI vs Human Transcription: Which Is Worth Paying For? — When AI transcription is good enough, when a person is worth $1.99 a minute, how to test accuracy on your own recordings, and the hybrid in between.
- AI Tools and AI Development: The Complete 2026 Guide — the AI hub, including every tutorial in the AI-for-business series.
Sources: series fact sheet (checked September 2026) for Zapier auto-pause and error-handler behaviour, Power Automate flow shut-off rules, Intercom Fin and Zendesk resolution definitions, the Microsoft 365 Copilot usage report, OpenAI's custom GPT retirement date, Microsoft's retirement of Excel's COPILOT function, and the Instagram hashtag cap.