How to Run a Monthly AI Quality Review in 30 Minutes

Coding Liquids tutorial cover featuring Sagnik Bhattacharya for How to Run a Monthly AI Quality Review in 30 Minutes.
Coding Liquids tutorial cover featuring Sagnik Bhattacharya for How to Run a Monthly AI Quality Review in 30 Minutes.

Run a monthly AI quality review as a fixed 30-minute meeting: pull five random examples from each AI job (emails, chatbot chats, posts, automations), score each for accuracy, tone, completeness and policy on a simple 0-2 scale, check automation error logs, then agree no more than three fixes with an owner and date. Prepare the sample beforehand.

Thirty minutes is enough only if the review deals in evidence rather than impressions. "The chatbot seems fine" is an impression. "Two of five chats promised same-day printing, which we don't offer" is evidence you can act on. Most monthly reviews fail by trying to cover everything; this one looks at a small random sample and a handful of numbers, and it ends with decisions.

Follow me on Instagram@sagnikteaches

The worked example is an illustrative print shop with six staff that uses AI in four places: drafting quote emails, a chat assistant on its website, social posts and product descriptions, and a Zapier automation with an AI step that turns online order forms into job tickets. The same structure works for any business with a few AI jobs running.

Connect on LinkedInSagnik Bhattacharya

The review sits between two other habits. Day-to-day, people check AI output before it goes out; when something slips through, it goes in an error log. The monthly review is where the log, the numbers and a fresh sample meet, so patterns get fixed instead of patched one at a time. If you don't have an error log yet, the tutorial on keeping an AI error log sets one up in an afternoon.

Subscribe on YouTube@codingliquids

What happens before the 30 minutes start

The meeting stays short because the preparation happens beforehand, ideally by a single person spending about 20 minutes on the last working day of the month. Prepare three things:

  1. The sample. Five random items from each AI job, pasted into one shared document. Random matters: if you pick examples yourself, you'll unconsciously pick the ones you remember, which are usually the best or the worst.
  2. The numbers. A few figures per job, the same ones every month, so trends show up: complaints mentioning AI output, chatbot handovers to a person, automation runs that errored, error-log entries.
  3. Vendor changes. Any notice from a tool you use about a price, feature, model or privacy change. These arrive by email and get ignored unless someone is looking for them.

For random picks without special software: export the month's items to a spreadsheet, add a column with =RAND(), sort by it and take the top five. For emails, the print shop exports the "AI draft" folder; for chats, its chat tool's transcript export; for the automation, the task history.

The agenda, minute by minute

MinutesWhat happensOutput
0-5Read the numbers against last month. Anything up sharply gets a question mark.One line per AI job: steady, better or worse
5-20Score the sample: five items per job, about 45 seconds each, using the scoring sheetA score per item and a note on every zero
20-25Pick the fixes: what caused each zero, and which three problems matter mostNo more than three actions
25-30Record: owner, due date and how you'll know the fix worked; glance at vendor changesThe review log updated

Keep the clock visible. The discipline of 45 seconds per item forces you to score what's in front of you rather than debate what the AI might do in other cases.

A scoring sheet you can copy

Four criteria, each scored 0, 1 or 2. Any item scoring 0 on accuracy or policy counts as a failure whatever its total, because a well-written wrong answer is still wrong.

Criterion210
AccuracyEvery fact, price, date and product detail correctA minor slip a customer wouldn't act onA wrong fact someone could act on
ToneSounds like your businessGeneric but acceptableWrong for the situation (cheery about a complaint, stiff with a regular)
CompletenessAnswers the actual question and says what happens nextAnswers it but leaves an obvious follow-upMisses the question or the key detail
PolicyWithin what you offer and promiseVague where it should be specificPromises something you don't offer, or breaks a rule

Filled in, three rows from the print shop's sheet look like this (illustrative):

ItemAcc.ToneComp.PolicyNote
Quote email: 250 A5 flyers2212Didn't mention artwork deadline for Friday delivery
Chat: "can you print today?"0220Promised same-day printing; we need 48 hours
Instagram post: wedding stationery1122Seven hashtags; Instagram caps posts at five

Two scorers should agree within a point on most items. If they don't, the criteria need a sentence of clarification, not a longer discussion. If your AI jobs are mostly support replies and you want a deeper weekly version for those, the weekly sampling routine for AI support replies goes into more detail.

One month at the print shop, scored

Here is the illustrative September review in full: 20 items sampled across four jobs, plus the numbers.

AI jobItemsFailures (a zero on accuracy or policy)Average score (out of 8)Numbers this month
Quote emails507.21 complaint (wrong paper weight quoted)
Website chat525.438% of chats handed to staff, up from 29%
Posts and product copy506.62 posts edited after publishing
Order-form automation5 runs1—23 errored runs, up from 2

Three findings came out of 15 minutes of scoring:

  • The chat assistant was promising same-day printing. Two of five sampled chats said "we can usually print the same day". The source was an old FAQ page from a promotion, still in the assistant's knowledge base. The rise in handovers had the same cause: customers who were told "same day" then asked staff to confirm.
  • The automation's errors had jumped from 2 to 23. Someone had renamed a field on the order form from "Paper stock" to "Paper type", so the AI step received nothing and guessed. One of the five sampled runs had created a job ticket for the wrong stock.
  • Posts were fine but careless with platform rules. Seven hashtags on one post, when Instagram has capped posts at five since December 2025.

The three actions agreed, with owners: remove the old FAQ from the chat assistant's sources and add an explicit "orders take 48 hours" rule (the owner, by Friday); map the renamed form field and replay the 23 failed runs (the office manager, by Wednesday); add "maximum five hashtags" to the social post prompt (whoever writes posts, today). Each had a check for next month: no "same day" in sampled chats, errored runs back under five, no post over five hashtags.

The quick sum on whether 30 minutes a month is worth it: that's six hours a year. The single wrong-stock job ticket cost the shop a reprint of 500 wedding invitations, roughly $180 in materials and three hours of press and finishing time. One catch like that a year pays for the review several times over.

Five causes behind a zero, and the fix each one needs

The five minutes set aside for picking fixes go faster if you know the usual causes. Almost every failure the print shop has found traces back to one of these, and the cause decides the fix. Rewriting the prompt is the reflex, but it's the right answer less than half the time.

CauseHow it shows up in the sampleThe fix
Stale source materialThe AI repeats an old price, policy or offer word for wordUpdate or remove the source document; the prompt is fine
Missing ruleOutput is reasonable but breaks something nobody wrote down (hashtag limits, lead times)Add one line to the prompt or saved instructions
Changed inputAn automation's output goes wrong from a specific dateFind what changed upstream (a form field, a column name) and map it again
Tool or model changeOutput gets longer, more formal or differently structured across the boardRe-run your standard test inputs and adjust instructions
Skipped human checkAn error that the day-to-day check should have caughtAsk why the check was skipped; usually time pressure or unclear ownership

In the September review above, the same-day promise was stale source material, the automation errors were a changed input and the hashtags were a missing rule. Three different fixes, none of which would have come from "improve the prompt".

Automations fail quietly: the checks that catch them

AI steps inside automations are the easiest to forget, because nobody reads their output directly. The monthly review should always look at run history, not just a sample of results, because the platforms behave in ways that hide problems:

  • Zapier automatically pauses a Zap when 95% of its runs error over seven days, which stops the damage but also stops the work. And if you've added an error handler, Zapier doesn't send the usual error emails when the handler runs, so a failing step can go unnoticed.
  • Power Automate switches flows off after 14 days of continuous failure, and flows that haven't triggered for 90 days are switched off too unless you have a Premium licence.
  • Make lets you attach error handlers (now named Skip, Retry, Resume, Commit and Rollback). Skip in particular quietly drops the item that failed, which is exactly what you need to see in a monthly review.

So the prep for any automation is: count errored runs this month versus last, open two or three failures to see why, and sample five successful runs to check the AI step's output made sense. A broader look at automations nobody owns is in the automation audit for Zapier and Make.

Chatbot numbers that look good and mean little

If a chat assistant is one of your AI jobs, be careful which number you track. Vendors report "resolutions", and their definitions are generous. Intercom counts an "assumed resolution" for Fin when the customer goes quiet for 24 hours after Fin's last answer. Zendesk's default closes a messaging conversation after two hours without activity, or 72 hours for email and web forms, then an AI check decides whether it counts as a verified, billable resolution (since 18 May 2026 hand-offs and unverified closes are free). A customer who gave up and phoned you instead counts as resolved in both.

For the review, track handovers to a person and read the sample. A rising handover rate, like the print shop's jump from 29% to 38%, often points to one specific wrong answer rather than a general decline. If you run a Shopify store with the free AI agent in Shopify Inbox, remember it can use web search as a secondary source, so include two or three of your own policy questions in the monthly sample and check its answers match your policies. The tutorial on measuring whether your AI chatbot is working covers the fuller set of metrics.

Vendor changes to scan for each month

The last five minutes include a quick look at what changed in your tools, because a vendor change can break a working setup overnight. Examples from the last year show how varied these are:

  • Retirements. Any custom GPT your team relies on will stop working on 11 December 2026, when OpenAI retires them. Microsoft retired Excel's =COPILOT() worksheet function on 14 September 2026, leaving cached results in place but allowing no new formulas. Anything built on either needs a replacement plan.
  • Renamed products. ChatGPT agent became ChatGPT Work; NotebookLM became Gemini Notebook. Renames rarely break things but confuse staff instructions.
  • Privacy defaults. Vendors sometimes change what they keep by default. When one does, your data handling rules may need updating.
  • Price and credit changes. A credit-based AI feature that gets more expensive per use shows up first as a larger bill.

If you use Microsoft 365 Copilot, the admin centre's Copilot usage report (active users over 7, 28, 90 or 180 days) is worth a glance at the same time: licences nobody uses are a cost issue, and a sudden drop in use sometimes means people have lost trust in the output.

Adapting the review for other businesses

The agenda stays the same; the AI jobs and the numbers change.

  • An online clothing shop samples chat answers about returns and sizing, AI product descriptions and the abandoned-cart emails. Its key number is returns with the reason "not as described", which rises when product copy drifts.
  • A software reseller samples AI-drafted renewal notices and the monthly licence summaries an AI step writes from its distributor reports. Accuracy dominates the scoring: one wrong licence count is a billing dispute.
  • An events company samples AI-drafted proposals and supplier emails. Its failures are usually completeness: dates, guest numbers or dietary notes missing from a proposal.
  • A handmade jewellery seller working alone does a 15-minute version: five product descriptions, five customer replies, one number (messages asking something the listing should have answered).

Turning findings into fixes without a backlog

The rule that keeps this to 30 minutes is the limit of three actions a month. Everything else found goes into the error log to be looked at next time. Three fixes that actually happen beat twelve that sit on a list.

Each action needs three things written down: who, by when, and how next month's review will tell whether it worked. The print shop's review log is a single spreadsheet tab with one row per action. Two rows from October, the month after the review above (illustrative):

MonthAI jobProblemFixOwner, dueCheckResult next month
SepWebsite chatPromised same-day printingOld FAQ removed; 48-hour rule addedOwner, Fri 4thNo "same day" in sampled chats0 of 5; handovers down to 30%
SepOrder automationRenamed form field broke AI stepField remapped; 23 runs replayedOffice manager, Wed 2ndErrored runs under 51 errored run

The last column is what makes the log useful: it shows whether fixes worked, not just that they were made. After a few months, that log is also the best evidence you have of what AI is doing well and badly in your business, which is useful when deciding whether to keep, change or cancel a tool. The tutorial on reviewing an AI tool after 90 days uses exactly that kind of record, and the one on checking that AI is doing good work, not just fast work covers the thinking behind the scoring.

Running the review: common questions

Who should attend a monthly AI quality review?

Two people is ideal in a small business: whoever owns the AI tools and one person who deals with customers directly. The customer-facing person spots wrong tone and missing information quickly. More than three people turns a 30-minute review into an hour of discussion.

Is five examples per AI job enough to find problems?

Five random examples won't catch rare errors, but they reliably catch problems that affect a meaningful share of output, and those are the ones worth fixing first. Rare but serious errors are what the error log and customer complaints are for. If one job keeps producing low scores, sample ten of it the following month.

What if we find nothing wrong for several months?

Keep the review, but shorten it or move one AI job to every other month. A clean run usually means the prompts and guardrails are working, and the review is what will tell you when a tool update or a new product breaks them.

Should the review include how much time AI is saving?

Once a quarter, yes: time a few tasks with and without AI, or compare output volumes. Monthly, stay focused on quality, because a tool that saves time while producing wrong answers is costing you more than it looks.

Further reads

Sources: series fact sheet (checked September 2026) for Zapier auto-pause and error-handler behaviour, Power Automate flow shut-off rules, Intercom Fin and Zendesk resolution definitions, the Microsoft 365 Copilot usage report, OpenAI's custom GPT retirement date, Microsoft's retirement of Excel's COPILOT function, and the Instagram hashtag cap.

Want a monthly AI review set up for your business?

On a 1:1 call we'll list where AI touches your customers and records, build the sample and scoring sheet around those jobs, and agree who runs the review and what happens to its findings.

Book a 1:1 call with me