How to Check AI Is Doing Good Work, Not Just Fast Work

Coding Liquids tutorial cover featuring Sagnik Bhattacharya for How to Check AI Is Doing Good Work, Not Just Fast Work.
Coding Liquids tutorial cover featuring Sagnik Bhattacharya for How to Check AI Is Doing Good Work, Not Just Fast Work.

Check AI work as you would a new employee's: write down what "good" means for the task in a short rubric, score a random sample of 10 to 20 outputs every week against it, track how heavily staff edit drafts before use, and count every error that reaches a customer. Speed only counts if those numbers hold.

The difficulty is that AI output looks good. It's fluent, tidy and confident even when it's wrong, so problems slip past a quick glance. Be wary of one reassuring number in particular: when staff stop editing AI drafts, it can mean the drafts got better or that nobody is reading them closely any more, and only a scored random sample tells you which.

Follow me on Instagram@sagnikteaches

None of this needs special software. A spreadsheet, the source documents the AI works from, and half an hour a week of someone who knows the subject are enough for most small businesses.

Connect on LinkedInSagnik Bhattacharya

Why fast AI work hides its mistakes

Four things make AI errors easier to miss than a person's:

Subscribe on YouTube@codingliquids
  • Fluency looks like knowledge. A person who doesn't know something usually sounds unsure. An AI writes a wrong answer in the same assured tone as a right one. AI hallucinations explained for business owners covers why this happens.
  • Errors cluster. Most outputs may be fine while one type of question goes wrong again and again. A general impression of "it's pretty good" hides that pattern completely.
  • Checkers relax. After fifty good drafts, the fifty-first gets skimmed. People come to trust automated output and check it less over time, which is exactly when a problem gets through. Handling staff who over-rely on AI deals with the people side of this.
  • Speed metrics reward volume. If the only number anyone watches is time saved, every incentive points towards approving faster.

Write a quality rubric for each task

A rubric is a short list of checks that define a good output for one specific task. Four to six checks is enough, and at least one should be marked critical, meaning any failure on it fails the whole item. Here's one for an illustrative garden centre whose AI drafts replies to plant-care emails:

RUBRIC: plant-care email replies                 Checker: ________

CRITICAL (any fail = item fails)
[ ] ACCURATE   Every fact matches our care sheets and stock list:
               plant names, light, watering, hardiness, prices, stock
[ ] SAFE       Anything about pets, children, eating plants or chemicals
               was passed to the plant team, not answered by AI
[ ] ON POLICY  Plant guarantee and returns terms stated correctly;
               no promises outside policy

STANDARD (note, but the item can still pass)
[ ] COMPLETE   Every question in the customer's email is answered
[ ] SOURCED    Uses our care sheet where we have one, not generic advice
[ ] OUR VOICE  Friendly, plain, short; sounds like us

Result:  PASS / FAIL        Error type: ______________

The safety line is specific to this business, and that's the point: your rubric should name the mistakes that would actually hurt you. For an optician it might be clinical advice; for a pet shop, feeding amounts; for a dry cleaner, care instructions for delicate fabrics.

A scored item makes the rubric concrete. Say a pet shop's rubric for AI-drafted web descriptions has two critical checks (feeding amounts match the supplier sheet; no health claims the supplier doesn't make) and two standard ones (the pack size and flavour are both stated; it reads like the shop). One sampled draft for a senior dog food said: "Gentle on older joints and proven to support mobility. Feed 250g a day for a 20kg dog. Available in 2kg and 12kg bags." Checked against the supplier sheet, the feeding amount was right, the pack sizes were right, and "proven to support mobility" appeared nowhere: the sheet said only that the food contained glucosamine. That's a critical fail on health claims, even though three of the four checks passed and the text read well. Error type: invented claim.

A weekly sampling routine, about 30 minutes

  1. Pick the sample at random. Take every tenth item from the week's log, or use random numbers. Don't let anyone choose, or you'll get the easy ones. Aim for 10 to 20 per workflow a week, or about 5% if volume is high.
  2. Check against the source, not memory. Open the care sheet, the price list or the policy. Most AI errors are plausible, which is why memory lets them through.
  3. Score each item against the rubric and log the error type for any failure: wrong fact, invented fact, unsafe answer, missed question, off-policy, tone.
  4. Look for repeats. The same error type twice in a month means a fix is needed in the instructions or reference files, not just a correction to that one reply.
  5. Record three numbers: pass rate, critical failures, and the share of items staff sent with no edits.

The random pick in step 1 takes a minute in a spreadsheet. If the week's log has 212 items and you want 15, add a column, fill it with =RAND(), sort the sheet by that column, and take the top 15 rows. Nobody chose them, so nobody can be accused of choosing easy ones, and the Monday that had the most emails will appear in the sample about as often as it should.

An AI assistant can take a first pass at scoring, as long as it gets the same source documents a person would use. A prompt along these lines works:

Score the draft reply below against this rubric. Use ONLY the
attached care sheets and stock list as the source of truth. For
each check, answer PASS or FAIL and quote the sentence that
decides it. If a fact can't be found in the sources, mark it FAIL.

[rubric] [draft] [customer email]

On one garden centre draft, the result came back like this (illustrative):

ACCURATE: PASS ("Water once a week in summer" matches the care sheet). SAFE: PASS (no pets, children or chemicals mentioned). ON POLICY: PASS. COMPLETE: FAIL (the customer also asked whether we sell a larger pot for it; not answered). SOURCED: PASS. OUR VOICE: PASS.

The completeness catch was right and saved the checker a careful re-read. But the draft also called the plant "fully hardy", and the reviewer passed it as accurate, because the care sheet for that plant hadn't been attached, and the instruction to fail anything unsourced was ignored. That's why the critical checks stay with a person: use the AI's scoring to spot missed questions and tone, then verify every fact line yourself against the sheet.

For customer support replies specifically, a weekly sampling routine for AI support replies goes further on sample design. Once a month, roll the weekly figures up; a monthly AI quality review in 30 minutes gives a format for that meeting.

Watch how much staff change, not just whether they approve

Ask whoever checks drafts to mark each one: no edits, minor edits, heavy edits, or discarded. The pattern over time tells you a lot:

PatternWhat it probably meansWhat to do
Heavy edits falling steadilyInstructions and reference files are improvingKeep going; keep sampling
Heavy edits stuck above about a third after a monthThe task or the instructions are wrong for AIRethink the setup or narrow the task
"No edits" jumps to nearly 100% overnightDrafts are being approved without being readSit with the checker for ten minutes; run a seeded test
Discards concentrated in one question typeThe AI shouldn't handle that typeRoute that type to people

A seeded test means occasionally slipping a deliberately flawed draft into the queue, such as one with a wrong watering instruction, to see whether it's caught. Tell staff in advance that this happens from time to time. It's a check on the process, not a trap for individuals, and knowing it might happen keeps everyone reading properly.

What a seeded test reveals is often about timing rather than people. An illustrative dry cleaner seeded two drafts in one week, each saying a silk blouse "can go through our standard wash". The one checked on a quiet Tuesday afternoon was caught at once. The one that landed at 5.40pm on a Saturday, with a queue at the counter, was approved and would have been sent. The fix wasn't a word with the checker. Drafts involving care instructions now wait for the Monday morning batch unless a customer is waiting on them, and the Saturday closing rush only handles collection and opening-hours replies.

Build a test set to re-run whenever anything changes

Sampling tells you how the AI is doing now. A test set tells you whether a change made it worse. Collect 20 to 30 real past cases where you know what a good answer looks like, including the awkward ones:

Case | Input (summary)                         | Good answer must...                   | Critical? | Last result
T01  | "Can my cat be near these lilies?"        | Pass to plant team, no AI answer      | Yes       | Pass
T02  | Watering a fern on a sunny windowsill     | Match care sheet; suggest moving it   | No        | Pass
T03  | Asks for a product we discontinued        | Say not stocked; offer the alternative| Yes       | Fail
T04  | Complaint phrased as a care question      | Pass to manager, no care advice       | Yes       | Pass
T05  | Two questions in one email                | Answer both                           | No        | Pass

Re-run the set whenever you edit the instructions, update reference files, or notice the tool behaving differently. AI providers update their models, and chat apps change which model sits behind them, so output can shift even when you've changed nothing. Any case that passed before and fails now needs investigating before the change goes live.

A re-run earns its keep on changes that look harmless. An illustrative lettings agency shortened its reply instructions because tenants said the emails were too long, adding "keep replies under 100 words". Before switching over, it re-ran its 25-case test set. Twenty-three still passed. Two failed: a question about getting a deposit back now left out the step-by-step of what happens after the check-out inspection, and a repair request no longer included the line telling the tenant what counts as an emergency and which number to ring. Both were exactly the sentences the shorter limit squeezed out. The agency kept the 100-word rule for everything else and exempted deposit and repair replies, then re-ran the set once more: 25 of 25.

Setting the bar: how good is good enough?

Decide your thresholds before you start sampling, so a disappointing week doesn't get explained away. The bar should depend on what a mistake costs, not on what the AI currently manages:

  • Critical checks: zero failures. Anything that could injure someone, cost real money or break a promise has no acceptable error rate. One critical failure means you act that week.
  • Overall pass rate: set it against your own team. If you scored staff-written replies before AI arrived and 85% passed, the AI-assisted process should match or beat that. Without a before figure, 90% is a reasonable starting target for customer-facing drafts that a person checks.
  • Two weeks below target: narrow the task (take the failing question types away from the AI), fix the instructions, and re-run the test set before widening it again.

The two-week rule exists because small samples are jumpy. In a sample of 15, each item is worth nearly 7 percentage points, so 13 of 15 is 87% and 14 of 15 is 93%: one reply is the difference between missing a 90% target and clearing it. A single week at 13 tells you little. Two weeks at 13, or three weeks sliding from 15 to 14 to 13, is a pattern worth acting on. Critical failures are the exception, because one is already too many.

When a sample turns up a critical failure

Treat it as a small incident rather than a correction. First, put it right with the customer. Second, search the rest of that week's output for the same error type, because a sample of 15 that contains one bad answer usually means there are others. Third, fix the cause, add the case to the test set, and note it in the log so the next review can confirm it hasn't come back.

A worked example: the garden centre's plant-care replies

The illustrative garden centre had been using AI to draft replies to about 50 plant-care emails a week, with plant-area staff checking and sending. After six weeks, speed looked excellent: time per reply was down from 7 minutes to 2.5. Then the owner started the weekly sample of 15.

Week 1: 12 of 15 passed. The three failures were a reply saying a plant would do well in partial shade when the care sheet says full sun; a recommendation for a slug control product the centre had stopped stocking; and, most seriously, a reply telling a customer a lily was "generally safe around cats in moderation". True lilies and daylilies are highly toxic to cats, and even small exposures can cause kidney failure. A member of staff had approved that draft. The edit log showed 92% of drafts going out with no edits at all, which explained how.

The first action wasn't a fix to the AI. The owner rang the customer that afternoon to correct the advice, then searched the sent folder for any other replies mentioning pets. There was one more, harmless as it turned out.

Fixes, made the same week:

  • A hard rule in the instructions: any question mentioning pets, children, eating plants or chemicals gets no draft, only "PASS TO PLANT TEAM".
  • The current stock list added as a reference file, updated every Monday.
  • The rubric printed and kept by the desk, and a seeded test once a week.
  • The lily case and the discontinued-product case added to the test set.

Weeks 2 to 6: pass rates of 13, 14, 14, 15 and 14 out of 15, with no critical failures. The "no edits" share fell to around 70%, a sign people were reading again, and time per reply rose to about 3 minutes.

So the real saving was 4 minutes a reply, not 4.5. The extra half-minute is what trustworthy output costs, and it's cheap compared with a customer's cat. That's the general lesson: the first speed figure after an AI rollout is often partly an illusion created by checking that stopped happening.

Signs quality is slipping between checks

Sampling catches most problems, but these day-to-day signals are worth watching too:

  • Customers reply with follow-up questions more often, or ask what a reply meant.
  • Replies grow longer and more generic, with less of your own information in them.
  • Complaints start clustering around one topic.
  • Staff quietly rewrite certain kinds of draft from scratch without mentioning it.
  • The style of output changes noticeably, which often follows a model update.
  • The redo rate rises: more customers writing back because something was missed.

When you see any of these, pull an extra sample that week and re-run the test set. For individual outputs that must be right before they leave the building, such as prices, dates and claims, a five-minute fact-check routine is a useful habit to add. And if the checking itself is slowing the team down, setting up human review without slowing down shows how to focus it on the items that matter.

Checking AI quality: follow-up questions

Who should do the quality checking?

Someone who knows the subject well enough to spot a wrong answer without looking everything up, and who didn't write or approve the items being checked. In a small team that's often the workflow owner checking a colleague's approved drafts. Rotate it every few months so fresh eyes catch what a familiar checker has learnt to skim past.

Can I use one AI to check another AI's work?

Partly. Given your rubric and source documents, an AI reviewer can flag missing answers, tone problems and obvious clashes with policy, which speeds up a first pass. It shares many of the same blind spots, though, and can confidently approve a wrong fact. Use it as a filter, and keep a person checking the critical items such as safety, prices and promises.

How do I check quality if I work alone?

Separate the checking from the doing in time. Once a week, take ten outputs you sent days earlier and score them against your rubric with the source documents open. Distance makes errors easier to see. Keep a small test set of tricky cases and re-run it whenever you change your instructions or notice the tool behaving differently.

Further reads

Sources: veterinary poison-control guidance on lily toxicity in cats (checked September 2026). The garden centre figures are illustrative.

Want a quality check built into your AI workflows?

On a 1:1 call we'll write the rubric for your highest-risk AI task, set up a sampling routine that fits your week, and build the test set you'll re-run whenever something changes.

Book a 1:1 call with me