How to Set Up Human Review for AI Work Without Slowing Down

Coding Liquids tutorial cover featuring Sagnik Bhattacharya for How to Set Up Human Review for AI Work Without Slowing Down.
Coding Liquids tutorial cover featuring Sagnik Bhattacharya for How to Set Up Human Review for AI Work Without Slowing Down.

Match the review to the risk instead of checking everything the same way. Sort each AI output by what a mistake would cost, then give high-stakes work a full check before it goes out, routine work a sample check afterwards, and let automatic rules catch the obvious errors first. Keep each check short, and adjust it using measured error rates.

Done this way, review stops being a queue in front of everything. In the bakery example below, weekly review time falls from about four and a half hours to about two and a half, while the checks on allergen information get stricter, not looser. The seven stages take an afternoon to set up and about ten minutes a week to run.

Follow me on Instagram@sagnikteaches

Stage 1: list every AI output and what a mistake costs (30 minutes)

Write down everything AI produces in your business, however small. For each one, note who sees it, the worst realistic mistake, and whether that mistake can be undone. A social post with a typo can be edited in a minute. An allergen statement a customer has already relied on cannot.

Connect on LinkedInSagnik Bhattacharya

Say a bakery uses AI for five things: replies to custom-cake enquiries (about 60 a week), allergen information for new products and order confirmations (about 5 new items a week), social posts (7 a week), summaries of supplier orders for the owner (3 a week) and replies to online reviews (about 10 a week). The figures are illustrative.

Subscribe on YouTube@codingliquids

Stage 2: give each output one of four review levels (20 minutes)

LevelWhat happensUse it for
A: full check before releaseA named person with the right knowledge checks every item against the sourceSafety, health, money over a set amount, legal wording, anything that can't be undone
B: quick approval before releaseA person reads it and approves, edits or rejects, aiming for under a minuteRoutine customer-facing messages
C: sample after releaseA set number are checked afterwards, looking for patternsInternal or low-stakes work that's easily corrected
D: automatic checks onlyRules catch obvious errors; a person looks occasionallyDrafts only you will read, internal notes

At the bakery: allergen information is Level A, always, checked by the baker against the recipe sheet. Enquiry replies and positive review replies are Level B. Replies to negative reviews move up to Level A, because a clumsy answer to an upset customer is public and hard to undo. Social posts are Level B. Supplier summaries are Level C.

Then add automatic rules in front of every level. These are simple checks a connector or the AI itself can run before anyone looks: the price in a draft must match the price list; any mention of nuts, gluten, dairy, egg, vegan or "allergy" sends an enquiry reply to Level A; drafts containing words from a banned list are held. Rules cost nothing per item and catch the errors that waste reviewers' time. If AI answers allergen questions directly anywhere in your business, read whether an AI chatbot should answer allergen questions first.

The banned list is specific to your trade, and writing it is the owner's job. For an illustrative small skincare brand that uses AI for product descriptions and replies to customer emails, it might hold any draft containing "cure", "treat", "heal", "eczema", "psoriasis", "clinically proven", "dermatologist approved" or "safe in pregnancy", because each is either a medical claim or a promise the brand can't back. A held draft (illustrative) might read: "Our oat balm is gentle enough for sensitive skin and can help treat eczema flare-ups." The rule catches "treat" and "eczema" before any person has to, and sends the draft to Level A, where the owner rewrites it as "Our oat balm is fragrance-free and made for dry, sensitive skin." The rule took five minutes to write and runs on every draft for nothing.

Stage 3: make each check fast (one hour to set up)

Give every output type a checklist of three to five items. Specific ones: "price matches the price list for this size and design", "date offered is one we have free", "no promise about delivery times". Not "check it's correct". A reviewer with a specific list moves faster and misses less.

Put the source next to the draft. The reviewer should see the customer's original message beside the drafted reply, or the recipe sheet beside the allergen text. Hunting for the source is where review time goes.

Review in windows, not constantly. Two or three set times a day (say 10:30, 13:30 and 16:00) beat interruptions every few minutes. Tell customers when they'll hear back, and keep that promise.

Use one-click approval where your tools offer it. The simplest version: AI writes drafts into your email's drafts folder and a person presses send. If you use Zapier, its Human in the Loop tool has a Request Approval action that pauses a workflow until a reviewer approves, declines or edits the data. You set a timeout, choose whether the run skips ahead or ends if nobody responds, and can send reminders. It's available on the Professional, Team and Enterprise plans, reviewers need a Zapier account with the workflow shared to them, and each successful approval step counts as a task. Setting up an approval step for AI-written quotes shows a worked build.

Ask the AI to mark what it isn't sure of. Add to its instructions: "Put CHECK: before any fact that isn't in the source material." It won't catch everything, but it points the reviewer at the likeliest errors.

Put together, a Level B check looks like this in an illustrative plumbing and heating firm. The customer asked for a boiler service and whether someone needs to be home:

CHECKLIST: service quote replies
1. Price matches the price list for this job
2. Slot offered is free in the diary
3. No promise about when parts will arrive
4. The customer's own question is answered

ILLUSTRATIVE DRAFT
Hi, thanks for getting in touch. A standard boiler
service is $95 and takes about an hour. We can come on
Tuesday 14th between 8am and midday. CHECK: If your
pressure valve needs replacing, we can usually fit the
part the same day. Let us know if Tuesday works.

With the price list and diary on screen, items 1 and 2 take ten seconds. The CHECK: marker leads straight to the item 3 problem, a same-day parts promise the firm can't keep, so that sentence goes. Item 4 fails too: the draft never says whether someone needs to be home. The reviewer adds one line and sends it, in under a minute, and ticks it as a material correction in the tally.

Stage 4: decide how many to sample (15 minutes)

For Level C, check at least ten items a week or one in ten, whichever is more. For the bakery's three supplier summaries a week, that means reading all three for the first month, then one a week.

At higher volumes the sum works the other way. An illustrative bookkeeping practice produces about 40 AI-drafted summaries of client calls a week, filed as internal notes. One in ten would be four, so the floor of ten applies: ten a week, drawn from different staff and different days rather than the first ten on Monday. After five clean weeks the practice has 50 checked summaries with no material errors, which is the evidence the rules in Stage 6 ask for before anything loosens.

Be careful what a clean sample proves. A handy rule of thumb from statistics, the "rule of three": if you check a number of items and find no errors, you can be reasonably confident the true error rate is below three divided by that number. Twenty clean checks means errors are probably under about 15%, or roughly one in seven. Fifty clean checks takes that to about 6%, and a hundred to about 3%. That's why a good first week is not evidence that review can stop.

Stage 5: measure the edit rate weekly (ten minutes a week)

For each Level A and B output, reviewers tick one of four boxes: approved unchanged, cosmetic edit (wording, tone), material correction (a wrong fact, price, date or promise), or rejected. The material-correction rate is the number that matters. A spreadsheet with four columns is enough, and checking AI is doing good work, not just fast work covers what else to track.

WEEKLY REVIEW TALLY            Week of: ________
Output type        | Reviewed | Unchanged | Cosmetic | Material | Rejected
Enquiry replies    |          |           |          |          |
Social posts       |          |           |          |          |
Review replies     |          |           |          |          |
Material-correction rate = Material / Reviewed = ____%

The bakery's tally for its third week, illustrative, reads:

WEEKLY REVIEW TALLY            Week of: 3rd week live
Output type        | Reviewed | Unchanged | Cosmetic | Material | Rejected
Enquiry replies    |    58    |    41     |    13    |    3     |    1
Social posts       |     7    |     5     |     2    |    0     |    0
Review replies     |    10    |     6     |     3    |    1     |    0
Material-correction rate = 4 / 75 = 5.3%

That's well above a 2% threshold, so nothing loosens this month. But read the notes behind the number before reacting to it: all three material corrections on enquiry replies were the same mistake, a three-tier cake priced at the two-tier rate, because the price list attached to the prompt was missing a row. One fix to the list, and the next week's rate falls to 1 in 71. A single rate hides whether you have one cause or many; a line of notes against each material correction tells you.

Stage 6: write rules for loosening and tightening

Decide these in advance, so nobody drops a check because it's a busy week.

  • Loosen one level (B to C, for example) only when at least 50 items have been reviewed at the current level and the material-correction rate has stayed under your threshold (say 2%) for four weeks in a row.
  • Tighten immediately after any serious error, and whenever something upstream changes: a new AI model or tool, rewritten instructions, a new price list or menu, a new reviewer. Changes are when errors return.
  • Never loosen Level A for anything involving health, safety or legal wording, however good the record.

The tightening rule is the one that gets skipped. Picture an illustrative dog-walking company that earns its move from Level B to C on review replies: 60 reviewed, one material correction, four clean weeks. Two weeks later it switches the drafting to a newer AI model because the old one is being retired, and nobody thinks of that as a change worth tightening for. The new model's replies to unhappy reviewers start offering "a free walk on us to make it right". The company has no such policy. The weekly sample of ten catches it three weeks later, by which time six public replies carry the offer and two customers have claimed it. Had the model switch sent the replies back to Level B for a fortnight, the first one would have been stopped.

Stage 7: keep reviewers alert

The quiet risk in any review step is that people approve what's usually right without really reading it. After a few hundred good drafts, a reviewer's attention drifts. Three habits help:

  • Plant a known error now and then. Once a fortnight, slip a draft with a deliberate mistake (a wrong price, a date you're closed) into the queue and see whether it's caught. Tell the team this happens; it keeps everyone reading. Address planted drafts to your own test inbox so nothing wrong can reach a customer. When the plumbing firm above planted a boiler service at $59 instead of $95, the 10:30 reviewer caught it in the morning queue, but the same plant a fortnight later went through at 16:00, when the reviewer was also answering the phone. The fix wasn't a lecture: the afternoon window moved to 15:00, before the engineers ring in, and the next two plants were caught.
  • Rotate reviewers where you can, so fresh eyes see the work.
  • Keep checklists short. A ten-item checklist gets skimmed. Five items get read.

For a quick fact-check sequence reviewers can run on any draft, a five-minute fact-check routine works alongside the checklist.

Who reviews, and what happens when they can't

A review step is only as quick as the person behind it. Three rules keep it moving.

The reviewer must know the source. Allergen text goes to whoever owns the recipes, quotes to whoever sets prices. A reviewer who has to ask someone else before approving has doubled the wait.

The reviewer must be allowed to say no. If a junior member of staff can approve but not reject, they'll approve. Give every reviewer the authority to reject a draft and write the reply by hand.

There's always a backup. If the owner is the only approver, every holiday and busy Saturday stalls the queue. Name a second person for each level and write down who covers when.

Then decide what happens when a review window is missed. A good fallback is a fixed holding message, written and approved once, not generated by AI each time: "Thanks for your enquiry. We'll reply with a quote and available dates by midday tomorrow." It buys time without letting an unchecked draft through. Set a limit too: if anything has waited longer than one working day, it goes to the backup reviewer automatically.

Plan for the day both reviewers are away, because it will come. Say an illustrative two-person picture-framing workshop is at a trade fair on Friday and Saturday. Level B enquiry replies can wait behind the holding message until Monday morning. Level A items, which for a framer means quotes for conservation framing of valuable pieces, don't get approved by the Saturday temp just because they're in the queue; the holding message for those says "we'll call you on Monday", and the temp's job is to make sure nobody is left without it.

The bakery's numbers, before and after

OutputBefore: owner reads everythingAfter: review by risk
Enquiry replies (60 a week)3 minutes each: 180 minRules first; 45-second approval for 60 (45 min) plus a 3-minute full check for about 12 flagged for allergens (36 min)
Allergen information (5 a week)5 minutes each: 25 minLevel A by the baker against recipe sheets, 8 minutes each: 40 min
Social posts (7 a week)3 minutes each: 21 min1-minute approval: 7 min
Supplier summaries (3 a week)5 minutes each: 15 minOne sampled a week: 5 min
Review replies (10 a week)3 minutes each: 30 min8 positive at 1 minute, 2 negative at 4 minutes: 16 min
Total271 minutes (about 4.5 hours)149 minutes (about 2.5 hours)

The time saved comes from routine work. The allergen check takes longer than before, and it's done by the person who knows the recipes rather than whoever happens to be free. That's the trade you want: less effort where mistakes are cheap, more where they aren't. Review the levels after a month using the tally, and expect to move one or two outputs up or down.

For support replies at higher volumes, a weekly sampling routine for AI support replies applies the same thinking to a helpdesk.

Further reads

Sources: Zapier help documentation on Human in the Loop and task counting (checked September 2026).

Want review steps that don't become a bottleneck?

On a 1:1 call we'll list what your AI produces, set a review level for each, and work out where an approval step fits in the tools you already use.

Book a 1:1 call with me