To pilot AI in shadow mode, run it on real work in parallel while staff carry on as normal, sending its answers to a private log that customers never see. Compare each AI answer with what staff actually did for two to four weeks, or at least 100 cases, and go live only when it meets a threshold you set beforehand.
Shadow mode is the most cautious way to test AI on customer-facing work, because nothing it produces can go wrong in public. It suits jobs where a mistake would be embarrassing or costly: replies to customers, booking changes, price questions. The price is double handling: staff still do every job themselves, so the AI saves no time until it graduates, which is why the test needs an end date and a pass mark agreed before it starts.
A note on names, because they're easily confused: shadow mode is a deliberate, controlled test. Shadow AI is staff using AI tools without the business knowing, which is a different problem covered in shadow AI in a small business.
How shadow mode differs from an ordinary pilot
Most AI pilots start in assisted mode: the AI drafts, staff edit and send. That's fine for low-risk work, but it has a blind spot. Once staff see the AI's draft, they tend to accept its framing, so you learn how much they edit rather than whether the AI would have been right on its own. In shadow mode, staff don't see the AI's output while they work, so the comparison is clean.
| Stage | Who acts on the AI's output | Risk to customers | What you learn |
|---|---|---|---|
| Shadow | Nobody; it's logged only | None | How often the AI would have been right, on real cases |
| Assisted | Staff edit and send | Low | Time saved and how much editing is needed |
| Supervised automatic | AI sends for low-risk categories; staff sample | Medium | Behaviour at full volume |
| Fully automatic | AI acts alone | Highest | Only suitable for well-proven, low-stakes jobs |
Shadow mode is a stage before assisted mode, not a replacement for it. For the general pilot structure (charter, baseline, decision meeting), see running your first AI pilot project; the same charter and decision meeting apply, with a private log standing in for live use.
Three ways to run AI in the shadow
1. Replay last month's cases (one day)
Export 50 to 100 past messages, run each through the AI with your draft instructions, and compare with what staff actually replied. It's the fastest way to catch obvious problems, like out-of-date prices, before investing in anything live. Its weakness: the AI doesn't have the context staff had at the time, such as what the diary looked like that afternoon.
2. End-of-day batch (10 to 15 minutes a day)
Each evening, someone pastes the day's messages into a business AI account with the saved instructions and copies the AI's answers into the log. No automation needed, and it runs on the same current information staff used. It's a good fit for 10 to 30 messages a day.
3. Live copy (an hour or two to set up)
An automation tool copies each incoming message to the AI as it arrives and writes the AI's answer into a spreadsheet row, and staff never see that sheet. In Zapier, for example, that's a new-email trigger, an AI step and a spreadsheet step. As of September 2026, Zapier offers its AI by Zapier step on the Professional, Team and Enterprise plans, with at most limited preview access on the free plan. If you connect to an AI provider directly, check its data terms; OpenAI, for example, says data sent to its API isn't used to train its models unless you opt in.
A sensible sequence is replay first to fix the obvious, then live copy or end-of-day batch for two to three weeks.
The comparison log
One row per case, filled in within 48 hours while the context is fresh. Keep personal details out of it: a short summary of the message is enough.
| Column | What goes in it |
|---|---|
| Case ID and date | A running number and the date received |
| Message type | Price, availability, cancellation, complaint, other |
| What staff did | A one-line summary of the actual reply or action |
| What the AI proposed | Its full draft, pasted in |
| Grade | A, B, C or D (see below) |
| Error type | Wrong fact, invented fact, missed question, wrong tone, should have passed to a person |
| Grader and notes | Who graded it, and anything odd |
Filled in, three rows from an illustrative self-storage site's log might look like this:
| Case | Type | What staff did | What the AI proposed (summary) | Grade | Error type |
|---|---|---|---|---|---|
| 041, 3 Mar | Price | Quoted the 50 sq ft monthly rate and the insurance add-on | Same rate; left out insurance, which is compulsory | C | Missed question |
| 042, 3 Mar | Availability | Said no 100 sq ft units until the 20th; offered two 75s | "We have plenty of 100 sq ft units available" | D | Invented fact |
| 043, 4 Mar | Cancellation | Explained the 14-day notice period and move-out checklist | Same, in slightly warmer wording | A | None |
Case 042 is the classic shadow-mode catch: the AI had no view of the unit list, so it filled the gap with something that sounded helpful. Live, that reply would have brought a customer with a van to a site with nothing to rent them. The grade and the error type together tell you the fix: availability questions need either a connection to the unit system or a rule that the AI never states availability.
The grading scale needs to separate harmless misses from harmful ones, because they lead to different decisions:
A Right outcome; would send as written
B Right outcome; minor wording edits needed
C Wrong, but harmless (a person would catch it; no cost if sent)
D Wrong and harmful: wrong price or time, a promise we can't keep,
rude or careless tone, personal data exposed, or it answered
something that should have gone to a person
Grading fairly, without fooling yourself
Staff replies aren't a perfect answer key. Sometimes the AI's version is better, and sometimes the staff reply was the wrong one. Four habits keep grading honest:
- Grade blind for part of the sample. Each week, put 20 staff replies and 20 AI replies for the same messages into one list without labels, and have the grader score them all. If AI answers score about as well as staff ones, that's strong evidence.
- Two graders for the first 30 cases. Compare grades and agree what separates a B from a C before one person grades the rest.
- Grade the outcome, not the style. A reply that's slightly more formal than your usual tone is a B, not a C.
- Budget the time. Grading takes about a minute a case, so 100 cases is under two hours spread across the pilot.
The two-grader session is where the scale gets its real meaning. In the barber shop example below, the owner and a senior barber split on an AI reply that said "we're open until 6pm on weekdays". The owner gave it a B, since the hours were right. The barber gave it a C, because the last walk-in cut is taken at 5.30pm and a customer arriving at 5.45 would be turned away. They settled it with a written rule, "hours answers must include the last-cut time", which turned a matter of opinion into something any grader would score the same way. Expect five or six rules like that from the first 30 cases.
Blind grading also catches the opposite problem: a staff reply that was wrong. In one week's blind list, a message asking about a skin fade with a beard trim scored lower for the staff answer than the AI's, because the staff member had quoted the old combined price from memory. The AI, reading the current price list, had it right. That case counts in the AI's favour, and it's worth passing to the owner too, since it shows the price list isn't reaching everyone.
A worked example: a barber shop's booking-message assistant
Take an illustrative four-chair barber shop that gets about 80 messages a week through email, its website form and its social media inbox: prices, "can I get in today?", beard trims, children's cuts and cancellations. The owner wants AI to draft replies, but a wrong price or a promised slot that doesn't exist would annoy exactly the regulars the shop depends on.
Week 0, replay. Sixty past messages went through the AI. Results: 28 A, 17 B, 9 C and 6 D, so 75% right outcomes and 10% harmful. The D grades had four causes: a beard-trim price taken from last year's price list, an invented Sunday opening, a reply saying a cancellation was "confirmed" when the AI can't cancel anything, and an offer of a 3pm slot with no access to the diary.
Fixes. A current, dated price list; opening hours as a fixed block in the instructions; and a hard rule:
Never confirm, cancel, move or offer a specific appointment time.
Instead write: "We'll check the diary and get straight back to you."
If a message mentions a bad cut, a refund, being unhappy or an
injury, do not draft a reply. Write only: PASS TO OWNER.
Weeks 1 to 3, live copy. 236 messages: 132 A, 71 B, 26 C and 7 D, so 86% right outcomes and 3% harmful. Five of the seven D grades were missed escalations (a customer unhappy with a fade, worded politely enough that the AI replied cheerfully), and two were children's prices applied to the wrong age band. The escalation rule gained more trigger phrases, and the price list gained explicit age bands.
Week 4. 78 messages, 71 graded A or B (91%) and no D grades. Price and opening-hours questions moved to assisted mode, with staff sending every reply. Cancellations and anything about a finished cut stayed human-only.
The decision was made type by type from a scorecard of the 314 live-copy cases (weeks 1 to 4), checked against the exit criteria further down:
| Message type | Cases | A or B, latest 100 | D grades, latest 50 | Decision |
|---|---|---|---|---|
| Prices | 112 | 93% | 0 | Assisted mode |
| Opening hours | 104 | 95% | 0 | Assisted mode |
| "Can I get in today?" | 39 | Too few cases | 0 | Human-only until the diary can be connected |
| Cancellations | 38 | Too few cases | 0 | Human-only by design |
| Unhappy with a cut | 21 | Too few cases | 5 | Always human |
Two of the seven D grades from weeks 1 to 3 were children's prices, yet prices still passed: both fell outside the latest 50 cases, both had a known cause, and the fix (explicit age bands) was re-tested before week 4. That's how the criteria are meant to work. A type isn't failed forever for an early error; it's failed until the fix has been proved on fresh cases.
How long to stay in the shadow
Length depends on volume, not the calendar. You want enough cases of each message type that one lucky or unlucky week can't decide the result. As a working rule, aim for at least 100 cases of any type you plan to take live, and never less than two weeks, so that different days and staff are represented.
Low-volume types are the awkward ones. If the barber shop only gets five cancellation messages a week, 100 cases would take five months. Don't wait that long; instead, keep those types human-only or move them to assisted mode with every reply checked, and let the high-volume types go ahead on their own evidence. Splitting the decision by type also keeps the pilot moving, which matters: shadow runs that drag past six weeks tend to lose the team's attention, and the log quietly stops being filled in.
Put a figure on the double handling before you commit, so the length is a choice rather than a drift. With the end-of-day batch method, pasting and logging takes about 15 minutes a day; over four weeks of five working days that's 5 hours. Grading 300 cases at about a minute each adds another 5 hours. So a four-week shadow run costs roughly 10 staff hours, plus an evening or two fixing instructions after the replay. If the job itself takes the team 6 hours a week, that's under two weeks of the time you're hoping to save, which is a fair price for going live with evidence rather than hope.
Exit criteria for leaving the shadow
Write these before the shadow run starts, and apply them per message type rather than overall, because an AI can be excellent at price questions and poor at complaints:
- At least 90% A or B grades across the most recent 100 cases of that type.
- No D grades in the most recent 50 cases of that type.
- Every D grade seen during the pilot has a known cause and a fix that was re-tested.
- The sample includes at least one busy period, such as a Saturday or a pre-holiday week.
Leaving the shadow means moving to assisted mode, not straight to automatic sending. If you later consider letting the AI act on its own, the extra safeguards are in piloting your first AI agent without risking customers, and a website chat assistant needs its own pre-launch testing, covered in how to test a customer chatbot before it goes live.
Where shadow testing can mislead you
- A quiet sample. Two weeks in a slow month won't show how the AI copes with a flood of Saturday-morning messages or seasonal requests.
- Missing context. In shadow, the AI may lack the diary or order system it would have live, or have information staff didn't. Grade with that in mind and note it in the log. A common version: the shop closes on a Thursday for staff training, and everyone knows because it's on the back-room whiteboard. The AI doesn't, so it cheerfully tells three customers to walk in on Thursday. Grade those as D, because live they'd be D, then add a short "this week's notices" block to the instructions and make updating it part of whoever writes the whiteboard.
- Grader drift. Graders get more lenient as the novelty wears off. Re-grade ten early cases in week 3 and see if the grades still match.
- Privacy is still in play. Copying customer messages into another tool is still processing their personal data, even if no reply is sent. Use a business account, keep only the fields you need in the log, and delete the log when the pilot ends.
Keep a written note of what went wrong in the shadow and how you fixed it. It's the first thing you'll need if, after go-live, the AI gets something wrong with a customer, and it's also the quickest way to brief the next person who has to maintain the instructions.
Further reads
- Did Your AI Pilot Work? How to Set Success Criteria That Hold Up — Set exit thresholds that will stand up at the decision meeting.
- How to Spot-Check AI Support Replies: A Weekly Sampling Routine — The sampling routine to keep after the AI goes live.
- How to Check AI Is Doing Good Work, Not Just Fast Work — Build the quality rubric your graders will use.
- How Barbers Can Win Back Lapsed Clients With Automated Messages — A barber-specific automation once booking replies are settled.
- AI Hallucinations Explained for Business Owners: Causes and Fixes — Why the AI invented opening hours, and how to prevent it.
- A Simple AI Risk Register for Small Businesses (With Template) — Record the D-grade error types as risks with owners.
- How to Brief a Developer on a Custom AI Workflow You Need — The eight parts of a developer brief for an AI workflow, a copyable requirements template, and a worked example from a dog-grooming salon.
- How to Set Up Human Review for AI Work Without Slowing Down — Four levels of human review matched to risk, how to make each check take under a minute, how many to sample, and when to relax or tighten.
- How to Tell If a Process Is Ready to Automate With AI — An eight-point readiness scorecard with two automatic vetoes, and a driving school's cancellation process scored 7, fixed, then rescored 13.
- Example AI Roadmap for a 12-Person Business, Month by Month — A 12-person driving school's first AI year, month by month: five workflows, about $30 a month in new tools, one idea postponed and why.
- How to Measure Customer Reaction After Introducing AI — Four signals, survey wording that doesn't lead, a conversation-sorting prompt and a decision rule for reading small-business numbers honestly.
- How to Talk to Staff Who Fear AI Will Take Their Job — Prepare your honest answer, then use a six-part conversation outline, better phrasing and a physiotherapy clinic example to talk it through.
- AI Use Case Template: Score Every Idea on One Page — A one-page template with scoring anchors, knock-out questions and a worked veterinary example for ranking AI ideas before you spend anything.
- AI Governance for a Small Business: Who Decides, Approves, Checks — Who says yes to AI in a small firm, and who looks back: a decision-rights table, three approval tiers, a 30-minute monthly check and a one-page register.
- Customer-Facing or Back-Office: Where Should AI Go First? — A side-by-side comparison of customer-facing and back-office AI as a first move, with a scoring sheet, the middle route, and two worked decisions.
- Why Your AI Pilot Stalled, and How to Get It Live — A one-hour diagnosis for a stuck AI pilot, the fix for each of five causes, a 30-day restart plan, and when stopping is the better call.
- AI Tools and AI Development: The Complete 2026 Guide — the AI hub, including every tutorial in the AI-for-business series.
Sources: Zapier help articles on AI by Zapier plan availability and model tiers; OpenAI platform data controls. Checked September 2026.