Pilot your first AI agent in stages: run it in shadow mode first, drafting answers staff compare but customers never see; then let it act only after a person approves; then go live for one narrow job in a limited window, with a tested kill switch, every conversation logged and stop rules written before launch. Expect six to eight weeks.
What makes an agent riskier than a chatbot is that it acts: it books, confirms, quotes or sends on your behalf. So the pilot isn't about whether the AI writes well. It's about whether you can see every action it takes, stop it in under a minute, and catch its mistakes before a customer does. Plan those three things first and the rest of the pilot is routine.
Costs during a pilot are modest if you stage it properly. Shadow mode can run in an assistant you already use, and the limited live phase on a per-resolution or per-conversation agent typically costs $100-$150 a month at small-business volumes. The expensive pilots are the ones that skip straight to live and spend the savings apologising.
Write the agent's job description before choosing anything
The worked example is a café, an illustrative one, that has grown a catering side: sandwich platters, cakes and canapés for offices and parties, about 60 enquiries a month by website form and email, handled by the manager between services. The owner wants an AI agent to answer catering enquiries and book tasting appointments, because enquiries that wait a day often go elsewhere.
The first step was a one-page job description, the same as for a new hire:
AGENT: Catering enquiries (pilot)
DOES
- Answers questions about the catering menu, prices per head,
minimum orders and lead times, from the approved files only
- Collects the details for a quote: date, numbers, budget,
dietary needs, delivery or collection
- Offers tasting appointments from the manager's calendar
(weekdays 2pm-4pm only)
ALWAYS HANDS OVER
- Allergen and dietary safety questions (beyond "noted for the chef")
- Orders over $600, discounts, complaints, refunds
- Any date inside 5 working days (kitchen capacity check)
NEVER
- Confirms an order or a price in writing as final
- Takes payment details
- Says anything about ingredients that isn't on the menu sheet
OWNER OF THE AGENT: cafe manager KILL SWITCH: manager or owner
The "never" list is the one that protects customers. Every item on it is something an agent could plausibly do if asked nicely, and each would be expensive or dangerous to get wrong. If you're unsure whether your job needs an agent at all, rather than a chatbot or a simple automation, settle that first; the risks specific to agents that take actions are laid out in what can go wrong when AI agents take actions for you.
Stage one: shadow mode for two weeks
In shadow mode the agent sees real enquiries but its output goes only to staff. Customers get the manager's normal replies. Shadow mode doesn't need the paid agent: the café loaded the menu sheet, price list, lead times and the job description into a project in the AI assistant it already used, and the manager pasted each real enquiry in, saved the draft, then wrote her own reply as usual. The general method is covered in piloting AI in shadow mode; for an agent, add one column: what action it would have taken.
Two weeks produced 28 enquiries and a clear picture (illustrative):
| Result | Count | Examples |
|---|---|---|
| Draft as good as the manager's | 17 | Menu questions, lead times, collecting quote details |
| Right but needed editing | 7 | Too long; missed the delivery charge; formal tone |
| Wrong | 4 | Quoted canapés for 30 people (minimum is 40); offered a Saturday tasting; said the vegan brownies were nut-free |
| Actions it would have taken | 9 tasting bookings | Two in slots the manager had blocked for a supplier visit |
The four wrong answers mattered more than the seventeen good ones. The minimum-order error came from a price list that didn't state the minimum; the Saturday tasting from a calendar that showed Saturdays as free because nobody worked them; the nut claim from a product description that said "no nuts added" without the kitchen's "may contain" warning. All three were fixed in the files and the calendar, not in the AI. The nut answer also moved "nut-free" and "allergy" into the always-hand-over list.
Stage two: act only after a person approves
For the next two weeks, the agent ran in the real chat tool on the website, but every reply and every booking waited for a person to approve it. Customers saw answers a few minutes slower than a bot's but faster than before, because the manager approved from her phone between orders. Approval stages exist in several tools, with limits worth knowing: ChatGPT Work, OpenAI's agent mode, asks for approval before important actions by default; Claude in Chrome offers an "Automatically approve" mode (the default) and a "Manually approve" mode, and Anthropic advises keeping it away from financial accounts, contracts and other people's personal information; on Zapier's Professional plan, human-in-the-loop approvals can only be sent to yourself; and Make's Human in the Loop app is Enterprise-only. If your tool has no approval step, stage two becomes "the agent drafts, a person sends", which works just as well for a pilot.
The manager approved 31 replies and rejected 3 in the fortnight. One rejection shows the value of the stage. A customer asked for "a tasting this Friday for our office party on the 12th". The agent offered Friday at 2pm, which was free, but the party was six working days away and the kitchen needs five to prepare a tasting menu. The rule "no date inside five working days" had been written for orders, not tastings. The manager added tastings to it before stage three.
Stage three: live in a narrow window, with a kill switch you've tested
Only now did the agent act on its own, and only within limits:
- When: evenings after 6pm and weekends, when enquiries previously waited longest. Daytime enquiries still went to the manager.
- What: the three jobs in its description. Everything else went to the manager's inbox with the chat attached.
- How much: a cap of 15 conversations a day; above that, the chat offered a callback instead. A cap protects you from a viral post or a misuse attempt as much as from cost.
- Disclosure: the first message said it was an automated assistant and how to reach a person.
The kill switch was defined before launch and rehearsed. In the café's chat tool it was the setting that pauses the AI agent so all chats route to the team inbox; the owner and manager both knew where it was, and a fallback message was ready: "Our catering team will reply to you by 10am tomorrow." They timed the drill: 40 seconds from phone in pocket to agent off. A kill switch nobody has practised is a hope, not a control. Find the equivalent setting in your tool, write down where it is, and test it with the whole team present.
Most of what the log showed in the first week was the agent doing its job, including knowing when to stop. A typical evening exchange (illustrative):
Customer: "Hi, we need lunch for 45 on the 20th, one coeliac and two vegans, budget about $15 a head. Can you deliver by 12?"
Agent: "Thanks, that sounds doable. For 45 people our sandwich and salad platters come to about $13-$16 a head, and we deliver from 11:30. I've noted the vegan and coeliac guests for the chef; our catering manager will confirm what we can safely prepare for coeliac diets, as that needs a kitchen check. She'll send your quote by 10am tomorrow. Is this the best email for it?"
Everything in that reply came from the approved files, the dietary question went to a person exactly as the rules said, and the customer knew when to expect an answer. That handover, not the menu information, is the part to check hardest in the logs.
In the second week live, the log showed the agent had offered a tasting at 3pm on a Monday that turned out to be a public-holiday closure. The calendar hadn't blocked it. The manager saw it in the next morning's log review, phoned the customer to rebook, and blocked closure days in the calendar for the rest of the year. The customer was fine. Without the daily review, they'd have arrived at a locked door.
Logs: what to keep and how to read them
Every tool worth piloting keeps a conversation log. What turns a log into safety is the review habit:
- Daily in the first week live, read every conversation the agent closed and every action it took.
- Then three times a week, sample at least 20 conversations, including every handover and every booking.
- Tag each failure by cause: missing or wrong knowledge, a rule that didn't cover the case, a tool or calendar error, or a customer deliberately trying to misuse it. Each cause has a different fix.
- Record fixes with the date, so you can see whether the same failure returns.
Customers probing the agent is a real category: one visitor tried to get a "staff discount" by claiming to be the owner's friend. The agent handed over, as its rules said. Hardening a public agent against that kind of attempt is covered in protecting a customer-facing chatbot from misuse.
Who does what, and how many hours the pilot takes
Pilots fail quietly when nobody owns them. The café named one owner for the agent (the manager) and budgeted the time honestly:
| Stage | Who | Time |
|---|---|---|
| Week zero: job description, rules, criteria | Owner and manager | About 3 hours |
| Shadow mode, two weeks | Manager | About 20 minutes a working day, roughly 3.5 hours |
| Approval stage, two weeks | Manager, owner as backup | About 15 minutes a day, roughly 3.5 hours |
| Live, first week | Manager | 15 minutes of log review every morning, under 2 hours |
| Live, weeks two to four | Manager | Three 30-minute reviews a week, 4.5 hours |
| Total over eight weeks | About 16 hours |
At an illustrative $25 an hour, that's roughly $400 of the manager's time, more than the agent's subscription for the whole pilot. It's worth saying out loud, because the most common way a pilot goes wrong is that the reviews quietly stop in week five. The café blocked the review slots in the manager's diary like any other shift.
Staff needed a short briefing too, since they'd receive the handovers. The manager's version fitted on a card by the till: what the agent can and can't do, that every handed-over chat arrives in the team inbox with the conversation attached, that a handover should get a reply within two working hours, and that anything the agent said which looked wrong goes to her with a screenshot rather than being corrected quietly with the customer. Errors reported that way are the ones that get fixed in the files, so they don't recur.
What the live stage costs on common agents
For about 120 agent conversations a month in the limited live window, with roughly half resolved without a person, list prices in USD as of September 2026 look like this:
| Agent | How it bills | Pilot month estimate |
|---|---|---|
| Intercom Fin | $0.99 per outcome, plus a seat ($39 billed monthly on Essential) | 60 outcomes: $59.40 + $39 = about $98 |
| Tidio Lyro | Blocks of conversations it replies in: 150 for $110 a month (monthly billing), plus a plan from $29 | About $139, whatever the resolution rate |
| HubSpot Customer Agent | 50 credits (about $0.50) per resolved conversation; Service Hub Professional includes about 3,000 credits a month, which don't roll over | 60 resolutions: 3,000 credits, the whole monthly pool; only sensible if you're already on the plan |
| Jobber Receptionist (for trades businesses on Jobber) | $29 for 30 conversations, then $0.79 each | 120 conversations: $29 + $71.10 = about $100 |
Two cost controls belong in the pilot plan. Put a volume cap on the agent, as the café did, and check how your tool behaves when an allowance runs out: some pause, some bill overage. And set the spending alert or limit before launch, not after the first invoice. The broader picture of what agents cost once they're fully live is in how much an AI agent costs a small business; HubSpot users should also read which HubSpot AI agents small teams should switch on.
Deciding at week eight: expand, fix or stop
The café wrote its criteria in week zero, so the decision at the end wasn't a matter of mood:
| Criterion (set before launch) | Target | Result (illustrative) |
|---|---|---|
| In-scope chats resolved without correction | 60% or more | 64% |
| Wrong prices or dietary claims sent to customers | None | None |
| Actions needing a manual fix | Under 5% | 2 of 48 bookings (4%) |
| Handover replies within 2 working hours | 90% | 93% |
| Complaints about the agent | No more than 1 | 0 |
Stop rules were written alongside: switch the agent off immediately if it states anything about allergens that isn't from the approved sheet, promises a price or date its rules don't allow, or collects payment details. Pause and fix if wrong answers exceed 5% of the weekly sample or handovers wait more than a day. None triggered, so the café extended the agent to daytime enquiries while keeping the same caps, and gave it one new job, following up quotes that customers hadn't answered after three days. That new job starts again at stage one. Setting criteria that survive contact with real results is covered in how to set AI pilot success criteria that hold up.
The pattern is worth keeping for every agent after this one: one job, written limits, shadow before approval before live, a switch you've practised, logs you actually read and criteria you set in advance. It's slower than switching an agent on and hoping, and it's the reason the café's customers never noticed the pilot happening.
Further reads
- What Is an AI Agent? A Plain-English Guide for Business Owners — A plain-English grounding in what makes something an agent.
- AI Agent vs Chatbot vs Automation: Which Does Your Business Need? — Check you need an agent rather than a simpler tool.
- How to Run Your First AI Pilot Project in a Small Business — General pilot planning for AI that isn't customer-facing.
- Why Your AI Pilot Stalled, and How to Get It Live — What to do if the pilot never makes it to live.
- How to Monitor AI That Talks to Customers: Hand-Offs and Errors — Ongoing monitoring once the agent is fully live.
- How to Prepare Your Small Business for AI Agents — Get processes and data ready before any agent arrives.
- Can AI Fill In Forms and Supplier Portals for You? — Browser agents, recorded workflows, integrations or AI-prepared data: how to choose for each form and portal, with a supervised-run prompt and a ten-form test.
- What Are AI Marketing Agents and Should a Small Business Use One? — What separates a marketing agent from a chatbot or automation, the agents small businesses can switch on now, what they cost and the guardrails to set.
- How to Test a Customer Chatbot Before It Goes Live — A filled-in 50-question test script for a customer chatbot, with pass, soft-fail and hard-fail rules and the launch threshold to aim for.
- What Is Copilot Studio and Does a Small Business Need It? — What Microsoft's agent-building tool actually does, how its credit billing adds up, and the four situations where a small firm genuinely needs it.
- AI Agent Use Cases for Small Businesses: 10 Worked Examples — Ten AI agent jobs in six small businesses, each with the before and after, the tool, the monthly cost at list prices and what a person still has to check.
- Can ChatGPT Agent Handle Business Admin Tasks Unsupervised? — ChatGPT agent became ChatGPT Work in July 2026. A task-by-task guide to what it can run alone and what must wait for your approval.
- What Is an AI Browser Agent, and Should Staff Use One? — What an AI browser agent does, which jobs it suits, how prompt injection can hijack it, and the settings and staff rules that keep mistakes small.
- How Much Does an AI Pilot Project Cost? — What a small-business AI pilot really costs, with a catering company's six-week pilot costed, three budget levels and the commitments that inflate it.
- AI Tools and AI Development: The Complete 2026 Guide — the AI hub, including every tutorial in the AI-for-business series.
Sources: Intercom and Tidio pricing pages and pricing FAQs; HubSpot credit, Customer Agent and Jobber Receptionist pricing, ChatGPT Work approvals, Claude in Chrome approval modes and Zapier and Make approval limits from the verified fact sheet. All checked September 2026.