The question changes. A proof of concept asks whether AI can do the task at all, on 20 to 100 real samples in a week or two. A pilot asks whether it works inside your workflow, with a few staff for four to eight weeks. Production asks whether it keeps working, adding an owner, monitoring, access controls and an off-switch.
Many small businesses can skip the proof of concept. If you're switching on a feature the vendor already sells, such as drafting in Gmail or Copilot Chat, you know it works in general, so go straight to a pilot. A proof of concept earns its place when nobody knows whether AI can handle your particular inputs: messy scans, unusual formats or your trade's own jargon.
The three stages side by side
| What changes | Proof of concept | Pilot | Production |
|---|---|---|---|
| The question | Can AI do this at all? | Does it work in our real workflow? | Will it keep working without attention? |
| Data | 20 to 100 real samples, copied out | All real cases of one type, as they arrive | Everything in scope, plus the odd cases |
| Who uses it | The person testing it | Two or three staff | Everyone who does the job |
| How long | Days to two weeks | Four to eight weeks | Open-ended, reviewed quarterly |
| Accounts | A business AI account; no connections | Business accounts; limited connections | Business accounts, named owner, least access needed |
| Checking | Every output compared with the right answer | A person checks everything before use | Checks on risky cases and a regular sample |
| When it's wrong | Note it; adjust the instructions | The checker catches it; log it | Alerts, an exceptions route and a fallback |
| Documentation | Notes on what worked | Instructions and a results log | A runbook anyone can follow |
| How it ends | A yes, a no or "only for these cases" | Pass, extend once, redesign or stop | A planned review, replacement or switch-off |
Read down each column and the shift is from "can it?" through "does it?" to "will it, reliably, when nobody's watching?" Each stage adds cost, but mostly in staff time and discipline rather than software.
Proof of concept: proving the hard part on real samples
Consider an illustrative pet shop with two branches and an online shop, stocking around 2,400 products from 11 suppliers. Suppliers send price updates in their own formats, some as spreadsheets and some as PDFs, about 12 a month with 150 to 600 lines each. A staff member spends around 6 hours a month finding the changes and typing them into the stock system, and a missed increase means selling at the old price, sometimes below cost.
The uncertain part was whether AI could read three suppliers' messy PDFs accurately enough, so that's all the proof of concept tested. The owner took 30 lines from each of three recent price lists, 90 in all, and ran them through a business AI workspace with this prompt:
From the attached supplier price list, extract every product line
as a table with these columns: supplier code, product description,
pack size, unit cost, case cost. Copy codes exactly as printed. If
a value is unclear or missing, write UNCLEAR. Don't calculate or
infer any figure that isn't printed.
An illustrative slice of the output:
Code Description Pack Unit cost Case cost
DF-2210 Adult dog food, chicken 12 x 400g UNCLEAR 14.40
CT-0871 Cat litter, clumping 10kg 6.95 UNCLEAR
DF-2214 Puppy food, lamb 12 x 400g 1.35 16.20
What the owner checked and fixed: 86 of the 90 lines were right. The four errors were all pack-size confusions, where a case price was reported as a unit price or the reverse. The instructions were changed to "report unit cost and case cost separately; if only one is printed, mark the other UNCLEAR", and a fresh 30 lines came back with no errors. The UNCLEAR markers are a feature: they show where a person must look rather than letting the AI guess.
Two weeks, about 6 hours and one month of a business plan. The proof of concept's result wasn't "AI works"; it was "AI can extract these three suppliers' formats reliably when told to separate unit and case prices". That narrower answer is the useful one. For more on reading documents like these, AI document processing covers the common formats and their traps.
A proof of concept that says no is a success
The stage is worth running precisely because it can fail cheaply. An illustrative picture framer wanted AI to estimate frame sizes and mount widths from the photos customers email of their artwork, so that rough quotes could go out without a visit to the shop. The proof of concept used 25 past enquiries where the true sizes were known. The AI's estimates were off by more than 3 cm in 11 of the 25, mostly because photos were taken at an angle with nothing in shot for scale.
That was a clear no for the idea as framed, after two afternoons and no contract. What the framer kept was the part that did work: AI drafting the reply that asks the customer for the measurements, in the shop's own wording, which is a drafting job with no proof of concept needed. Had the framer skipped straight to a pilot, customers would have received wrong quotes for weeks before anyone noticed the pattern.
Pilot: the same job inside real work, with a person checking
The pilot moves from samples to every price list as it arrives, for six weeks, done by the staff member who normally does the job. The AI extracts; a spreadsheet compares the result with current costs; the staff member checks every line that changed by more than 10%, plus a random ten others, before anything goes into the stock system.
The pilot is where the workflow, not the AI, gets tested. The pet shop found three things the proof of concept couldn't have shown:
- One supplier sends scanned price lists as photographs of paper. Extraction quality dropped badly, so that supplier was routed to manual entry as a named exception.
- Two suppliers changed their layout mid-pilot. The UNCLEAR markers jumped, which is how the change was noticed.
- The check took longer than expected in week 1 and settled at about 20 minutes per list by week 4.
The pilot ended against criteria written before it started: time on price updates fell from about 6 hours a month to about 2.5, no wrong cost reached the till, and the staff member preferred the new way. How to write criteria that can fail, and judge them fairly, is covered in setting AI pilot success criteria that hold up.
Production: the dull additions that keep it running
A pilot works because people are paying attention. Production has to work when they aren't. For the pet shop, moving to production meant adding:
- An automation for the routine part. Supplier emails with attachments go to a dedicated address; a no-code automation passes each attachment to an AI step and writes proposed changes to a spreadsheet for approval. The approval stays human.
- A named owner and a deputy, each with about 30 minutes a week for the approval queue and a monthly look at the run history.
- Failure handling. Zapier auto-pauses a Zap only when 95% of its runs error over seven days, and if you add an error handler, Zapier sends no error email when the handler runs, so the handler must alert someone itself. Make's error handlers (Skip, Retry, Resume, Commit and Rollback) need the same thought: decide which failures retry and which go to a person.
- Least access. The automation can read the price-list mailbox and write to one spreadsheet. It can't change prices in the stock system on its own.
- A runbook: one page on what the automation does, where it lives, how to pause it and how to do the job by hand if it's off.
- Business accounts throughout. Business AI plans don't train on business content by default, and neither does OpenAI's API; a pilot run on someone's personal account has to move before production.
The runbook sounds like paperwork, but it's what lets production survive a holiday or a staff change. The pet shop's fits on a page:
PRICE-LIST AUTOMATION - RUNBOOK
What it does: reads supplier price lists sent to prices@[shop],
extracts the lines and writes proposed cost changes to the
"Price changes" sheet for approval. It never changes the till.
Owner: [name]. Deputy: [name].
Weekly (30 min): approve or reject rows in "Price changes";
check every row marked UNCLEAR; import approved rows.
Monthly (15 min): open the run history; any errors? any supplier
with no list this month?
Known exceptions: [supplier] sends scans - enter by hand.
If it stops: pause the automation; forward new lists to [owner];
enter changes by hand from the PDF, as before.
Where things live: automation in the shop's Make account;
instructions in "Price list prompt" doc; this runbook in Drive.
Test the runbook by asking the deputy to follow it for one week while the owner watches without helping. Whatever they get stuck on is a missing line.
Production also inherits ongoing costs that don't show up in a pilot: someone updating the instructions when a supplier changes format, renewals, and the occasional rebuild when a tool changes. AI maintenance costs after go-live lists them so you can budget before, not after.
What it costs at each stage, in the pet shop's numbers
Here's where the money and hours go, with illustrative figures and list prices:
| Stage | Software | Staff time | What you learn |
|---|---|---|---|
| Proof of concept (2 weeks) | A business AI plan: two seats at about $25 each, so about $50 for the month (ChatGPT Business and Claude Team both have a two-seat minimum) | About 6 hours | Whether extraction is accurate on real samples |
| Pilot (6 weeks) | The same plan, about $75 for the period | About 3 hours a week | Whether the workflow works and saves time |
| Production (monthly) | Automation plan from about $9 (Make) to $29.99 (Zapier Professional, monthly billing), plus AI usage | About 2 hours a month on approvals and checks | Whether it keeps working unattended |
The AI part of the production bill is small. Say each price list is about 8,000 tokens in and 6,000 out. At Claude Haiku 4.5's API prices of $1 per million input tokens and $5 per million output, 12 lists a month cost about $0.10 for input and $0.36 for output: under 50 cents a month. On Zapier, an "AI by Zapier" step uses 1, 3 or 5 tasks per run depending on the model tier chosen, which 12 lists barely dent within Professional's 750 tasks. The automation platform, and above all the staff time, are the real costs.
Against that, the saving of about 3.5 hours a month in the pilot, plus fewer items sold below cost, made the decision easy. The numbers usually look like this: the software is cheap, and the time spent checking and maintaining is what decides whether production is worth it.
Gates between the stages: what must be true to move on
A gate is a short checklist you pass before spending more. Write both before the proof of concept starts.
From proof of concept to pilot
- The hard part worked on real samples, including awkward ones, at an accuracy you'd accept with a person checking.
- You know which case types failed, and there's a plan for them: better instructions, an exception route or exclusion.
- You can name the people who'll use it in the pilot and the hours they'll give.
- You've written the pilot's success criteria and stop rule.
From pilot to production
- The pilot met its written criteria in its later weeks, not just its best week.
- Every account involved belongs to the business, and access is limited to what the job needs.
- There's a named owner and a deputy, with the time in their week.
- Failures alert a person, and there's a manual fallback someone has actually tried.
- A runbook exists and someone other than its author has followed it.
- Full-volume costs have been estimated, including maintenance hours.
If a gate check fails, the answer isn't to skip it. Go back one stage with a specific fix, or narrow the scope to the part that passes.
When to skip a stage, and when skipping backfires
Skipping the proof of concept is often right. If the tool is a mature, general feature, such as drafting replies, summarising meetings or rewriting text, it's already in wide everyday use. What you don't know is how it fits your workflow, which is a pilot question. A proof of concept is for inputs nobody has tried: your supplier's odd PDFs, your handwritten forms, your specialist terms.
Skipping the pilot is rarely right, and it's never right for anything customers see. An illustrative tutoring agency saw a convincing demo of AI turning tutors' session notes into parent reports, and switched it on for every family the following week. Within days, one parent received a report describing the wrong subject, because the tutor had two students with the same first name. A two-week pilot with a coordinator checking each report would have caught it quietly. Instead, the agency spent a month rebuilding trust with parents who now read every report looking for mistakes. For AI that talks to customers, monitoring customer-facing AI covers the hand-offs and error checks production needs.
Skipping production discipline is the skip nobody notices making. The pilot "just keeps running" on the pilot's accounts and the pilot's checks, with the pilot's enthusiast holding it together. It works until that person is on holiday.
Signs you're stuck between stages
- The proof of concept keeps being re-run with new samples because nobody set an accuracy target. Set one and decide.
- The pilot has been "nearly done" for three months. Pilots need an end date and a single permitted extension. Beyond that, it's an unmanaged production system.
- Production has no owner's name on it. If nobody can say who fixes it when it breaks, it's still a pilot, however long it's been running.
- The tool is "in production" but staff route around it. Usage has quietly dropped; check the run history against the job's volume.
- A custom build is being planned before a no-code pilot has run. Build only after the pilot proves the job is worth it and shows exactly what the build needs to do; no-code AI versus a custom build explains where the line usually falls.
Each stage exists to answer one question cheaply before the next, more expensive question is asked. Keep them distinct and write down the answer at each gate, and even a failed idea costs you two weeks and a subscription rather than a year and a contract.
Further reads
- How to Run Your First AI Pilot Project in a Small Business — The step-by-step guide to running the middle stage.
- How to Pilot AI in Shadow Mode Before Customers See It — A safer pilot design for work customers will eventually see.
- Why Your AI Pilot Stalled, and How to Get It Live — Unstick a pilot that never reached its production gate.
- How Much Does an AI Pilot Project Cost? — What a pilot costs in fees, tools and staff time.
- How to Scale AI From One Workflow to the Whole Business — What comes after the first job reaches production.
- Automation Audit: Find the Zaps and Scenarios Nobody Owns — Keep production automations owned and documented over time.
- What a Typical AI Consulting Engagement Looks Like, Week by Week — A six-week AI consulting engagement explained stage by stage: what the consultant does, what you do, what you should have each Friday, and why weeks slip.
- How to Choose Your First AI Project: 7 Tests Before You Commit — Put each candidate for your first AI project through seven pass-or-fail tests, then commit to the one that passes all seven, with a one-page note.
- AI Tools and AI Development: The Complete 2026 Guide — the AI hub, including every tutorial in the AI-for-business series.
Sources: Anthropic and OpenAI API pricing pages; Zapier pricing and help centre (AI step task usage, auto-pause, error handlers); Make pricing and help centre (credits, error handlers); OpenAI and Anthropic business data policies.