AI Proof of Concept vs Pilot vs Production: What Changes?

Coding Liquids tutorial cover featuring Sagnik Bhattacharya for AI Proof of Concept vs Pilot vs Production: What Changes?
Coding Liquids tutorial cover featuring Sagnik Bhattacharya for AI Proof of Concept vs Pilot vs Production: What Changes?

The question changes. A proof of concept asks whether AI can do the task at all, on 20 to 100 real samples in a week or two. A pilot asks whether it works inside your workflow, with a few staff for four to eight weeks. Production asks whether it keeps working, adding an owner, monitoring, access controls and an off-switch.

Many small businesses can skip the proof of concept. If you're switching on a feature the vendor already sells, such as drafting in Gmail or Copilot Chat, you know it works in general, so go straight to a pilot. A proof of concept earns its place when nobody knows whether AI can handle your particular inputs: messy scans, unusual formats or your trade's own jargon.

Follow me on Instagram@sagnikteaches

The three stages side by side

What changesProof of conceptPilotProduction
The questionCan AI do this at all?Does it work in our real workflow?Will it keep working without attention?
Data20 to 100 real samples, copied outAll real cases of one type, as they arriveEverything in scope, plus the odd cases
Who uses itThe person testing itTwo or three staffEveryone who does the job
How longDays to two weeksFour to eight weeksOpen-ended, reviewed quarterly
AccountsA business AI account; no connectionsBusiness accounts; limited connectionsBusiness accounts, named owner, least access needed
CheckingEvery output compared with the right answerA person checks everything before useChecks on risky cases and a regular sample
When it's wrongNote it; adjust the instructionsThe checker catches it; log itAlerts, an exceptions route and a fallback
DocumentationNotes on what workedInstructions and a results logA runbook anyone can follow
How it endsA yes, a no or "only for these cases"Pass, extend once, redesign or stopA planned review, replacement or switch-off

Read down each column and the shift is from "can it?" through "does it?" to "will it, reliably, when nobody's watching?" Each stage adds cost, but mostly in staff time and discipline rather than software.

Connect on LinkedInSagnik Bhattacharya

Proof of concept: proving the hard part on real samples

Consider an illustrative pet shop with two branches and an online shop, stocking around 2,400 products from 11 suppliers. Suppliers send price updates in their own formats, some as spreadsheets and some as PDFs, about 12 a month with 150 to 600 lines each. A staff member spends around 6 hours a month finding the changes and typing them into the stock system, and a missed increase means selling at the old price, sometimes below cost.

Subscribe on YouTube@codingliquids

The uncertain part was whether AI could read three suppliers' messy PDFs accurately enough, so that's all the proof of concept tested. The owner took 30 lines from each of three recent price lists, 90 in all, and ran them through a business AI workspace with this prompt:

From the attached supplier price list, extract every product line
as a table with these columns: supplier code, product description,
pack size, unit cost, case cost. Copy codes exactly as printed. If
a value is unclear or missing, write UNCLEAR. Don't calculate or
infer any figure that isn't printed.

An illustrative slice of the output:

Code     Description               Pack        Unit cost  Case cost
DF-2210  Adult dog food, chicken   12 x 400g   UNCLEAR    14.40
CT-0871  Cat litter, clumping      10kg        6.95       UNCLEAR
DF-2214  Puppy food, lamb          12 x 400g   1.35       16.20

What the owner checked and fixed: 86 of the 90 lines were right. The four errors were all pack-size confusions, where a case price was reported as a unit price or the reverse. The instructions were changed to "report unit cost and case cost separately; if only one is printed, mark the other UNCLEAR", and a fresh 30 lines came back with no errors. The UNCLEAR markers are a feature: they show where a person must look rather than letting the AI guess.

Two weeks, about 6 hours and one month of a business plan. The proof of concept's result wasn't "AI works"; it was "AI can extract these three suppliers' formats reliably when told to separate unit and case prices". That narrower answer is the useful one. For more on reading documents like these, AI document processing covers the common formats and their traps.

A proof of concept that says no is a success

The stage is worth running precisely because it can fail cheaply. An illustrative picture framer wanted AI to estimate frame sizes and mount widths from the photos customers email of their artwork, so that rough quotes could go out without a visit to the shop. The proof of concept used 25 past enquiries where the true sizes were known. The AI's estimates were off by more than 3 cm in 11 of the 25, mostly because photos were taken at an angle with nothing in shot for scale.

That was a clear no for the idea as framed, after two afternoons and no contract. What the framer kept was the part that did work: AI drafting the reply that asks the customer for the measurements, in the shop's own wording, which is a drafting job with no proof of concept needed. Had the framer skipped straight to a pilot, customers would have received wrong quotes for weeks before anyone noticed the pattern.

Pilot: the same job inside real work, with a person checking

The pilot moves from samples to every price list as it arrives, for six weeks, done by the staff member who normally does the job. The AI extracts; a spreadsheet compares the result with current costs; the staff member checks every line that changed by more than 10%, plus a random ten others, before anything goes into the stock system.

The pilot is where the workflow, not the AI, gets tested. The pet shop found three things the proof of concept couldn't have shown:

  • One supplier sends scanned price lists as photographs of paper. Extraction quality dropped badly, so that supplier was routed to manual entry as a named exception.
  • Two suppliers changed their layout mid-pilot. The UNCLEAR markers jumped, which is how the change was noticed.
  • The check took longer than expected in week 1 and settled at about 20 minutes per list by week 4.

The pilot ended against criteria written before it started: time on price updates fell from about 6 hours a month to about 2.5, no wrong cost reached the till, and the staff member preferred the new way. How to write criteria that can fail, and judge them fairly, is covered in setting AI pilot success criteria that hold up.

Production: the dull additions that keep it running

A pilot works because people are paying attention. Production has to work when they aren't. For the pet shop, moving to production meant adding:

  • An automation for the routine part. Supplier emails with attachments go to a dedicated address; a no-code automation passes each attachment to an AI step and writes proposed changes to a spreadsheet for approval. The approval stays human.
  • A named owner and a deputy, each with about 30 minutes a week for the approval queue and a monthly look at the run history.
  • Failure handling. Zapier auto-pauses a Zap only when 95% of its runs error over seven days, and if you add an error handler, Zapier sends no error email when the handler runs, so the handler must alert someone itself. Make's error handlers (Skip, Retry, Resume, Commit and Rollback) need the same thought: decide which failures retry and which go to a person.
  • Least access. The automation can read the price-list mailbox and write to one spreadsheet. It can't change prices in the stock system on its own.
  • A runbook: one page on what the automation does, where it lives, how to pause it and how to do the job by hand if it's off.
  • Business accounts throughout. Business AI plans don't train on business content by default, and neither does OpenAI's API; a pilot run on someone's personal account has to move before production.

The runbook sounds like paperwork, but it's what lets production survive a holiday or a staff change. The pet shop's fits on a page:

PRICE-LIST AUTOMATION - RUNBOOK
What it does: reads supplier price lists sent to prices@[shop],
  extracts the lines and writes proposed cost changes to the
  "Price changes" sheet for approval. It never changes the till.
Owner: [name]. Deputy: [name].
Weekly (30 min): approve or reject rows in "Price changes";
  check every row marked UNCLEAR; import approved rows.
Monthly (15 min): open the run history; any errors? any supplier
  with no list this month?
Known exceptions: [supplier] sends scans - enter by hand.
If it stops: pause the automation; forward new lists to [owner];
  enter changes by hand from the PDF, as before.
Where things live: automation in the shop's Make account;
  instructions in "Price list prompt" doc; this runbook in Drive.

Test the runbook by asking the deputy to follow it for one week while the owner watches without helping. Whatever they get stuck on is a missing line.

Production also inherits ongoing costs that don't show up in a pilot: someone updating the instructions when a supplier changes format, renewals, and the occasional rebuild when a tool changes. AI maintenance costs after go-live lists them so you can budget before, not after.

What it costs at each stage, in the pet shop's numbers

Here's where the money and hours go, with illustrative figures and list prices:

StageSoftwareStaff timeWhat you learn
Proof of concept (2 weeks)A business AI plan: two seats at about $25 each, so about $50 for the month (ChatGPT Business and Claude Team both have a two-seat minimum)About 6 hoursWhether extraction is accurate on real samples
Pilot (6 weeks)The same plan, about $75 for the periodAbout 3 hours a weekWhether the workflow works and saves time
Production (monthly)Automation plan from about $9 (Make) to $29.99 (Zapier Professional, monthly billing), plus AI usageAbout 2 hours a month on approvals and checksWhether it keeps working unattended

The AI part of the production bill is small. Say each price list is about 8,000 tokens in and 6,000 out. At Claude Haiku 4.5's API prices of $1 per million input tokens and $5 per million output, 12 lists a month cost about $0.10 for input and $0.36 for output: under 50 cents a month. On Zapier, an "AI by Zapier" step uses 1, 3 or 5 tasks per run depending on the model tier chosen, which 12 lists barely dent within Professional's 750 tasks. The automation platform, and above all the staff time, are the real costs.

Against that, the saving of about 3.5 hours a month in the pilot, plus fewer items sold below cost, made the decision easy. The numbers usually look like this: the software is cheap, and the time spent checking and maintaining is what decides whether production is worth it.

Gates between the stages: what must be true to move on

A gate is a short checklist you pass before spending more. Write both before the proof of concept starts.

From proof of concept to pilot

  • The hard part worked on real samples, including awkward ones, at an accuracy you'd accept with a person checking.
  • You know which case types failed, and there's a plan for them: better instructions, an exception route or exclusion.
  • You can name the people who'll use it in the pilot and the hours they'll give.
  • You've written the pilot's success criteria and stop rule.

From pilot to production

  • The pilot met its written criteria in its later weeks, not just its best week.
  • Every account involved belongs to the business, and access is limited to what the job needs.
  • There's a named owner and a deputy, with the time in their week.
  • Failures alert a person, and there's a manual fallback someone has actually tried.
  • A runbook exists and someone other than its author has followed it.
  • Full-volume costs have been estimated, including maintenance hours.

If a gate check fails, the answer isn't to skip it. Go back one stage with a specific fix, or narrow the scope to the part that passes.

When to skip a stage, and when skipping backfires

Skipping the proof of concept is often right. If the tool is a mature, general feature, such as drafting replies, summarising meetings or rewriting text, it's already in wide everyday use. What you don't know is how it fits your workflow, which is a pilot question. A proof of concept is for inputs nobody has tried: your supplier's odd PDFs, your handwritten forms, your specialist terms.

Skipping the pilot is rarely right, and it's never right for anything customers see. An illustrative tutoring agency saw a convincing demo of AI turning tutors' session notes into parent reports, and switched it on for every family the following week. Within days, one parent received a report describing the wrong subject, because the tutor had two students with the same first name. A two-week pilot with a coordinator checking each report would have caught it quietly. Instead, the agency spent a month rebuilding trust with parents who now read every report looking for mistakes. For AI that talks to customers, monitoring customer-facing AI covers the hand-offs and error checks production needs.

Skipping production discipline is the skip nobody notices making. The pilot "just keeps running" on the pilot's accounts and the pilot's checks, with the pilot's enthusiast holding it together. It works until that person is on holiday.

Signs you're stuck between stages

  • The proof of concept keeps being re-run with new samples because nobody set an accuracy target. Set one and decide.
  • The pilot has been "nearly done" for three months. Pilots need an end date and a single permitted extension. Beyond that, it's an unmanaged production system.
  • Production has no owner's name on it. If nobody can say who fixes it when it breaks, it's still a pilot, however long it's been running.
  • The tool is "in production" but staff route around it. Usage has quietly dropped; check the run history against the job's volume.
  • A custom build is being planned before a no-code pilot has run. Build only after the pilot proves the job is worth it and shows exactly what the build needs to do; no-code AI versus a custom build explains where the line usually falls.

Each stage exists to answer one question cheaply before the next, more expensive question is asked. Keep them distinct and write down the answer at each gate, and even a failed idea costs you two weeks and a subscription rather than a year and a contract.

Further reads

Sources: Anthropic and OpenAI API pricing pages; Zapier pricing and help centre (AI step task usage, auto-pause, error handlers); Make pricing and help centre (credits, error handlers); OpenAI and Anthropic business data policies.

Not sure which stage your AI project is really at?

On a 1:1 call we'll look at what you've built or tested so far, decide whether it needs a proof of concept, a proper pilot or the production basics, and list what's missing for the next gate.

Book a 1:1 call with me