How to Run a Paid Discovery Phase Before a Full AI Project

Coding Liquids tutorial cover featuring Sagnik Bhattacharya for How to Run a Paid Discovery Phase Before a Full AI Project.
Coding Liquids tutorial cover featuring Sagnik Bhattacharya for How to Run a Paid Discovery Phase Before a Full AI Project.

Run it as a short, fixed-fee piece of work with one question to answer: should we build this, and exactly what? Agree the question and deliverables in writing, give the provider access and real samples in week one, test the idea on 20 to 50 real cases, then finish with a costed build scope and a go or no-go.

The reason to pay is independence. A discovery you pay for can honestly conclude "don't build this", which a free sales call rarely will. Keep it small next to the build it informs, time-boxed to one to three weeks for a small business, and owned by you: every map, test result and scope document should be yours to keep and to show other providers, whoever does the build.

Follow me on Instagram@sagnikteaches

Discovery, audit and a free call are three different things

People use "discovery" loosely, so pin the meaning down before you buy one. A free introductory call tells you whether you and a provider can work together. An audit looks across the whole business for places AI could help, or checks the AI you already use. A discovery phase starts after you have picked one idea and answers whether that idea works, in your tools, with your data, at a cost that makes sense.

Connect on LinkedInSagnik Bhattacharya
Free intro callAI audit or readiness checkPaid discovery phase
Question it answersCan we work together?Where could AI help, and are we ready?Should we build this one thing, and exactly what?
Typical length15 to 60 minutesDays to a few weeksOne to three weeks for a small business
Uses your real dataNoSometimesYes: a sample of real cases is essential
What you walk away withImpressionsA ranked list of opportunities or gapsA test result, costed options and a build scope

If you haven't chosen the idea yet, a discovery is premature. Start with a wider check such as an audit or a readiness assessment, then come back with one question.

Subscribe on YouTube@codingliquids

The seven stages below follow a running example. A subscription box company with about 1,800 subscribers and four staff receives roughly 900 customer emails a month. Two people spend most of a working week each month answering them. The owner wants to know whether AI can take a real share of that load.

Stage 1: write the one question and the conditions for saying no

Time: about an hour of the owner's time. A good discovery question is narrow enough to test and written so the answer can be yes or no. "How can we use AI in customer service?" is an audit question. This is a discovery question:

"Can AI sort our customer emails into skip, pause, cancel, address change, damaged item, billing and other, and draft replies good enough to send with light edits, inside the helpdesk and tools we already have?"

Write the no-go conditions at the same time, before anyone has an incentive to soften them. For the box company: stop if fewer than 44 of 50 test emails are sorted correctly after two rounds of prompt changes; stop if running costs would exceed $150 a month; stop if the drafts would need the support lead's review for more than 30 minutes a day. Conditions written in advance turn the readout into a check rather than a debate.

Stage 2: agree the brief, the fee and who owns the results

Time: a 30-minute call and a one-page document. The brief is the discovery's contract in plain words. Here is the box company's version, filled in:

DISCOVERY BRIEF

Question: Can AI sort and draft replies for our customer emails
well enough to send with light edits, in our current tools?

In scope: support inbox emails from the last 90 days; the helpdesk;
the order system export; the refund and skip policy document.
Out of scope: churn prediction, marketing emails, website chat.

Deliverables:
 1. Map of the current email process, with timings
 2. Email volumes by type, from a sample of 300
 3. Feasibility test on 50 held-back emails, with scores
 4. Two or three build options, each with one-off and monthly costs
 5. Written build scope and acceptance criteria for the best option
 6. Risk list (data, privacy, customer impact, supplier)
 7. Go or no-go recommendation against the conditions in Stage 1

Time box: 3 weeks. Provider time: up to 6 days.
Client time: owner 4 hours, support lead 5 hours.
Fee: fixed. 50% on start, 50% at the readout.
Credited against a build with the same provider: [yes / no].
Ownership: all documents, test sets and prompts belong to the
client and may be shared with other providers to get build quotes.
Data: anonymised samples only; read-only access; provider deletes
copies within 30 days of the readout.

Two lines deserve attention. The ownership line lets you take the build scope to other providers, which keeps the discovery honest and the build quotes competitive. The credit line is worth asking about, but a credit that only applies if you hire the same provider is a mild form of lock-in; weigh it against the value of a free choice.

Fees for discovery vary widely with the provider's size. In September 2026, every firm on the first page of Clutch's directory of AI consultancies listed a minimum project of $10,000 or more, so many directory-listed firms are not set up for a small discovery. Independent consultants and automation specialists usually are. Zapier's Solution Partner directory lets you filter by project budget and minimum spend, and many listed partners offer a free 15- or 30-minute call: use that call to judge fit, not as the discovery itself. Compare discovery quotes on the provider days included and the deliverables list, not the headline fee.

Stage 3: hand over access and real samples in week one

Time: two to three hours of your time, usually the biggest cause of delay. A discovery run on invented examples tells you nothing. The box company pulled 350 real emails from the last 90 days: 300 for counting volumes and 50 held back for the feasibility test. Before handing them over, the support lead replaced names, addresses and order numbers with placeholders such as [first name] and [order no.], which took about 90 minutes with find-and-replace.

Give access, not passwords. Create a separate login for the provider with read-only rights where the tool allows it, and remove it after the readout. The detail of what a provider needs, and what they don't, is set out in what an AI consultant needs from you.

Week one is also where hidden set-up work surfaces. The box company's support inbox turned out to be a consumer @gmail.com address rather than a business workspace account. Make, one of the automation tools under consideration, needs @gmail.com users to set up their own Google Cloud sign-in client before it can connect to Gmail, Drive or Sheets. That is an hour or two of fiddly work that no build quote would have mentioned, and discovery found it for free. Data quality problems surface here too: the order export used three different names for the same box, which is exactly the kind of mess a data clean-up checklist is for.

Stage 4: watch the job being done, and time it

Time: one to two hours with the person who does the work. The provider should sit with the support lead, in person or on a shared screen, while 20 emails are handled for real. Every lookup, copy-paste and judgement call goes on the map. If you haven't mapped a process before, how to map a business process before you automate it shows the format.

From the box company's sample of 300 emails:

Email typeShare of sampleAverage handling timeNeeds a lookup?
Skip next box21%1.5 minSubscription status
Pause for one or more months13%2 minSubscription status
Address change18%2 minOrder system
Damaged or missing item14%6 minCourier tracking, photos
Cancel9%3 minSubscription status
Billing question11%4 minPayment system
Other14%3 minVaries

That works out at roughly 2.7 minutes an email, or about 40 hours a month across 900 emails. This is the baseline every later saving is measured against; setting a baseline before you introduce AI explains why skipping it makes any result impossible to prove.

Stage 5: test the idea on real cases

Time: two to three provider days, plus an hour of scoring from you. This is the heart of the discovery. The provider builds a rough prototype, often just a prompt run over the 50 held-back emails, and you score the results. The prompt tested at the box company looked like this:

You sort customer emails for a subscription box company.
Categories: SKIP (skip the next box only), PAUSE (stop for one or
more months, then restart), CANCEL, ADDRESS, DAMAGED (item damaged
or missing), BILLING, OTHER.

Return JSON with: category, confidence (high/medium/low),
second_issue (a second category if the email raises two things,
otherwise null), needs_human (true/false), draft_reply.

Rules: if there is a second issue, needs_human is true. Draft
replies under 120 words, friendly, signed "The [box name] team".
Only offer remedies that appear in the policy text. Never promise
a refund, replacement or date that the policy does not state.

Policy: [paste policy text]
Email: [paste anonymised email]

Here is what one test email ("Hi, can you skip my July box? Also the jar in June's box was smashed") produced, shown as an illustration:

{"category": "PAUSE", "confidence": "medium",
 "second_issue": "DAMAGED", "needs_human": true,
 "draft_reply": "Hi [first name], thanks for letting us know. We've
 paused your subscription for July and we'll send a replacement
 jar with your next box..."}

Two problems, both typical. The customer asked to skip one box, which the company treats as SKIP, not PAUSE; customers use the words interchangeably, so the fix was two worked examples in the prompt showing the difference. And the draft promised a replacement jar, which the policy only allows after a photo is received. Adding "ask for a photo before offering any remedy for DAMAGED" fixed it. Consistency comes from fixed instructions, examples and a review step like this, not from a settings tweak.

First run: 41 of 50 sorted correctly, 29 drafts usable with light edits. Second run, after the fixes: 47 of 50 sorted correctly, 38 drafts usable. The three misses all mixed two issues, and every one was flagged needs_human, which is the safe failure. That clears the 44-of-50 condition from Stage 1.

Stage 6: cost the options, including doing nothing new

Time: one provider day. A good discovery offers two or three options, and one of them should be cheap and dull. For the box company, with the figures worked from vendor list prices:

OptionMonthly running costEstimated time savedMain risk
1. Saved replies and helpdesk rules, no AI$0 extraAbout 8 hours a monthStaff still read and sort every email
2. AI sorts and drafts; staff approve every reply (built in Make with the company's own model key)Make's entry plan is $9 a month for 5,000 credits; model usage well under $1 a monthAbout 20 hours a monthCredits run short if the trigger polls too often
3. AI agent replies to customers directly, priced per resolutionAt Intercom Fin's $0.99 per resolved outcome, 500 resolutions would be about $495 a monthThe most, if resolutions are genuineA wrong answer goes straight to a customer

The quick sums behind option 2 are worth checking yourself. Four modules per email at one credit each is 3,600 credits for 900 emails. A trigger that checks the inbox every 15 minutes uses a credit per check even when nothing has arrived, about 2,880 credits a month, which pushes the total past the 5,000 included. The fix is to check less often (hourly checks use about 720 credits a month, bringing the total to roughly 4,320) or to buy extra credits, and the discovery should say which. Model costs are tiny: 900 emails at about 1,500 input and 250 output tokens each, on gpt-6-luna at $0.10 per million input tokens and $0.50 per million output tokens, comes to roughly 25 cents a month.

Option 3 needs one caution in the report: per-resolution billing counts what the vendor defines as resolved. Intercom, for example, bills an "assumed resolution" once a customer has said nothing for 24 hours after Fin's final reply, which is not the same thing as a happy customer.

Stage 7: the readout and the go or no-go call

Time: a 60-minute meeting. Run it against the conditions you wrote in Stage 1, in this order:

  1. Did the feasibility test meet the threshold? (Box company: 47 of 50 against a condition of 44. Yes.)
  2. Do running costs fit the limit? (Option 2: about $10 a month with hourly inbox checks. Yes.)
  3. Can the team give the review time the recommended option needs? (Support lead: about 20 minutes a day approving drafts. Yes.)
  4. Is there a risk on the list that nobody can live with? (None for option 2; option 3's customer-facing risk was parked.)

Four yeses mean go, with option 2. The final deliverable is the build scope, written with acceptance criteria the builder must meet; the method is in how to scope an AI project. Because you own it, you can send it to two or three providers, including the one that ran the discovery, and compare like with like. With a scope this tight, a fixed price becomes realistic, as explained in fixed-scope versus hourly AI consulting.

A food truck discovery that ended in a no

A no is a useful result, and cheaper than a failed build. A two-truck street food business wanted an AI phone line to take lunch orders, having seen AI receptionists advertised. The discovery question: "Can an AI phone agent take and confirm pre-orders accurately enough to cut queue times at lunch?"

Stage 4 changed the picture. Timing a week of trade showed about 12 phone calls a day, most of them asking where the truck would be, while more than 90% of orders happened at the window. The queue problem came from payment and cooking time, not phone calls. The recommendation: no AI phone agent; publish the weekly location schedule on the website and social profiles, and add a simple pre-order form for office groups. The owner spent a few days of discovery fee and avoided a monthly phone subscription that would have answered location questions a web page could handle.

Signs a discovery is a sales pitch in disguise

Most providers run discovery in good faith, but the format can be bent into a long sales meeting. Watch for these:

  • Every option needs the provider's own product or platform. A genuine discovery includes at least one option they don't profit from, such as doing it with tools you have.
  • No real data is ever requested. Without your samples, the "feasibility test" is a demo on friendly examples.
  • The no-go conditions are missing or vague. If nobody wrote down what failure looks like, the answer will be yes.
  • Deliverables stay in the provider's accounts. A prototype you can't see and a slide deck without numbers are hard to take elsewhere.
  • The build quote arrives before the test results. The price should follow the evidence, not precede it.
  • The question grows. "While we're here, let's look at churn prediction too" doubles the time and blurs the answer. Park new ideas for a second discovery.

What you should hold at the end

Whatever the answer, check the handover against this list before paying the final instalment:

  • The process map with timings, in a format you can edit
  • The test set of real cases, the prompt used and the scored results for each run
  • Volumes and the baseline figure (hours per month, cost per case)
  • The options table with one-off and monthly costs, and the vendor pages the prices came from
  • The written build scope with acceptance criteria, or the reasons for a no
  • The risk list, with an owner for each risk
  • Confirmation that the provider's access has been removed and copies of your data deleted

If all seven are in your hands, the discovery did its job, whether it ended in a build, a smaller change or a well-evidenced no.

Further reads

Sources: Make pricing and credits help pages; Make help on connecting Google accounts; Intercom Fin pricing; OpenAI API pricing; Clutch AI consulting directory and Zapier Solution Partner directory (checked September 2026).

Want help framing your discovery question?

On a 1:1 call we can narrow your idea to one testable question, agree what the discovery should hand you, and decide which real samples to gather first.

Book a 1:1 call with me