How to Scope an AI Project: Deliverables and Acceptance Criteria

Coding Liquids tutorial cover featuring Sagnik Bhattacharya for How to Scope an AI Project: Deliverables and Acceptance Criteria.
Coding Liquids tutorial cover featuring Sagnik Bhattacharya for How to Scope an AI Project: Deliverables and Acceptance Criteria.

Scope an AI project by writing down five things before anyone quotes: the single job the AI will do, the inputs it receives, the outputs it must produce, the deliverables you'll own at the end, and acceptance criteria with a pass mark tested on real examples. Then list what's out of scope and what each side must provide.

AI changes one thing about ordinary project scoping: the output is never right every time, so "it works" can't be the acceptance test. Criteria need a pass mark on a set of real cases the builder hasn't seen, rules for what must never happen, and a plan for re-testing when the model underneath changes. Get those three right and most disputes never start.

Follow me on Instagram@sagnikteaches

The sections below build one real-looking scope in the order you would write it. A three-person bespoke furniture maker receives about 40 enquiries a month through its website form and email, many with photos and rough measurements. The owner spends around 30 minutes on each one, replying with a price range and a list of questions. The project: an AI step reads each enquiry, pulls out the details, drafts a reply with a price range from the workshop's price guide, and asks for anything missing. The owner approves every reply before it goes.

Connect on LinkedInSagnik Bhattacharya

One job, one judge and the number that justifies the spend

A scope that covers "AI for enquiries, quotes and follow-ups" is three projects wearing one name. Pick the job that costs the most time or loses the most money and write it as a single sentence with a verb, an input and an output:

Subscribe on YouTube@codingliquids

"For every new enquiry, extract the piece type, wood, finish and dimensions, and draft a reply with a price range and questions for any missing details, for the owner to approve."

Name the person on your side who will judge the results (here, the owner) and the measure that justifies the spend: 40 enquiries at 30 minutes is 20 hours a month, and the target is under 8 minutes each including review. Without a baseline like that, nobody can later say whether the project paid off.

Inputs described exactly as they arrive

Builders scope against tidy examples; your real inputs are untidy. Collect ten recent ones and describe their variety honestly. The furniture maker's list:

  • Website form: name, email, piece type from a drop-down, a free-text box, up to three photos
  • Emails forwarded from the owner's personal address, sometimes with a PDF drawing attached
  • Dimensions in centimetres, millimetres or inches, and sometimes "about two metres" or "to fit the alcove in the photo"
  • References to past work: "like the oak table on your Instagram but longer"
  • Roughly one enquiry in eight isn't a new commission at all: repairs, trade enquiries, job applications

Every item on that list is a potential defect if the builder never saw it coming. Hand over 30 anonymised real examples for building and tuning, and keep a separate set back for testing (the acceptance criteria below explain why).

Outputs specified down to the field names

"A draft reply" is not an output specification. Say what fields come out, in what format, and where they land. Structured output, meaning the AI fills named fields rather than writing free text, makes the result checkable and far more consistent. The furniture maker's output, one row per enquiry in a shared sheet plus a draft in the inbox:

FieldExample valueRule
piece_typeDining tableOne of the 12 types in the price guide, or "Other"
wood / finishOak / hard wax oil"Not stated" if missing; never guessed
dimensions_cm220 x 95 x 76Converted to centimetres; original text kept alongside
price_low / price_high$3,400 / $4,100Calculated only from price-guide rows, never invented
missing_infoFinish; delivery accessEvery field the price guide needs that is absent
needs_ownerFalseTrue for repairs, trade, unclear or unusual requests
draft_replyUnder 150 wordsSaved as a draft; never sent automatically

Say how consistency will be achieved, too. A common assumption is that setting the model's "temperature" to zero makes answers repeatable, but Anthropic has deprecated that setting on its newer Claude models and OpenAI's reasoning models don't support it. Fixed instructions, worked examples and structured output fields, with a person reviewing the result, are what make answers repeatable now. If you're curious why AI varies at all, why AI gives different answers each time covers the mechanics.

Deliverables: what you hold when the provider leaves

Deliverables are the things you hold when the provider leaves, not the activities they perform. Write each one as a noun you could point at:

  1. The working workflow, built in accounts registered to your business (automation tool, AI model key, shared sheet)
  2. The prompt text and the price-guide lookup, in an editable document with a version number
  3. The tuning set and the acceptance test set, with the scored results of every test run
  4. A run-book of two pages or fewer: how to pause the workflow, how to edit the prompt, where the log is, what to do when it errors
  5. A handover session of about an hour, recorded, in which you pause, edit and restart it yourself
  6. Removal of the provider's access on a stated date, confirmed in writing

Item 1 matters more than it looks. A workflow built in the provider's own account is theirs in practice, whatever the contract says, and moving it later is a rebuild. The full list of what to collect before a provider leaves is in the AI consultant handover checklist.

One deliverable to question: if a proposal offers a custom GPT, note that OpenAI is retiring them. They stop running on 11 December 2026, and OpenAI's migration turns each one into a plugin. For shared business context in ChatGPT, a Project is the current equivalent, and a deliverable built on something being retired is a deliverable with a short life.

Acceptance criteria that allow for AI's misses

This is the step most scopes get wrong. "The AI accurately extracts enquiry details" can't be tested, because it doesn't say how accurate, on what, or measured by whom. Good criteria for AI work come in four kinds: a pass mark on held-back real cases, must-never rules with zero tolerance, correct handling of planted awkward cases, and practical limits on speed and cost. The furniture maker's set:

CriterionHow it's measuredPass mark
Extraction accuracy40 held-back enquiries; 5 key fields each checked against the owner's own readingAt least 190 of 200 fields correct
Price rangeCompared with the range the owner would give for the same 40Overlaps the owner's range in at least 36 of 40
Questions for missing detailsOwner marks whether the draft asks for everything neededAt least 34 of 40 need no added question
Must-never rulesAny reply sent without approval; any price below the guide's minimum; any promised delivery dateZero occurrences in all testing
Planted awkward cases6 cases: repair request, trade order, job application, abusive message, blank form, photo-only emailAll 6 flagged needs_owner
SpeedTime from enquiry arriving to draft appearingUnder 5 minutes
Running costSubscriptions plus model usage at 40 enquiries a monthUnder $60 a month

Two rules make these numbers mean something. First, the acceptance set must be held back: if the prompt is tuned on the same enquiries it is tested on, it can score 40 out of 40 and still stumble on next month's post. Second, the test set must look like real life, awkward cases included in the proportion they actually occur. A test set of 40 easy enquiries passes anything.

The quick sum on the running-cost line: at 40 enquiries a month, even a workflow that uses seven Zapier tasks per enquiry needs 280 tasks, well inside the 750 included in Zapier's Professional plan at $29.99 a month on monthly billing, and model usage for 40 short enquiries costs cents. Under $60 a month leaves room for the photo-reading step if it is added later.

Acceptance criteria overlap with pilot success criteria but aren't the same thing: acceptance asks "was it built as agreed?", while a pilot asks "does it pay off in daily use?". The second question is covered in setting success criteria for an AI pilot that hold up, and a safe way to gather that evidence is piloting in shadow mode before customers see anything.

The out-of-scope list and each side's duties

The out-of-scope list is the cheapest insurance in the whole document. For the furniture maker:

  • Out of scope: sending replies automatically; taking deposits; booking site visits; estimating dimensions from photos; social media messages; changes to the website form; moving enquiries into a CRM.
  • Client provides: price guide v2 as one spreadsheet before build starts; 70 anonymised past enquiries (30 for tuning, 40 held back); accounts in the business's name; two hours of testing time within five working days of being asked.
  • Provider provides: the build; all six deliverables listed earlier; fixes for defects against the criteria for 30 days after acceptance.

When someone asks mid-build for "just one more thing", this list is where the conversation starts. A good process for those requests is in how to handle change requests in AI projects.

How the acceptance test actually runs

Criteria without a process still end in argument. Settle these points in the scope:

  1. Who runs the test. The provider runs the 40 held-back cases through the final version while you watch, or you run them yourself; either way the results go into a shared sheet.
  2. How defects are graded. Critical: breaks a must-never rule, so fix and re-run all 40. Major: misses a pass mark, so fix and re-run the affected criterion. Minor: wording or formatting, so list for the warranty period.
  3. How many re-test rounds are included. Two is common; further rounds caused by new requirements are change requests.
  4. How long you have to test. Five working days is typical. Be wary of "deemed acceptance" wording that treats silence as sign-off; at least require a written reminder first.
  5. What happens after acceptance. A warranty period for defects against these criteria, and who re-runs the test (at what cost) if the model, the prompt or a connected app changes later.

Here is an excerpt from the furniture maker's first test round, scored in the shared sheet:

CaseWhat happenedGrade
T07Dimensions given in inches and converted correctly, but the draft forgot to ask about stair access for deliveryMinor
T19Wood recorded as oak from the photo, though the customer wrote "walnut or oak, not sure"Major: guessed a field the rules say must never be guessed
T26Price range below the guide's minimum for a 2.4-metre tableCritical: a must-never rule broken
T33Planted repair request flagged needs_ownerPass

Round one scored 186 of 200 fields (the pass mark is 190), 35 of 40 overlapping price ranges and one critical defect. The builder fixed the size-band lookup behind T26 and added the instruction "if the customer names more than one wood, record all of them and ask which", then re-ran all 40 cases, as the rules require after a critical defect. Round two: 194 of 200 fields, 38 of 40 ranges, no critical defects, and all six planted cases flagged. The project was accepted, with T07's missing question logged for the warranty period. Notice that nobody argued about whether it "worked": the sheet said so.

The furniture maker's one-page scope, assembled

PROJECT: Enquiry reader and reply drafter
OWNER (client side): [owner's name]

JOB: For every new enquiry, extract piece type, wood, finish and
dimensions, and draft a reply with a price range and questions for
missing details. The owner approves every reply before sending.

INPUTS: website form (fields + up to 3 photos); forwarded emails
with attachments; mixed units; ~1 in 8 not a new commission.

OUTPUTS: one row per enquiry in the shared sheet (fields listed in
Annex A) and a draft reply in the shared inbox. Nothing sent
automatically.

DELIVERABLES: workflow in client-owned accounts; prompt and lookup
document (versioned); tuning and test sets with scored results;
two-page run-book; recorded handover (~1 hour); provider access
removed on [date].

ACCEPTANCE: criteria and pass marks in Annex B, tested on 40
held-back enquiries plus 6 planted cases. Critical defects: re-run
all tests. Two re-test rounds included. Client testing window: 5
working days after notice.

OUT OF SCOPE: auto-sending, deposits, site-visit booking, dimension
estimates from photos, social messages, form changes, CRM migration.

CLIENT PROVIDES: price guide v2 (one sheet) before build; 70
anonymised past enquiries; accounts in client's name; testing time.

AFTER ACCEPTANCE: 30 days of defect fixes against Annex B. Re-test
after any model or app change: [who / how charged].

It fits on one page, yet it answers the questions that usually surface as disputes in week four.

Drafting the scope with an AI assistant, and what to strike out

An assistant can turn your rough notes into this structure quickly, provided you police the result. A workable prompt:

Here are my notes about an AI project for my business: [notes].
Draft a one-page scope with these sections: job, inputs, outputs,
deliverables, acceptance criteria, out of scope, client duties,
after acceptance. Every acceptance criterion must include how it is
measured and a numeric pass mark. Do not add any system, feature or
integration that is not in my notes; if something is unclear, list
it as a question at the end instead.

The first draft that came back, in this illustrative case, had two typical faults:

"Acceptance criteria: The system will achieve high accuracy in extracting enquiry details. Responses will be generated promptly. Integration with the accounting system will allow quotes to become invoices."

"High accuracy" and "promptly" have no numbers, despite the instruction, so they go back as "at least 190 of 200 fields" and "under 5 minutes". The accounting integration appeared from nowhere; it wasn't in the notes, and left in, it would have become an expensive line in every quote. Strike anything you didn't ask for, and treat the assistant's draft as a checklist of headings rather than a finished scope.

Scoping gaps that surface after sign-off

Three realistic failures, each traceable to one missing line.

No named source of truth. A sports equipment shop accepted a returns-question chatbot as delivered. Three weeks later it was quoting a 60-day returns window from an old PDF the builder had loaded, although the current policy on the website said 30 days. The scope never said which document the bot answers from or who updates it. The missing line: "Knowledge source: the returns page at [address], re-read nightly; owner: [name]."

Tested on the tuning set. A catering company's quote drafter scored 30 out of 30 at acceptance, because the test used the same 30 enquiries the prompt had been refined on. On new enquiries it mispriced canapé orders by guest count about one time in five. The missing line: "Acceptance uses 40 enquiries withheld from the provider until testing."

No baseline, so no verdict. An e-commerce homeware brand automated product-question replies with a clear build scope and a passed test, but nobody had timed the old process. Six months later, asked whether it was worth renewing, nobody could say. The missing line: "Baseline: current handling time per enquiry, measured over two weeks before build starts."

A scope with a named job, specified inputs and outputs, owned deliverables, numeric acceptance criteria, an out-of-scope list and a testing process is not bureaucracy. It is what lets you compare quotes fairly, accept work with confidence and, if it comes to it, walk away from a build that didn't meet what was agreed.

Scoping and acceptance testing: your questions

How many test cases do I need for acceptance testing?

Enough that one lucky or unlucky case can't swing the result. With 10 cases, a single case moves the score by 10 points, so aim for 40 to 60 real examples for a workflow that runs dozens of times a month. Include the awkward ones in the proportion they really occur, plus a handful of planted edge cases that must be routed to a person.

Should the builder see the acceptance test set?

No. Give the builder a separate set of real examples for tuning and keep the acceptance set back until testing. If the prompt is tuned on the same cases it is tested on, it can pass perfectly and then perform worse on new work. Sharing the test set after acceptance, to help with fixes, is fine.

What happens if the AI passes acceptance and then gets worse?

Write this into the scope before you sign. A warranty period covers defects against the agreed criteria for a set time, often 30 days. Beyond that, agree who re-runs the acceptance test when a model, prompt or connected app changes, and what it costs. Keeping the test set and the scoring sheet means anyone can repeat the check later.

Who should write the scope, me or the provider?

Draft it yourself first, even roughly, because you know the job, the exceptions and what a good result looks like. Then let the provider tighten it: they know what can be built and tested. A scope written only by the provider tends to describe what they plan to deliver rather than what you need to accept.

Further reads

Sources: OpenAI help article on custom GPT retirement and migration; Anthropic and OpenAI documentation on temperature and structured outputs; Zapier pricing page.

Want a second pair of eyes on your AI scope?

On a 1:1 call we can turn your description of the job into inputs, outputs and acceptance criteria with real pass marks, and spot what a quote would otherwise leave out.

Book a 1:1 call with me