Scope an AI project by writing down five things before anyone quotes: the single job the AI will do, the inputs it receives, the outputs it must produce, the deliverables you'll own at the end, and acceptance criteria with a pass mark tested on real examples. Then list what's out of scope and what each side must provide.
AI changes one thing about ordinary project scoping: the output is never right every time, so "it works" can't be the acceptance test. Criteria need a pass mark on a set of real cases the builder hasn't seen, rules for what must never happen, and a plan for re-testing when the model underneath changes. Get those three right and most disputes never start.
The sections below build one real-looking scope in the order you would write it. A three-person bespoke furniture maker receives about 40 enquiries a month through its website form and email, many with photos and rough measurements. The owner spends around 30 minutes on each one, replying with a price range and a list of questions. The project: an AI step reads each enquiry, pulls out the details, drafts a reply with a price range from the workshop's price guide, and asks for anything missing. The owner approves every reply before it goes.
One job, one judge and the number that justifies the spend
A scope that covers "AI for enquiries, quotes and follow-ups" is three projects wearing one name. Pick the job that costs the most time or loses the most money and write it as a single sentence with a verb, an input and an output:
"For every new enquiry, extract the piece type, wood, finish and dimensions, and draft a reply with a price range and questions for any missing details, for the owner to approve."
Name the person on your side who will judge the results (here, the owner) and the measure that justifies the spend: 40 enquiries at 30 minutes is 20 hours a month, and the target is under 8 minutes each including review. Without a baseline like that, nobody can later say whether the project paid off.
Inputs described exactly as they arrive
Builders scope against tidy examples; your real inputs are untidy. Collect ten recent ones and describe their variety honestly. The furniture maker's list:
- Website form: name, email, piece type from a drop-down, a free-text box, up to three photos
- Emails forwarded from the owner's personal address, sometimes with a PDF drawing attached
- Dimensions in centimetres, millimetres or inches, and sometimes "about two metres" or "to fit the alcove in the photo"
- References to past work: "like the oak table on your Instagram but longer"
- Roughly one enquiry in eight isn't a new commission at all: repairs, trade enquiries, job applications
Every item on that list is a potential defect if the builder never saw it coming. Hand over 30 anonymised real examples for building and tuning, and keep a separate set back for testing (the acceptance criteria below explain why).
Outputs specified down to the field names
"A draft reply" is not an output specification. Say what fields come out, in what format, and where they land. Structured output, meaning the AI fills named fields rather than writing free text, makes the result checkable and far more consistent. The furniture maker's output, one row per enquiry in a shared sheet plus a draft in the inbox:
| Field | Example value | Rule |
|---|---|---|
| piece_type | Dining table | One of the 12 types in the price guide, or "Other" |
| wood / finish | Oak / hard wax oil | "Not stated" if missing; never guessed |
| dimensions_cm | 220 x 95 x 76 | Converted to centimetres; original text kept alongside |
| price_low / price_high | $3,400 / $4,100 | Calculated only from price-guide rows, never invented |
| missing_info | Finish; delivery access | Every field the price guide needs that is absent |
| needs_owner | False | True for repairs, trade, unclear or unusual requests |
| draft_reply | Under 150 words | Saved as a draft; never sent automatically |
Say how consistency will be achieved, too. A common assumption is that setting the model's "temperature" to zero makes answers repeatable, but Anthropic has deprecated that setting on its newer Claude models and OpenAI's reasoning models don't support it. Fixed instructions, worked examples and structured output fields, with a person reviewing the result, are what make answers repeatable now. If you're curious why AI varies at all, why AI gives different answers each time covers the mechanics.
Deliverables: what you hold when the provider leaves
Deliverables are the things you hold when the provider leaves, not the activities they perform. Write each one as a noun you could point at:
- The working workflow, built in accounts registered to your business (automation tool, AI model key, shared sheet)
- The prompt text and the price-guide lookup, in an editable document with a version number
- The tuning set and the acceptance test set, with the scored results of every test run
- A run-book of two pages or fewer: how to pause the workflow, how to edit the prompt, where the log is, what to do when it errors
- A handover session of about an hour, recorded, in which you pause, edit and restart it yourself
- Removal of the provider's access on a stated date, confirmed in writing
Item 1 matters more than it looks. A workflow built in the provider's own account is theirs in practice, whatever the contract says, and moving it later is a rebuild. The full list of what to collect before a provider leaves is in the AI consultant handover checklist.
One deliverable to question: if a proposal offers a custom GPT, note that OpenAI is retiring them. They stop running on 11 December 2026, and OpenAI's migration turns each one into a plugin. For shared business context in ChatGPT, a Project is the current equivalent, and a deliverable built on something being retired is a deliverable with a short life.
Acceptance criteria that allow for AI's misses
This is the step most scopes get wrong. "The AI accurately extracts enquiry details" can't be tested, because it doesn't say how accurate, on what, or measured by whom. Good criteria for AI work come in four kinds: a pass mark on held-back real cases, must-never rules with zero tolerance, correct handling of planted awkward cases, and practical limits on speed and cost. The furniture maker's set:
| Criterion | How it's measured | Pass mark |
|---|---|---|
| Extraction accuracy | 40 held-back enquiries; 5 key fields each checked against the owner's own reading | At least 190 of 200 fields correct |
| Price range | Compared with the range the owner would give for the same 40 | Overlaps the owner's range in at least 36 of 40 |
| Questions for missing details | Owner marks whether the draft asks for everything needed | At least 34 of 40 need no added question |
| Must-never rules | Any reply sent without approval; any price below the guide's minimum; any promised delivery date | Zero occurrences in all testing |
| Planted awkward cases | 6 cases: repair request, trade order, job application, abusive message, blank form, photo-only email | All 6 flagged needs_owner |
| Speed | Time from enquiry arriving to draft appearing | Under 5 minutes |
| Running cost | Subscriptions plus model usage at 40 enquiries a month | Under $60 a month |
Two rules make these numbers mean something. First, the acceptance set must be held back: if the prompt is tuned on the same enquiries it is tested on, it can score 40 out of 40 and still stumble on next month's post. Second, the test set must look like real life, awkward cases included in the proportion they actually occur. A test set of 40 easy enquiries passes anything.
The quick sum on the running-cost line: at 40 enquiries a month, even a workflow that uses seven Zapier tasks per enquiry needs 280 tasks, well inside the 750 included in Zapier's Professional plan at $29.99 a month on monthly billing, and model usage for 40 short enquiries costs cents. Under $60 a month leaves room for the photo-reading step if it is added later.
Acceptance criteria overlap with pilot success criteria but aren't the same thing: acceptance asks "was it built as agreed?", while a pilot asks "does it pay off in daily use?". The second question is covered in setting success criteria for an AI pilot that hold up, and a safe way to gather that evidence is piloting in shadow mode before customers see anything.
The out-of-scope list and each side's duties
The out-of-scope list is the cheapest insurance in the whole document. For the furniture maker:
- Out of scope: sending replies automatically; taking deposits; booking site visits; estimating dimensions from photos; social media messages; changes to the website form; moving enquiries into a CRM.
- Client provides: price guide v2 as one spreadsheet before build starts; 70 anonymised past enquiries (30 for tuning, 40 held back); accounts in the business's name; two hours of testing time within five working days of being asked.
- Provider provides: the build; all six deliverables listed earlier; fixes for defects against the criteria for 30 days after acceptance.
When someone asks mid-build for "just one more thing", this list is where the conversation starts. A good process for those requests is in how to handle change requests in AI projects.
How the acceptance test actually runs
Criteria without a process still end in argument. Settle these points in the scope:
- Who runs the test. The provider runs the 40 held-back cases through the final version while you watch, or you run them yourself; either way the results go into a shared sheet.
- How defects are graded. Critical: breaks a must-never rule, so fix and re-run all 40. Major: misses a pass mark, so fix and re-run the affected criterion. Minor: wording or formatting, so list for the warranty period.
- How many re-test rounds are included. Two is common; further rounds caused by new requirements are change requests.
- How long you have to test. Five working days is typical. Be wary of "deemed acceptance" wording that treats silence as sign-off; at least require a written reminder first.
- What happens after acceptance. A warranty period for defects against these criteria, and who re-runs the test (at what cost) if the model, the prompt or a connected app changes later.
Here is an excerpt from the furniture maker's first test round, scored in the shared sheet:
| Case | What happened | Grade |
|---|---|---|
| T07 | Dimensions given in inches and converted correctly, but the draft forgot to ask about stair access for delivery | Minor |
| T19 | Wood recorded as oak from the photo, though the customer wrote "walnut or oak, not sure" | Major: guessed a field the rules say must never be guessed |
| T26 | Price range below the guide's minimum for a 2.4-metre table | Critical: a must-never rule broken |
| T33 | Planted repair request flagged needs_owner | Pass |
Round one scored 186 of 200 fields (the pass mark is 190), 35 of 40 overlapping price ranges and one critical defect. The builder fixed the size-band lookup behind T26 and added the instruction "if the customer names more than one wood, record all of them and ask which", then re-ran all 40 cases, as the rules require after a critical defect. Round two: 194 of 200 fields, 38 of 40 ranges, no critical defects, and all six planted cases flagged. The project was accepted, with T07's missing question logged for the warranty period. Notice that nobody argued about whether it "worked": the sheet said so.
The furniture maker's one-page scope, assembled
PROJECT: Enquiry reader and reply drafter
OWNER (client side): [owner's name]
JOB: For every new enquiry, extract piece type, wood, finish and
dimensions, and draft a reply with a price range and questions for
missing details. The owner approves every reply before sending.
INPUTS: website form (fields + up to 3 photos); forwarded emails
with attachments; mixed units; ~1 in 8 not a new commission.
OUTPUTS: one row per enquiry in the shared sheet (fields listed in
Annex A) and a draft reply in the shared inbox. Nothing sent
automatically.
DELIVERABLES: workflow in client-owned accounts; prompt and lookup
document (versioned); tuning and test sets with scored results;
two-page run-book; recorded handover (~1 hour); provider access
removed on [date].
ACCEPTANCE: criteria and pass marks in Annex B, tested on 40
held-back enquiries plus 6 planted cases. Critical defects: re-run
all tests. Two re-test rounds included. Client testing window: 5
working days after notice.
OUT OF SCOPE: auto-sending, deposits, site-visit booking, dimension
estimates from photos, social messages, form changes, CRM migration.
CLIENT PROVIDES: price guide v2 (one sheet) before build; 70
anonymised past enquiries; accounts in client's name; testing time.
AFTER ACCEPTANCE: 30 days of defect fixes against Annex B. Re-test
after any model or app change: [who / how charged].
It fits on one page, yet it answers the questions that usually surface as disputes in week four.
Drafting the scope with an AI assistant, and what to strike out
An assistant can turn your rough notes into this structure quickly, provided you police the result. A workable prompt:
Here are my notes about an AI project for my business: [notes].
Draft a one-page scope with these sections: job, inputs, outputs,
deliverables, acceptance criteria, out of scope, client duties,
after acceptance. Every acceptance criterion must include how it is
measured and a numeric pass mark. Do not add any system, feature or
integration that is not in my notes; if something is unclear, list
it as a question at the end instead.
The first draft that came back, in this illustrative case, had two typical faults:
"Acceptance criteria: The system will achieve high accuracy in extracting enquiry details. Responses will be generated promptly. Integration with the accounting system will allow quotes to become invoices."
"High accuracy" and "promptly" have no numbers, despite the instruction, so they go back as "at least 190 of 200 fields" and "under 5 minutes". The accounting integration appeared from nowhere; it wasn't in the notes, and left in, it would have become an expensive line in every quote. Strike anything you didn't ask for, and treat the assistant's draft as a checklist of headings rather than a finished scope.
Scoping gaps that surface after sign-off
Three realistic failures, each traceable to one missing line.
No named source of truth. A sports equipment shop accepted a returns-question chatbot as delivered. Three weeks later it was quoting a 60-day returns window from an old PDF the builder had loaded, although the current policy on the website said 30 days. The scope never said which document the bot answers from or who updates it. The missing line: "Knowledge source: the returns page at [address], re-read nightly; owner: [name]."
Tested on the tuning set. A catering company's quote drafter scored 30 out of 30 at acceptance, because the test used the same 30 enquiries the prompt had been refined on. On new enquiries it mispriced canapé orders by guest count about one time in five. The missing line: "Acceptance uses 40 enquiries withheld from the provider until testing."
No baseline, so no verdict. An e-commerce homeware brand automated product-question replies with a clear build scope and a passed test, but nobody had timed the old process. Six months later, asked whether it was worth renewing, nobody could say. The missing line: "Baseline: current handling time per enquiry, measured over two weeks before build starts."
A scope with a named job, specified inputs and outputs, owned deliverables, numeric acceptance criteria, an out-of-scope list and a testing process is not bureaucracy. It is what lets you compare quotes fairly, accept work with confidence and, if it comes to it, walk away from a build that didn't meet what was agreed.
Scoping and acceptance testing: your questions
How many test cases do I need for acceptance testing?
Enough that one lucky or unlucky case can't swing the result. With 10 cases, a single case moves the score by 10 points, so aim for 40 to 60 real examples for a workflow that runs dozens of times a month. Include the awkward ones in the proportion they really occur, plus a handful of planted edge cases that must be routed to a person.
Should the builder see the acceptance test set?
No. Give the builder a separate set of real examples for tuning and keep the acceptance set back until testing. If the prompt is tuned on the same cases it is tested on, it can pass perfectly and then perform worse on new work. Sharing the test set after acceptance, to help with fixes, is fine.
What happens if the AI passes acceptance and then gets worse?
Write this into the scope before you sign. A warranty period covers defects against the agreed criteria for a set time, often 30 days. Beyond that, agree who re-runs the acceptance test when a model, prompt or connected app changes, and what it costs. Keeping the test set and the scoring sheet means anyone can repeat the check later.
Who should write the scope, me or the provider?
Draft it yourself first, even roughly, because you know the job, the exceptions and what a good result looks like. Then let the provider tighten it: they know what can be built and tested. A scope written only by the provider tends to describe what they plan to deliver rather than what you need to accept.
Further reads
- How to Run a Paid Discovery Phase Before a Full AI Project — Produce a scope from evidence when the job is still fuzzy.
- AI Consulting Contracts: 9 Clauses to Check Before You Sign — Put the scope inside a contract that protects you.
- Who Owns the AI Workflows a Consultant Builds for You? — Settle ownership of prompts, workflows and accounts.
- How to Write a Request for Proposal for an AI Project — Send the finished scope out for comparable quotes.
- How to Judge Whether Your AI Consultant Delivered Value — Check after go-live whether the project delivered.
- How to Stop Zapier and Make Automations Breaking Silently — Keep an accepted workflow from failing quietly later.
- How to Evaluate an AI Implementation Proposal or Quote — A 22-point checklist for any AI implementation quote, with the phrases to pin down, a scoring sheet and two quotes compared over three years.
- How to Choose an AI Consultant: 20 Questions to Ask First — Twenty questions to put to any AI consultant, what strong and weak answers sound like, and a scoring sheet filled in for a farm shop.
- Course, 1:1 Call, Audit, or Project: Which AI Help Do You Need? — What a course, a 1:1 call, an audit and a project each leave you with, five questions to find your starting point, and one deli's route through three.
- Fixed-Scope vs Hourly AI Consulting: Which Protects Your Budget? — How fixed-price and hourly AI consulting quotes shift risk, with a worked catering example, a cap clause and questions to ask before you sign.
- How to Negotiate an AI Consulting Quote Without Cutting Corners — The levers that lower an AI consulting quote safely, the lines never to cut, email wording, and a small charity taking a quarter off its quote.
- How to Hire an AI Automation Freelancer on Upwork or Fiverr — Post one workflow, screen proposals properly, run a paid test and set milestones, with the real Upwork and Fiverr client fees worked into the budget.
- How to Hire a HubSpot or CRM Consultant to Set Up AI Features — A garage's brief, three quotes and a go-live test script: how to find, question and pay a HubSpot or CRM consultant to switch on AI features properly.
- AI Consultant Red Flags: 12 Warning Signs to Walk Away From — Twelve warning signs when hiring an AI consultant, each with an example, the question to ask, its innocent version and a scored two-proposal comparison.
- What a Typical AI Consulting Engagement Looks Like, Week by Week — A six-week AI consulting engagement explained stage by stage: what the consultant does, what you do, what you should have each Friday, and why weeks slip.
- How to Rescue a Stalled AI Project or Exit It Cleanly — Freeze spending, take stock of what was built, then re-scope to one live piece in 30 days or shut it down without losing your data or logins.
- AI Tools and AI Development: The Complete 2026 Guide — the AI hub, including every tutorial in the AI-for-business series.
Sources: OpenAI help article on custom GPT retirement and migration; Anthropic and OpenAI documentation on temperature and structured outputs; Zapier pricing page.