Scope Creep in AI Projects: How to Handle Change Requests

Coding Liquids tutorial cover featuring Sagnik Bhattacharya for Scope Creep in AI Projects: How to Handle Change Requests.
Coding Liquids tutorial cover featuring Sagnik Bhattacharya for Scope Creep in AI Projects: How to Handle Change Requests.

Agree a written baseline before work starts (the tasks, data sources, volumes and an accuracy target with a test set), then put every new idea through a one-page change request that states the extra hours, cost and delay. Approve it, move it to a phase-two list, or decline it, in writing, before anyone builds it.

AI projects creep through routes ordinary software projects don't: "can it also read the attachments?", a new channel such as WhatsApp, an accuracy target that drifts from "mostly right" to "never wrong", and changes forced by the platforms themselves. OpenAI's custom GPTs stop running on 11 December 2026, and Excel's =COPILOT() function was retired on 14 September 2026. Deciding up front who pays when a vendor shifts the ground under a project prevents the bitterest arguments.

Follow me on Instagram@sagnikteaches

Why AI projects creep in ways ordinary IT projects don't

Scope creep is the slow growth of a project beyond what was agreed, one reasonable-sounding request at a time. Every project suffers it. AI projects suffer some extra kinds:

Connect on LinkedInSagnik Bhattacharya
  • New inputs. A workflow built for email text is asked to read PDFs, photos or voice notes. Each new input type brings new failure cases and extra model costs.
  • New channels. Adding WhatsApp or web chat means new message formats, new rules and sometimes new fees. Meta's WhatsApp Business Platform, for instance, starts charging for free-form service replies from 1 October 2026 once a number passes its first 1,000 service messages a month.
  • Accuracy drift. The target quietly rises. Getting from 90% to 95% correct is usually a matter of rules and examples; getting from 95% to 99% can take far more testing than the rest of the build.
  • Edge cases found in testing. Night-time emergencies, enquiries in the wrong inbox, customers who reply to a year-old thread. Each is small; together they double the work.
  • Data surprises. The price list turns out to be three spreadsheets that disagree.
  • Platform changes. A model is retired, a feature renamed or a price structure changed mid-project, as when ChatGPT agent became ChatGPT Work on 9 July 2026.

None of these is anyone's fault. They become disputes only when nobody wrote down what was in and what was out.

Subscribe on YouTube@codingliquids

Write a baseline you can test against (half a day)

A baseline is the agreed description of what will be built, precise enough that both sides can tell whether a request is inside it. For AI work it needs two things most scope documents leave out: volumes and a measurable accuracy target with a named test set. How to scope an AI project covers deliverables and acceptance criteria in full; here's the baseline an illustrative property maintenance firm agreed for an automation that turns emailed enquiries into draft job sheets:

Project: AI job booking from email enquiries
In scope:
- Read enquiries arriving in the jobs inbox (email body text only)
- Classify trade (plumbing, electrical, roofing, general) and urgency
- Create a draft job sheet in the job-management system for a
  scheduler to approve
- Draft a reply to the customer, sent only after approval
Out of scope: attachments, phone and WhatsApp enquiries, pricing,
invoicing, changes to the job-management system itself
Volume: about 400 enquiries a month
Accuracy target: trade and urgency correct on 90% of a fixed test set
of 100 past enquiries, chosen by the firm before the build starts
Acceptance: five working days of live drafts in which schedulers edit
no more than 10% of job sheets for longer than a minute
Effort: 80 hours; contingency 12 hours (15%), used only by agreement

The out-of-scope line does more work than the in-scope list. Every item on it is something the firm had mentioned in early conversations. Writing it down made clear they were ideas for later, not promises.

The contingency matters too. A reserve of 10-20% of the estimated effort, spent only with both sides' agreement, absorbs the small discoveries that every build throws up without a formal request each time.

A one-page change request both sides fill in

Anyone can raise a change: the owner, a scheduler, the builder. What matters is that it goes on the same form and nothing gets built until it's decided. The form:

CHANGE REQUEST No. __          Raised by: ______   Date: ______
1. What is being asked for (one or two sentences)
2. Why: the problem it solves or the value it adds
3. Is it inside the baseline? (yes / no / unclear, with the line)
4. Extra effort: build ___ h, testing ___ h, documentation ___ h
5. Running costs per month (model tokens, automation tasks,
   messaging fees, new licences)
6. Effect on the delivery date and the acceptance test
7. Risks and what it displaces
8. Decision: approve (contingency / extra budget) | phase two | decline
   Signed: client ______  builder ______

Line 3 settles half of all disputes on its own. If the request is inside the baseline, it's the builder's job to deliver at no extra cost. If it's outside, it's a change, however small. "Unclear" means the baseline needs a sentence added, which both sides agree on before going further.

Size the change before anyone says yes (about an hour each)

Sizing means estimating four things, in hours and money rather than adjectives:

  1. Build effort for the new behaviour, including the automation steps and prompts.
  2. Test effort, often the biggest line for AI changes, because new inputs need new test cases and a fresh run of the old ones.
  3. Running costs each month from now on: extra model tokens, automation tasks, messaging fees or seats.
  4. Knock-on effects: a later delivery date, a harder acceptance test, or work it pushes aside.

Express effort in hours and let each side multiply by the agreed rate; it keeps the conversation about the work rather than about the rate. A general assistant can produce a useful first draft of the sizing, as long as a person checks it. The prompt the firm's builder used:

You are helping size a change request for a small automation project.
Baseline: [paste baseline]. Requested change: [paste request].
List: (a) whether the change is inside the baseline, quoting the line;
(b) build, test and documentation tasks, each with an hour estimate and
the assumption behind it; (c) new monthly running costs, naming what
drives them; (d) effects on the acceptance test. Mark every estimate as
an assumption to confirm. Do not invent prices you weren't given.

An illustrative extract for the request "read photo attachments to judge urgency":

(a) Outside baseline: "email body text only".
(b) Build: fetch attachments, send images to model, merge result: 5 h
    Test: 30 new test cases with photos, rerun 100 existing: 4 h
    Docs: 1 h                                    Total: 10 h
(c) Model tokens for images; roughly 150 photos a month (assumed).
(d) Test set grows from 100 to 130; target unchanged at 90%.

The builder fixed two things. The photo volume assumption was checked against a month of real enquiries: 140 photos, close enough. And the running cost needed a number, which the model had rightly declined to invent. On Anthropic's published figures a one-megapixel photo is about 1,300 input tokens, roughly a quarter of a cent each on Claude Sonnet 5, so 140 photos add about 36 cents a month. Small enough that the decision rested on the ten hours, not the running cost.

Decide, log and park the rest in phase two

Every request ends in one of three outcomes, recorded in a change log that both sides can see:

  • Approve, funded from contingency or with an agreed addition to budget and timeline.
  • Phase two: a good idea, but not now. Keeping a visible phase-two list lets people say yes to an idea without saying yes to building it this month.
  • Decline, with the reason, so it doesn't return next week in different words.

Declining or deferring is easier with wording ready to hand. An email the firm's builder used for a deferral:

Subject: Change request 1 (WhatsApp enquiries): proposed for phase two

Thanks for raising this; it's a sensible addition. Sized at 18 hours,
plus WhatsApp's own messaging charges from 1 October, it would push
acceptance back by about two weeks. I suggest we finish and accept the
email version first, then start WhatsApp as phase two with its own
test set. If you'd rather do it now, I'll send a revised timeline.

Defect or change? Judging AI errors against the baseline

AI output is probabilistic: even a well-built workflow gets some cases wrong. That creates a dispute ordinary software rarely has. When the client finds a mistake, is it a defect the builder must fix for free, or a request for better-than-agreed performance, which is a change? The baseline's accuracy target answers it, provided you measure against the agreed test set rather than against the latest annoying example.

What the client foundJudged againstVerdict
Roofing jobs classed as "general" in 14 of 100 test enquiriesTarget: 90% correct on the test set; actual 86%Defect: below the agreed target, fixed at the builder's cost
One urgent leak marked routine in a live demoTest set still at 93%Not a defect on its own; add the case to the test set and watch the rate
Photos of damage ignored"Email body text only"Change: outside the baseline
A customer replying to a year-old thread creates a duplicate jobBaseline silent on old threadsUnclear: agree a sentence for the baseline, then decide
Draft replies sound stiffNo tone requirement in the baselineChange, though often a cheap one: a tone guide and examples

Two habits keep this fair. First, freeze the test set at the start and add new cases only by agreement, so the target can't move silently. Second, when a live mistake worries you, add it to a separate "watch" list and review the pattern weekly. One wrong answer is an anecdote; five of the same kind in a fortnight is a finding, and possibly a defect.

A property maintenance firm's job-booking build, change by change

Here is how the illustrative firm's 80-hour build went. Three change requests arrived in the first four weeks:

No.RequestRaised byExtra hoursRunning costDecision
1Handle WhatsApp enquiries tooOwner18WhatsApp platform fees above the free allowance; about 360 service messages a month, inside the first 1,000Phase two
2Read photo attachments to judge urgencySenior scheduler10About 36 cents a month in model tokensApproved from contingency
3Raise the accuracy target from 90% to 98%Owner24 or more, uncertainNoneDeclined as worded; replaced by a 95% target for trade only

The third request shows the most common AI-specific creep. After a scheduler found two misrouted jobs in a demo, the owner asked for 98% accuracy. The builder's sizing showed why that was expensive: urgency is partly a judgement call that the firm's own schedulers disagreed on in about one enquiry in ten, so no model could reach 98% against their labels. The agreed replacement kept the schedulers' approval step for urgency, which was already in the baseline, and raised the target for trade classification alone to 95%, costing six hours of extra test cases from the contingency's remaining two plus four approved extra hours.

The totals: 80 baseline hours, 10 contingency hours for photos, 6 hours for the trade target (2 from contingency, 4 added by agreement), so 96 hours delivered against 92 planned including contingency, with the WhatsApp work sitting on the phase-two list at 18 hours. No invoice was disputed, because every hour traced back to a signed line.

Phase two started a month after acceptance, once the email version had run cleanly. The WhatsApp request became a small baseline of its own: WhatsApp messages to the firm's business number, the same trade and urgency targets, a fresh test set of 50 past WhatsApp conversations, and a line naming Meta's messaging fees as the firm's cost rather than the builder's. Scoped separately, it could be judged on its own results, and the first project's acceptance didn't wait for it. Plenty of phase-two items elsewhere never get built at all, because the problem they were meant to solve goes away while they wait. That's the list doing its job, not failing at it.

Compare what happens without the process. An estate agency's sales manager asked, on a call, whether an enquiry-reply assistant "could also do lettings". The builder said "should be fine" and added it. Lettings enquiries needed different rules, a second reference document and forty new test cases; the build overran by three weeks, and the extra invoice arrived with nothing either side had signed. Both had acted in good faith, and both felt cheated. A one-page request would have cost ten minutes.

A fifteen-minute weekly change review

Requests that arrive by phone, chat and corridor conversation are how creep slips past a process. A short weekly review catches them. The firm's version, every Friday:

  1. New requests: anything raised during the week goes on a form, whoever raised it and however small. The builder brings the ones mentioned in passing.
  2. Sizing due: requests awaiting an estimate get a date for one.
  3. Decisions: one named decision-maker on each side approves, defers or declines. For a small firm that's usually the owner and the lead builder; nobody else can approve a change.
  4. Numbers: hours used against the baseline and contingency, and the latest test-set score.

That last item doubles as an early warning. Creep shows up in the numbers before anyone raises a formal request. Three signs to watch:

  • Hours running ahead of results. As a rough guide, once half the hours are used, a good part of the acceptance test should already pass. At 60% of hours with the test score still well short of target, something unplanned is eating time.
  • A growing test set. If the 100-case set has quietly become 180, someone has been adding requirements through the back door.
  • "Quick questions" from the builder about situations the baseline never mentioned. Each one is either an edge case worth a sentence in the baseline or the first sign of a change request.

At the firm, the Friday review in week three spotted that the test set had grown to 118 cases, most of them night-time emergency enquiries a scheduler had added. That became change request 3's conversation about accuracy, a week earlier than it would otherwise have surfaced.

Contract wording that keeps change requests civil

Most of this belongs in the contract before the project starts, so nobody is inventing the rules mid-argument. The clauses to check in an AI consulting contract cover the full list, and fixed-scope versus hourly consulting explains how the pricing model changes the stakes. Two clauses matter most for creep. Example wording to adapt with your own adviser:

Change control. Work outside the Baseline will be carried out only
under a Change Request signed by both parties, stating effort, running
costs and the effect on timeline and acceptance. Contingency of [15]%
of the estimated effort may be used for changes by mutual agreement.

Platform changes. If a third-party product, model or feature used in
the Deliverables is retired or materially changed during the project,
the parties will agree a Change Request for the necessary rework. The
Builder will notify the Client within [5] working days of becoming
aware of the change.

The platform clause stops a common stand-off, where the builder says a vendor's change isn't their fault and the client says a half-working system isn't theirs. Sharing the cost of a change neither side caused is usually the fair answer, but agree it before it happens.

When the creep comes from the builder's side

Creep isn't always the client's doing. Builders, whether freelancers, agencies or consultants, can expand a project too, and the same process protects you. Watch for:

  • Gold-plating: features nobody asked for, such as a dashboard, billed as necessary.
  • "Discovered complexity" with no evidence. Ask for the specific examples that turned out harder than expected.
  • Open-ended research: hours spent evaluating tools after the tools were agreed.
  • Changes made without a request, then presented as done.

Apply the same rule in both directions: no request, no build, no bill. These checks apply to anyone you hire for this work, me included. If a project has already drifted a long way from its baseline, rescuing a stalled AI project or exiting cleanly covers the recovery options, and a paid discovery phase is the best way to make the next baseline realistic from the start.

Further reads

Sources: OpenAI help centre on custom GPT retirement; Microsoft Support on the retired =COPILOT() function; Meta's WhatsApp Business Platform pricing changes from October 2026; Anthropic documentation on image token costs. Checked September 2026.

Is your AI project growing faster than its budget?

On a 1:1 call we'll compare what's being built with what was agreed, size the open change requests, and decide what belongs in this phase and what can wait.

Book a 1:1 call with me