Small Business AI Implementation Checklist: Before, During, After

Coding Liquids tutorial cover featuring Sagnik Bhattacharya for Small Business AI Implementation Checklist: Before, During, After.
Coding Liquids tutorial cover featuring Sagnik Bhattacharya for Small Business AI Implementation Checklist: Before, During, After.

It should cover three stages. Before: a named job, a baseline, an owner, business accounts, data rules and written success criteria. During: a test set, a person checking anything customers see, an error log and a weekly review. After: results against the baseline, a runbook, monitoring, renewal dates and a way to switch it off.

That's 30 checks in all, and each one exists because skipping it causes a specific, predictable problem later. Use the list to run your own implementation, or hand it to a supplier as the definition of "done". Either way, tick an item only when you can point to the evidence, not when it feels handled. The timing of the stages is covered in the 90-day AI implementation roadmap; this list is the detail inside it.

Follow me on Instagram@sagnikteaches

Before you start: 12 checks

#CheckWhy it mattersEvidence to tick it
B1One named job, with its weekly hoursA vague goal can't succeed or failA sentence: job, who does it, hours a week
B2A baseline over two normal weeksWithout it, "it saves time" is a feelingDated figures: volume, minutes per item, redo rate
B3A named owner with diary time, and a sponsor who decidesProjects without an owner stall at the first snagTwo to four hours a week blocked in the owner's diary
B4Business accounts, admin held by the business, two-step sign-in onEverything later depends on access you controlYou can open each admin page yourself today
B5A data map for this jobYou can't protect data you haven't locatedA list of the personal data involved and where it's stored
B6A one-page AI rule, sharedStaff need to know what never goes inThe rule, and a note of when staff saw it
B7Vendor terms checkedTraining, retention and access defaults differ by planNotes on training default, retention period and plan tier
B8Notice and consent where neededCustomers may need telling, or askingThe wording you'll use, and your adviser's view if data is sensitive
B9Success criteria and a stop rule, signedCriteria written afterwards always passA dated, signed one-page criteria sheet
B10Three months of budget, including hoursTrials judged in week two are wastedA monthly figure for tools and a figure for time
B11An exit planSuppliers change terms, prices and productsWhere your data and instructions would go, and how to switch off
B12Staff told, and the people doing the job askedTools imposed on people get ignoredA dated announcement, and notes from the conversations

Three of these deserve more than a line.

Connect on LinkedInSagnik Bhattacharya

B7, vendor terms. Business plans such as ChatGPT Business, Claude Team, Microsoft 365 Copilot and Gemini in Workspace don't train on business content by default, while consumer plans rely on each user switching off the model-training setting. Retention periods and who at the vendor can see your data vary too. Read the privacy and data terms for the exact plan you'll use; what to check in an AI tool's privacy terms lists the clauses that matter.

Subscribe on YouTube@codingliquids

B8, notice and consent. If the AI will talk to customers directly and you sell to customers in the EU, the EU AI Act's transparency duty has applied since 2 August 2026: people must be told they're dealing with AI. If you'll record calls or process health or children's information, agree the wording and the lawful basis with your data-protection adviser before go-live, not after a complaint.

AI can draft the notice, but check its claims line by line. Here's a prompt used by the tutoring agency in the filled-in example below:

Draft three sentences for our termly letter to parents explaining
that weekly progress updates are now drafted with AI from tutors'
notes and checked by a coordinator before sending. Plain and
honest, no marketing language.

An illustrative draft:

"From this term, your child's weekly update will be drafted with the help of AI, using their tutor's session notes. Every update is read and approved by a coordinator before it reaches you. Your child's information is never stored and is completely secure."

What you'd fix: the first two sentences are accurate. The third is a promise no business can make: the notes are stored, in the agency's own account, and "completely secure" isn't a claim to put in writing. The agency replaced it with "Notes stay in our own business account, and our AI plan doesn't use them to train its models." Assistants reach for reassuring absolutes; notices need true, checkable statements.

B11, the exit plan. This is the check owners find oddest to write before starting. Real suppliers show why it matters: the AI calendar tool Clockwise shut down on 27 March 2026 and deleted user data rather than transferring it. If your instructions, examples and data exist only inside one tool, a closure notice becomes a rebuild. What to check in case an AI vendor shuts down covers the questions to ask before you commit.

During the build and pilot: 9 checks

#CheckWhy it mattersEvidence to tick it
D1A test set of 10 to 20 real, anonymised cases, including awkward onesYou'll re-run it after every change and every vendor updateA saved file of cases, with the right answer for each
D2Instructions and reference files kept outside the tool tooTools change; your work shouldn't vanish with themA copy in your own documents, dated
D3Least access: connections only to what the job needsAn automation that can read everything can leak everythingA list of each connection and what it can reach
D4A person checks everything customers seeAI states guesses as factsThe checking step written into the process, with a name
D5An error and grade logPatterns in mistakes tell you what to fixA log with date, case type, grade and note
D6One change at a time, loggedSeveral changes at once hide which one helpedA change log with dates
D7A weekly 15-minute reviewProblems surface while they're smallA recurring diary slot, and brief notes
D8An exceptions routeSome cases must always go to a personA written list of case types AI never handles, and who gets them
D9Failure alerts on any automationAutomations fail silently more often than loudlyA test failure that reached a named person

D9 catches people out because each platform fails differently. Zapier auto-pauses a Zap only when 95% of its runs error over seven days, and when an error handler runs, Zapier sends no error email, so the handler itself must alert someone. Make's error handlers (Skip, Retry, Resume, Commit and Rollback) each behave differently, and choosing Skip for a failure you needed to hear about hides it. Power Automate turns off any flow that has failed continuously for 14 days. The only real test is to break it on purpose, for example with a malformed test case, and confirm the right person hears about it.

D1 is the check that pays off longest. Keep the test set, with the correct answer for each case, somewhere you'll find it in a year. When the vendor updates its model or you change the instructions, re-running the set takes 20 minutes and tells you immediately whether anything drifted.

After go-live: 9 checks

#CheckWhy it mattersEvidence to tick it
A1Results compared with the baseline, same measuresOnly like-for-like numbers prove a changeBefore-and-after figures on one page
A2Full cost countedUpkeep hours often exceed the subscriptionTools, set-up hours and weekly upkeep, totalled
A3A runbook, followed once by someone elseKnowledge in one head leaves with that personA one-page runbook, and the name of who tested it
A4Owner and deputy named, monthly check in the diaryUnowned systems decay quietlyNames in the runbook; a recurring diary slot
A5Usage monitoredFalling usage is the earliest warning signA monthly figure from the tool's report or your own tally
A6Renewal dates and notice periods diarisedAnnual plans renew whether or not anyone uses themReminders 60 and 30 days before each renewal
A7A leaver processEx-staff keeping access is a common, avoidable riskLeaver steps listing each AI account and shared project
A8Re-testing after vendor changesModels, features and defaults change without your sayThe test set re-run, with the date and result
A9An annual keep, fix or stop reviewWhat was worth it last year may not be nowA dated decision, with the numbers behind it

A2 often changes the verdict, so do the sum properly. Here are the tutoring agency's illustrative figures: three ChatGPT Business Standard seats on monthly billing cost $75 a month. Keeping the facts and examples current, re-running the test set and handling exceptions took the senior coordinator about an hour a week, roughly $100 a month at an illustrative $25 an hour. The saving was about 1.5 minutes on each of 140 updates a week, around 3.5 hours a week or about $375 a month at the same rate. Net, about $200 a month in the agency's favour, but note that the upkeep cost more than the subscription. Leave the hours out and the project looks twice as good as it is.

For A5, use whatever the tool reports. The Microsoft 365 admin centre's Copilot usage report shows active users over 7, 28, 90 or 180 days, and automation platforms show run history. For a drafting job with no report, a two-column tally of cases handled with and without AI is enough.

A8 matters more each year, because vendors retire things on their own timetable. OpenAI's custom GPTs stop running on 11 December 2026, and Excel's =COPILOT() worksheet function was retired on 14 September 2026. Vendors also change defaults: SimplePractice's Note Taker began opting new users in to keeping de-identified transcripts from 16 June 2026. A quarterly look at your vendors' change notices, plus your test set, catches these before customers do.

The tutoring agency's checklist, filled in

Here's how the list looks in use. An illustrative tutoring agency, with the owner, three coordinators and about 70 tutors, introduced AI drafting of weekly progress updates to parents from tutors' session notes. Coordinators had spent about 6 hours a week writing them. An extract from its checklist at go-live, with honest gaps left open:

[x] B1  Weekly parent updates from session notes; 3 coordinators;
        ~6 hrs/week.
[x] B2  Baseline: 140 updates/week, median 2.5 min each (2 weeks).
[x] B5  Data: session notes hold pupils' first names, subjects and
        progress; stored in the agency's Drive. Parent emails in
        the booking system. No surnames or dates of birth in notes.
[x] B7  Business plan; no training on business content by default;
        retention period noted from vendor terms.
[x] B8  Parents told in the term letter that updates are drafted
        with AI and checked by a coordinator before sending.
[ ] B11 Exit plan drafted but instructions not yet copied outside
        the tool. OWNER: senior coordinator, by Friday.
[x] D1  Test set: 15 anonymised notes incl. 3 with two pupils of
        the same first name.
[x] D4  Coordinator reads every update before it's sent.
[x] D8  Never drafted by AI: anything about wellbeing, behaviour
        or safeguarding. Goes to the owner.
[ ] A3  Runbook written; not yet followed by a second person.
[x] A6  Renewal in 11 months; reminders set at 60 and 30 days.
[x] A7  Leaver steps updated: remove from AI workspace and shared
        project on the last day.

Two items are unticked, each with a named owner and a date, which is how the checklist is meant to be used. A list where everything is ticked on day one usually means some ticks are optimistic. Note D1's awkward cases: two pupils with the same first name was exactly the situation most likely to produce an update about the wrong child, so it went into the test set on purpose.

Checks people skip, and how it showed up

  • A6, renewals. An illustrative picture framer tried an AI image tool for product photos, stopped using it after a month and forgot it was on an annual plan. The renewal charge eleven months later was the first reminder it existed.
  • D1 and A8, the test set. A pet shop's AI-drafted replies gradually became longer and more formal after a model update. With no saved test cases to compare against, it took weeks and a customer's comment before anyone noticed.
  • A7, leavers. A language school found a teacher who had left three months earlier still had access to the shared AI workspace, including student work pasted for feedback. Removing them took two minutes; noticing took a term.
  • D9, alerts. An optician's reminder automation stopped sending when a password changed. Nobody was told, because nobody had tested what happens on failure, and the gap showed up as a rise in missed appointments.

Each of these took minutes to prevent and weeks to discover. That's the pattern the whole checklist is built around: the expensive failures in small-business AI are rarely the AI getting something wrong once; they're the dull checks nobody did, failing quietly for months. If you want those risks tracked in one place, a simple AI risk register gives each one an owner and a review date.

Using the checklist with an outside supplier

If a consultant, freelancer or agency is doing the work, the checklist doubles as your definition of "done". Split it before the contract is signed:

  • Yours, whoever builds it: B1 to B3, B5, B6, B8, B10, B12, A4 and A9. These are decisions and duties that can't be outsourced.
  • Theirs, as deliverables: D1 to D3, D5, D8, D9, A3 and the technical half of B4 and B11: the test set, the stored instructions, the access list, the failure alerts, the runbook and the exit route.
  • Shared: B9, where you set the criteria and they confirm they can be measured, and A1, where they supply the numbers and you judge them.

Then make payment for the final stage depend on their items being evidenced, not just demonstrated. The AI consultant handover checklist covers what you should hold before a supplier leaves, which overlaps heavily with the "after" stage here.

Scaling the list to the size of the job

Thirty checks is right for a workflow that touches customers or runs automatically. For smaller uses, a subset does the job:

  • One person drafting internal documents with a business AI account: B4, B6, B7, D4 and A6. Five checks, about an hour.
  • A team using AI drafting with a person checking everything: add B1, B2, B3, B9, D1, D5, D7, A1, A5 and A7.
  • Anything automated or customer-facing: all 30, plus B8 read carefully.

Whichever size applies, keep the evidence with the checklist in one folder: the baseline figures, the signed criteria, the test set, the runbook and the renewal reminders. A year from now, when you're deciding whether to keep, change or replace the tool, that folder will answer most of the questions in minutes.

Further reads

Sources: Zapier help centre (auto-pause and error handlers); Make help centre (error handlers); Microsoft Learn (Power Automate flow suspension; Copilot usage report); OpenAI help centre (custom GPT retirement); Microsoft support (Excel COPILOT function retirement); EU AI Act Articles 4 and 50; vendor notices on Clockwise and SimplePractice Note Taker.

Want this checklist applied to your AI project?

On a 1:1 call we'll go through the checks that matter for your specific job, mark which you've already covered, and turn the gaps into a short plan with owners and dates.

Book a 1:1 call with me