It should cover three stages. Before: a named job, a baseline, an owner, business accounts, data rules and written success criteria. During: a test set, a person checking anything customers see, an error log and a weekly review. After: results against the baseline, a runbook, monitoring, renewal dates and a way to switch it off.
That's 30 checks in all, and each one exists because skipping it causes a specific, predictable problem later. Use the list to run your own implementation, or hand it to a supplier as the definition of "done". Either way, tick an item only when you can point to the evidence, not when it feels handled. The timing of the stages is covered in the 90-day AI implementation roadmap; this list is the detail inside it.
Before you start: 12 checks
| # | Check | Why it matters | Evidence to tick it |
|---|---|---|---|
| B1 | One named job, with its weekly hours | A vague goal can't succeed or fail | A sentence: job, who does it, hours a week |
| B2 | A baseline over two normal weeks | Without it, "it saves time" is a feeling | Dated figures: volume, minutes per item, redo rate |
| B3 | A named owner with diary time, and a sponsor who decides | Projects without an owner stall at the first snag | Two to four hours a week blocked in the owner's diary |
| B4 | Business accounts, admin held by the business, two-step sign-in on | Everything later depends on access you control | You can open each admin page yourself today |
| B5 | A data map for this job | You can't protect data you haven't located | A list of the personal data involved and where it's stored |
| B6 | A one-page AI rule, shared | Staff need to know what never goes in | The rule, and a note of when staff saw it |
| B7 | Vendor terms checked | Training, retention and access defaults differ by plan | Notes on training default, retention period and plan tier |
| B8 | Notice and consent where needed | Customers may need telling, or asking | The wording you'll use, and your adviser's view if data is sensitive |
| B9 | Success criteria and a stop rule, signed | Criteria written afterwards always pass | A dated, signed one-page criteria sheet |
| B10 | Three months of budget, including hours | Trials judged in week two are wasted | A monthly figure for tools and a figure for time |
| B11 | An exit plan | Suppliers change terms, prices and products | Where your data and instructions would go, and how to switch off |
| B12 | Staff told, and the people doing the job asked | Tools imposed on people get ignored | A dated announcement, and notes from the conversations |
Three of these deserve more than a line.
B7, vendor terms. Business plans such as ChatGPT Business, Claude Team, Microsoft 365 Copilot and Gemini in Workspace don't train on business content by default, while consumer plans rely on each user switching off the model-training setting. Retention periods and who at the vendor can see your data vary too. Read the privacy and data terms for the exact plan you'll use; what to check in an AI tool's privacy terms lists the clauses that matter.
B8, notice and consent. If the AI will talk to customers directly and you sell to customers in the EU, the EU AI Act's transparency duty has applied since 2 August 2026: people must be told they're dealing with AI. If you'll record calls or process health or children's information, agree the wording and the lawful basis with your data-protection adviser before go-live, not after a complaint.
AI can draft the notice, but check its claims line by line. Here's a prompt used by the tutoring agency in the filled-in example below:
Draft three sentences for our termly letter to parents explaining
that weekly progress updates are now drafted with AI from tutors'
notes and checked by a coordinator before sending. Plain and
honest, no marketing language.
An illustrative draft:
"From this term, your child's weekly update will be drafted with the help of AI, using their tutor's session notes. Every update is read and approved by a coordinator before it reaches you. Your child's information is never stored and is completely secure."
What you'd fix: the first two sentences are accurate. The third is a promise no business can make: the notes are stored, in the agency's own account, and "completely secure" isn't a claim to put in writing. The agency replaced it with "Notes stay in our own business account, and our AI plan doesn't use them to train its models." Assistants reach for reassuring absolutes; notices need true, checkable statements.
B11, the exit plan. This is the check owners find oddest to write before starting. Real suppliers show why it matters: the AI calendar tool Clockwise shut down on 27 March 2026 and deleted user data rather than transferring it. If your instructions, examples and data exist only inside one tool, a closure notice becomes a rebuild. What to check in case an AI vendor shuts down covers the questions to ask before you commit.
During the build and pilot: 9 checks
| # | Check | Why it matters | Evidence to tick it |
|---|---|---|---|
| D1 | A test set of 10 to 20 real, anonymised cases, including awkward ones | You'll re-run it after every change and every vendor update | A saved file of cases, with the right answer for each |
| D2 | Instructions and reference files kept outside the tool too | Tools change; your work shouldn't vanish with them | A copy in your own documents, dated |
| D3 | Least access: connections only to what the job needs | An automation that can read everything can leak everything | A list of each connection and what it can reach |
| D4 | A person checks everything customers see | AI states guesses as facts | The checking step written into the process, with a name |
| D5 | An error and grade log | Patterns in mistakes tell you what to fix | A log with date, case type, grade and note |
| D6 | One change at a time, logged | Several changes at once hide which one helped | A change log with dates |
| D7 | A weekly 15-minute review | Problems surface while they're small | A recurring diary slot, and brief notes |
| D8 | An exceptions route | Some cases must always go to a person | A written list of case types AI never handles, and who gets them |
| D9 | Failure alerts on any automation | Automations fail silently more often than loudly | A test failure that reached a named person |
D9 catches people out because each platform fails differently. Zapier auto-pauses a Zap only when 95% of its runs error over seven days, and when an error handler runs, Zapier sends no error email, so the handler itself must alert someone. Make's error handlers (Skip, Retry, Resume, Commit and Rollback) each behave differently, and choosing Skip for a failure you needed to hear about hides it. Power Automate turns off any flow that has failed continuously for 14 days. The only real test is to break it on purpose, for example with a malformed test case, and confirm the right person hears about it.
D1 is the check that pays off longest. Keep the test set, with the correct answer for each case, somewhere you'll find it in a year. When the vendor updates its model or you change the instructions, re-running the set takes 20 minutes and tells you immediately whether anything drifted.
After go-live: 9 checks
| # | Check | Why it matters | Evidence to tick it |
|---|---|---|---|
| A1 | Results compared with the baseline, same measures | Only like-for-like numbers prove a change | Before-and-after figures on one page |
| A2 | Full cost counted | Upkeep hours often exceed the subscription | Tools, set-up hours and weekly upkeep, totalled |
| A3 | A runbook, followed once by someone else | Knowledge in one head leaves with that person | A one-page runbook, and the name of who tested it |
| A4 | Owner and deputy named, monthly check in the diary | Unowned systems decay quietly | Names in the runbook; a recurring diary slot |
| A5 | Usage monitored | Falling usage is the earliest warning sign | A monthly figure from the tool's report or your own tally |
| A6 | Renewal dates and notice periods diarised | Annual plans renew whether or not anyone uses them | Reminders 60 and 30 days before each renewal |
| A7 | A leaver process | Ex-staff keeping access is a common, avoidable risk | Leaver steps listing each AI account and shared project |
| A8 | Re-testing after vendor changes | Models, features and defaults change without your say | The test set re-run, with the date and result |
| A9 | An annual keep, fix or stop review | What was worth it last year may not be now | A dated decision, with the numbers behind it |
A2 often changes the verdict, so do the sum properly. Here are the tutoring agency's illustrative figures: three ChatGPT Business Standard seats on monthly billing cost $75 a month. Keeping the facts and examples current, re-running the test set and handling exceptions took the senior coordinator about an hour a week, roughly $100 a month at an illustrative $25 an hour. The saving was about 1.5 minutes on each of 140 updates a week, around 3.5 hours a week or about $375 a month at the same rate. Net, about $200 a month in the agency's favour, but note that the upkeep cost more than the subscription. Leave the hours out and the project looks twice as good as it is.
For A5, use whatever the tool reports. The Microsoft 365 admin centre's Copilot usage report shows active users over 7, 28, 90 or 180 days, and automation platforms show run history. For a drafting job with no report, a two-column tally of cases handled with and without AI is enough.
A8 matters more each year, because vendors retire things on their own timetable. OpenAI's custom GPTs stop running on 11 December 2026, and Excel's =COPILOT() worksheet function was retired on 14 September 2026. Vendors also change defaults: SimplePractice's Note Taker began opting new users in to keeping de-identified transcripts from 16 June 2026. A quarterly look at your vendors' change notices, plus your test set, catches these before customers do.
The tutoring agency's checklist, filled in
Here's how the list looks in use. An illustrative tutoring agency, with the owner, three coordinators and about 70 tutors, introduced AI drafting of weekly progress updates to parents from tutors' session notes. Coordinators had spent about 6 hours a week writing them. An extract from its checklist at go-live, with honest gaps left open:
[x] B1 Weekly parent updates from session notes; 3 coordinators;
~6 hrs/week.
[x] B2 Baseline: 140 updates/week, median 2.5 min each (2 weeks).
[x] B5 Data: session notes hold pupils' first names, subjects and
progress; stored in the agency's Drive. Parent emails in
the booking system. No surnames or dates of birth in notes.
[x] B7 Business plan; no training on business content by default;
retention period noted from vendor terms.
[x] B8 Parents told in the term letter that updates are drafted
with AI and checked by a coordinator before sending.
[ ] B11 Exit plan drafted but instructions not yet copied outside
the tool. OWNER: senior coordinator, by Friday.
[x] D1 Test set: 15 anonymised notes incl. 3 with two pupils of
the same first name.
[x] D4 Coordinator reads every update before it's sent.
[x] D8 Never drafted by AI: anything about wellbeing, behaviour
or safeguarding. Goes to the owner.
[ ] A3 Runbook written; not yet followed by a second person.
[x] A6 Renewal in 11 months; reminders set at 60 and 30 days.
[x] A7 Leaver steps updated: remove from AI workspace and shared
project on the last day.
Two items are unticked, each with a named owner and a date, which is how the checklist is meant to be used. A list where everything is ticked on day one usually means some ticks are optimistic. Note D1's awkward cases: two pupils with the same first name was exactly the situation most likely to produce an update about the wrong child, so it went into the test set on purpose.
Checks people skip, and how it showed up
- A6, renewals. An illustrative picture framer tried an AI image tool for product photos, stopped using it after a month and forgot it was on an annual plan. The renewal charge eleven months later was the first reminder it existed.
- D1 and A8, the test set. A pet shop's AI-drafted replies gradually became longer and more formal after a model update. With no saved test cases to compare against, it took weeks and a customer's comment before anyone noticed.
- A7, leavers. A language school found a teacher who had left three months earlier still had access to the shared AI workspace, including student work pasted for feedback. Removing them took two minutes; noticing took a term.
- D9, alerts. An optician's reminder automation stopped sending when a password changed. Nobody was told, because nobody had tested what happens on failure, and the gap showed up as a rise in missed appointments.
Each of these took minutes to prevent and weeks to discover. That's the pattern the whole checklist is built around: the expensive failures in small-business AI are rarely the AI getting something wrong once; they're the dull checks nobody did, failing quietly for months. If you want those risks tracked in one place, a simple AI risk register gives each one an owner and a review date.
Using the checklist with an outside supplier
If a consultant, freelancer or agency is doing the work, the checklist doubles as your definition of "done". Split it before the contract is signed:
- Yours, whoever builds it: B1 to B3, B5, B6, B8, B10, B12, A4 and A9. These are decisions and duties that can't be outsourced.
- Theirs, as deliverables: D1 to D3, D5, D8, D9, A3 and the technical half of B4 and B11: the test set, the stored instructions, the access list, the failure alerts, the runbook and the exit route.
- Shared: B9, where you set the criteria and they confirm they can be measured, and A1, where they supply the numbers and you judge them.
Then make payment for the final stage depend on their items being evidenced, not just demonstrated. The AI consultant handover checklist covers what you should hold before a supplier leaves, which overlaps heavily with the "after" stage here.
Scaling the list to the size of the job
Thirty checks is right for a workflow that touches customers or runs automatically. For smaller uses, a subset does the job:
- One person drafting internal documents with a business AI account: B4, B6, B7, D4 and A6. Five checks, about an hour.
- A team using AI drafting with a person checking everything: add B1, B2, B3, B9, D1, D5, D7, A1, A5 and A7.
- Anything automated or customer-facing: all 30, plus B8 read carefully.
Whichever size applies, keep the evidence with the checklist in one folder: the baseline figures, the signed criteria, the test set, the runbook and the renewal reminders. A year from now, when you're deciding whether to keep, change or replace the tool, that folder will answer most of the questions in minutes.
Further reads
- How to Set a Baseline Before You Introduce AI — Do check B2 properly, with records you already keep.
- Did Your AI Pilot Work? How to Set Success Criteria That Hold Up — Write the criteria behind check B9 so they can actually fail.
- How to Write an AI Usage Policy for Your Small Business — Wording for the one-page AI rule in check B6.
- What to Check in an AI Vendor's Data Processing Agreement — What to look for in the vendor terms behind check B7.
- How to Review an AI Tool After 90 Days: Keep, Fix or Cancel — Run check A9 as a scored keep, fix or cancel review.
- AI Maintenance Costs: What You Pay After an Automation Goes Live — Budget the upkeep that check A2 asks you to count.
- AI Literacy Requirements: What Your Staff Need to Know — What staff need to know, and the EU literacy duty.
- What AI Implementation Actually Involves for a Small Business — The six pieces of work behind a small-business AI implementation, with rough hours for each and one workflow followed from baseline to handover.
- 10 AI Implementation Mistakes Small Businesses Make — Ten setup, billing and upkeep mistakes that turn a sensible AI project into a year-long cost, each with how it shows up and the fix.
- AI Tools and AI Development: The Complete 2026 Guide — the AI hub, including every tutorial in the AI-for-business series.
Sources: Zapier help centre (auto-pause and error handlers); Make help centre (error handlers); Microsoft Learn (Power Automate flow suspension; Copilot usage report); OpenAI help centre (custom GPT retirement); Microsoft support (Excel COPILOT function retirement); EU AI Act Articles 4 and 50; vendor notices on Clockwise and SimplePractice Note Taker.