Shortlist three to five jobs, then run each through seven pass-or-fail tests: it's weekly and takes two-plus hours, you can measure it now, it runs on tools you have, it fits one sentence, its worst mistake is cheap, it can be live in six weeks, and you can walk away cleanly. Commit to one that passes all seven.
Most first projects that fail would have flunked test 5 or test 7: they started with something customers see, or they signed up for a year before anyone knew whether it worked. If nothing on your list passes all seven, take the candidate that fails only test 6 and cut its scope until it passes. For the general traits of a good first job, see what to automate first with AI; the tests here are for deciding between the specific candidates in front of you.
Build a shortlist worth testing
You need three to five real candidates, each named as a job someone does now, not as a technology. "Use a chatbot" isn't a candidate; "answering parents' questions about lesson times" is. If you don't have a list, an afternoon of auditing your workflows for AI opportunities will produce one with hours attached.
Take an illustrative music studio: the owner, who teaches, plus two part-time teachers, with about 90 pupils, most of them children. The owner's shortlist ran to five jobs:
- Practice summaries: turning each teacher's lesson notes into a short weekly practice summary for parents.
- Fee reminders: chasing unpaid termly fees.
- Website chatbot: answering enquiries about lessons, prices and availability.
- Recital posts: social posts about concerts and exam results.
- Make-up lessons: rescheduling lessons after cancellations.
Test 1: it happens weekly and takes two hours or more
A first project needs enough volume for the saving to be felt and for you to learn quickly. The threshold: it happens at least weekly, and across everyone who does it, it takes two hours a week or more. A painful peak, such as 20 hours at the end of each term, can count instead, but then your pilot has to wait for the peak.
Check it in ten minutes: count last week's instances from email, the diary or your system, and multiply by a rough time each. At the studio, practice summaries came to about 4.5 hours a week (90 pupils at roughly 3 minutes each, split across three teachers), make-up lessons about 2, the chatbot's enquiries about 1.5, recital posts about 1. Fee reminders were a termly burst, about 4 hours in the fortnight after fees fell due. Recital posts and the chatbot failed; fee reminders were marked "peak only".
Test 2: you can put a number on it this week
If you can't measure the job before you start, you can't prove the project worked, and a first project that can't prove itself rarely gets a second. Pass if you can get a baseline within five working days: minutes per item, volume, turnaround, or how often it's redone.
A baseline doesn't need special software. The studio's teachers each timed ten summaries with a phone stopwatch over two weeks and wrote the minutes in a shared sheet: the average was a little under 3 minutes, with the longest, about 6 minutes, for pupils preparing for exams. That spread matters later, because the exam pupils are where AI drafts are most likely to need rewriting.
Check it: ask where the number would come from. Practice summaries and make-up lessons passed, since both leave a trail in email and the diary. The chatbot failed a second time, because most enquiries came by phone and nobody counted them.
Test 3: it runs on tools you already have, plus one at most
A first project should teach you about AI, not about migrating systems. Pass if it runs on software you use now, plus at most one new subscription that you could cancel monthly. The studio runs on Google Workspace, whose plans include Gemini in Gmail from Business Starter upwards, and the owner already had a $20-a-month chat assistant.
Practice summaries, fee reminders and posts passed easily. The chatbot needed a new tool and changes to the website. Make-up lessons needed the booking system to expose free slots, which it couldn't. Both failed.
Test 4: it fits in one sentence
Write the job as: when [input arrives], produce [output], which [person] checks before [it goes where]. If you need "and" twice, it's two projects.
- Pass: "When a teacher saves lesson notes, draft a five-line practice summary for the parent, which the teacher checks before sending."
- Pass: "Two weeks after fees fall due, draft a polite reminder for each unpaid family, which the owner checks against the bank statement before sending."
- Fail: "Answer whatever parents ask on the website, book trial lessons and take payments."
The failing sentence isn't wrong as an ambition. It's three projects, and the hardest one, answering anything, is buried in the middle.
Test 5: its worst likely mistake is cheap and caught early
Imagine the most likely bad output, not the most dramatic one, and ask what it would cost and who would catch it. Pass if a person checks before anything leaves the business and the mistake costs minutes, not money, trust or safety.
At the studio, a practice summary with a wrong piece named would be caught by the teacher; pass. A fee reminder sent to a family who had already paid is embarrassing, so the check against the bank statement is written into the sentence; pass. A chatbot quoting last year's prices or promising a slot that doesn't exist reaches the parent unchecked; fail. The general case for starting behind the scenes is set out in whether AI should go customer-facing or back-office first.
Here's how failing this test tends to show up. An illustrative pet shop chose a website chatbot as its first project because it looked the most modern. Within a fortnight it had given a customer a feeding amount for a puppy that the staff would never have recommended, and staff were spending more time checking the bot's conversations than they'd ever spent answering messages. The owner switched it off, and nobody wanted to hear about AI for months.
Test 6: a working version can be in daily use within six weeks
Six weeks of part-time effort is long enough to build and adjust something simple and short enough to keep attention. Pass if a first version could be in real, daily use within six weeks, with the owner or one person giving about two hours a week.
This is the test most often failed by good ideas that are too big. The fix is to cut the scope, not to extend the deadline. Make-up lessons failed as "reschedule cancellations automatically", but "draft the message offering the three free slots the teacher has picked" would pass.
Test 7: you can walk away at week six with nothing lost
You'll only judge a pilot honestly if stopping is cheap. Pass if there's no annual contract, your data and instructions stay in accounts you own, and customers haven't been promised anything that depends on the tool.
Watch the billing terms. Some AI products are annual-only; Microsoft 365 Copilot Business, for instance, requires an annual commitment. Annual versus monthly billing for AI tools explains when the discount is worth the lock-in, which for a first project is almost never. The pet shop's chatbot was on an annual plan, so the owner paid for eleven months of a tool that sat switched off: a failed test 7 on top of a failed test 5.
Five candidates put through all seven tests
| Test | Practice summaries | Fee reminders | Website chatbot | Recital posts | Make-up lessons |
|---|---|---|---|---|---|
| 1 Weekly, 2+ hours | Pass | Peak only | Fail | Fail | Pass |
| 2 Measurable now | Pass | Pass | Fail | Pass | Pass |
| 3 Tools you have | Pass | Pass | Fail | Pass | Fail |
| 4 One sentence | Pass | Pass | Fail | Pass | Fail |
| 5 Cheap mistakes | Pass | Pass | Fail | Pass | Pass |
| 6 Live in six weeks | Pass | Pass | Fail | Pass | Fail |
| 7 Clean exit | Pass | Pass | Fail | Pass | Pass |
Practice summaries pass all seven and become the first project. Fee reminders pass everything except the weekly volume and go second, timed for the next fee date. Recital posts pass six but save too little to be worth a pilot; the owner can just use the chat assistant for them. The chatbot fails everything that matters and goes on a "later, if ever" list. Make-up lessons need the booking system fixed first.
It's normal for only one candidate to pass everything. The tests are strict on purpose, because a first project has a second job: proving to you and your team that AI can be introduced without drama. Pick the candidate most likely to succeed, not the one with the biggest possible payoff.
Breaking a tie between two candidates that pass
If two candidates pass all seven, choose using these tie-breakers, in order:
- The person who does the job wants help with it. A willing teacher beats a sceptical one with twice the hours.
- The saving lands on your most stretched person. An hour freed for the owner is often worth more than an hour freed elsewhere.
- It leaves something reusable behind. A house-style guide, a set of example outputs or a shared project the next job can use.
- Its results are easy to show. A before-and-after anyone can read makes the second project easier to start.
Here's a tie in practice. An illustrative tutoring agency found two candidates that passed all seven tests. The first was turning tutors' session notes into short progress updates for parents, about 5 hours a week across two coordinators. The second was drafting reminders to tutors whose timesheets were late, about 2.5 hours a week, all the owner's.
- Tie-breaker 1: both coordinators had asked for help with the updates; the owner was indifferent about timesheets. Point to the updates.
- Tie-breaker 2: the owner was the most stretched person, and the timesheet reminders were entirely theirs. Point to the reminders.
- Tie-breaker 3: the updates would leave behind a house style for writing to parents, which the agency's enquiry replies could reuse later. The reminders would leave a template, not much more. Point to the updates.
- Tie-breaker 4: parents' replies to clearer updates would be visible to everyone; nobody would notice smoother timesheets. Point to the updates.
The updates went first. The owner handled the timesheet reminders with a plain email template in the meantime, which took ten minutes to write and needed no AI at all. Tie-breakers rarely produce a 4-0 result; the point is to make the reasoning visible, so nobody feels their job was passed over on a whim.
One hour on ten real cases before you commit
The tests say a candidate is sensible; they don't say AI does it well in your business. Before signing anything, spend an hour trying it by hand on ten real, recent cases with names removed. For the studio, the prompt was:
You write weekly practice summaries for parents of music pupils.
Use ONLY what is in the teacher's notes below. Five lines at most:
what we worked on, what to practise, how long, one encouragement.
Plain, warm, no jargon. If the notes don't say how long to
practise, write "[ASK TEACHER]" instead of guessing.
Teacher's notes: [paste notes, pupil's first name only]
An illustrative output for one pupil:
"This week Maya worked on the first page of her grade piece and her G major scale, hands together. Please practise the scale slowly every day, then the tricky bars 9 to 12 of the piece. Aim for 20 minutes a day. She's really improving her rhythm, well done!"
What you'd fix: the notes never mentioned 20 minutes, and the instruction to write [ASK TEACHER] was ignored, which is exactly the kind of invented detail to catch in a desk test rather than in a parent's inbox. The teacher added the real practice time and tightened the instruction to "never state a practice time the notes don't give". Across the ten cases, eight summaries needed only small edits, and the time per summary fell from about 3 minutes to about 1. That's enough to commit.
The commitment note: a page to sign before money moves
Write one page and have whoever owns the project sign it. It turns a good idea into a decision with limits. The studio's, filled in:
FIRST AI PROJECT - COMMITMENT NOTE
Job: weekly practice summaries for parents.
One sentence: when a teacher saves lesson notes, draft a five-line
summary, which the teacher checks before sending.
Baseline: about 3 minutes per summary, 90 a week, about 4.5 hours
a week across three teachers (timed over two weeks).
Target: about 1.5 minutes per summary including the check, with
no more than 1 in 10 needing a rewrite.
Owner: studio owner, 2 hours a week for six weeks.
Tools: existing Workspace plan; owner's chat assistant ($20 a
month). No new contracts.
Safety: first names only; teacher reads every summary.
Review: 15 minutes each Friday. Decision on [date]: keep,
adjust or stop.
Exit: stop sending AI drafts; notes stay in our Drive.
Not in scope: chatbot, booking, payments.
A quick sum shows why the studio was happy to sign. Saving about 1.5 minutes on each of 90 summaries frees roughly 2.25 hours a week, or about 80 hours over a 36-week teaching year. At an illustrative $30 an hour for teaching staff, that's around $2,400 a year of time, against about $240 a year for the chat assistant. Even if the saving turned out at half the estimate, the project would pay for itself several times over, and the commitment note caps the downside at six weeks and one subscription.
The last line matters as much as the first. Writing down what the project won't do is how you stop it growing into the multi-part project that failed test 4. The target line is the start of your pass mark; setting pilot success criteria that hold up shows how to make it firm enough to judge the pilot fairly at week six.
When nothing on the list passes
Sometimes every candidate fails a test. That's useful information, and the failing tests tell you what to do:
- Everything fails test 2: you're not measuring anything. Spend two weeks counting, then re-test.
- Everything fails test 3: your core systems can't share data. The first project may be fixing that, with no AI involved.
- Everything fails test 5: your list is all customer-facing. Look for the internal step behind each one, such as drafting the reply a person sends, rather than the reply itself.
- Everything fails test 6: your candidates are too big. Cut each to its first step and test again.
The seven tests are a filter, not a verdict on whether AI suits your business. A list that fails today usually has a passing candidate hidden inside it, one step smaller and one step further from the customer.
Further reads
- How to Run Your First AI Pilot Project in a Small Business — Run the project you picked as a four to six-week pilot.
- How to Write a One-Page AI Business Case for a Small Business — Turn the commitment note into a business case if others must agree.
- How to Use AI on Your Own Admin First, Then Roll It Out — Why the owner's own admin often makes the safest first project.
- How to Tell If a Process Is Ready to Automate With AI — Score the chosen job's process before anyone builds.
- How to Run a Pre-Mortem Before You Launch an AI Project — Imagine the project failed, then fix those causes in advance.
- AI Proof of Concept vs Pilot vs Production: What Changes? — Know which stage your first project is really at.
- What Happens in a 1:1 AI Implementation Consultation? — What to send beforehand, how a consultant maps your work and scores AI jobs, the tool check, and the one-page plan you should leave with.
- How to Run a Paid Discovery Phase Before a Full AI Project — A stage-by-stage plan for a paid AI discovery phase, with a filled-in brief, a feasibility test on real emails and a go or no-go decision rule.
- How Soon Should AI Pay for Itself? Payback Periods by Project — Payback targets for seven kinds of AI project, the sum with ramp-up months left in, and a toy shop testing four ideas against them.
- After an AI Consultation: Turn the Advice Into a 30-Day Plan — A one-page AI action plan template, a filled-in 30-day plan for a boutique hotel, the owner's weekly check and how to judge the result at day 30.
- How to Prepare for an AI Consultation and Leave With a Plan — The processes, volumes, timings and tool plans to gather before an AI consultation, a filled-in pre-read, the questions to ask and a written plan to leave with.
- AI Readiness Checklist: Score Your Business in 20 Minutes — Fifteen statements, scored 0 to 2 with evidence, tell you whether to pilot AI now, fix a few gaps first or do groundwork before spending anything.
- AI Implementation Roadmap for Small Businesses: 5 Phases in 90 Days — Groundwork, design, pilot, roll-out and review: a five-phase, 90-day AI implementation roadmap with the hours, costs and exit gate for each phase.
- What Can AI Realistically Do for a Small Business? 15 Tasks — Fifteen jobs AI can realistically take on in a small business, rated by how much you can hand over, with tools, costs and the catch for each.
- AI Tools and AI Development: The Complete 2026 Guide — the AI hub, including every tutorial in the AI-for-business series.
Sources: OpenAI, Anthropic and Google plan pages (individual and business plan prices); Google Workspace pricing page (Gemini in Starter and above); Microsoft 365 Copilot Business terms (annual commitment); Zapier plan limits.