How to Choose Your First AI Project: 7 Tests Before You Commit

Coding Liquids tutorial cover featuring Sagnik Bhattacharya for How to Choose Your First AI Project: 7 Tests Before You Commit.
Coding Liquids tutorial cover featuring Sagnik Bhattacharya for How to Choose Your First AI Project: 7 Tests Before You Commit.

Shortlist three to five jobs, then run each through seven pass-or-fail tests: it's weekly and takes two-plus hours, you can measure it now, it runs on tools you have, it fits one sentence, its worst mistake is cheap, it can be live in six weeks, and you can walk away cleanly. Commit to one that passes all seven.

Most first projects that fail would have flunked test 5 or test 7: they started with something customers see, or they signed up for a year before anyone knew whether it worked. If nothing on your list passes all seven, take the candidate that fails only test 6 and cut its scope until it passes. For the general traits of a good first job, see what to automate first with AI; the tests here are for deciding between the specific candidates in front of you.

Follow me on Instagram@sagnikteaches

Build a shortlist worth testing

You need three to five real candidates, each named as a job someone does now, not as a technology. "Use a chatbot" isn't a candidate; "answering parents' questions about lesson times" is. If you don't have a list, an afternoon of auditing your workflows for AI opportunities will produce one with hours attached.

Connect on LinkedInSagnik Bhattacharya

Take an illustrative music studio: the owner, who teaches, plus two part-time teachers, with about 90 pupils, most of them children. The owner's shortlist ran to five jobs:

Subscribe on YouTube@codingliquids
  1. Practice summaries: turning each teacher's lesson notes into a short weekly practice summary for parents.
  2. Fee reminders: chasing unpaid termly fees.
  3. Website chatbot: answering enquiries about lessons, prices and availability.
  4. Recital posts: social posts about concerts and exam results.
  5. Make-up lessons: rescheduling lessons after cancellations.

Test 1: it happens weekly and takes two hours or more

A first project needs enough volume for the saving to be felt and for you to learn quickly. The threshold: it happens at least weekly, and across everyone who does it, it takes two hours a week or more. A painful peak, such as 20 hours at the end of each term, can count instead, but then your pilot has to wait for the peak.

Check it in ten minutes: count last week's instances from email, the diary or your system, and multiply by a rough time each. At the studio, practice summaries came to about 4.5 hours a week (90 pupils at roughly 3 minutes each, split across three teachers), make-up lessons about 2, the chatbot's enquiries about 1.5, recital posts about 1. Fee reminders were a termly burst, about 4 hours in the fortnight after fees fell due. Recital posts and the chatbot failed; fee reminders were marked "peak only".

Test 2: you can put a number on it this week

If you can't measure the job before you start, you can't prove the project worked, and a first project that can't prove itself rarely gets a second. Pass if you can get a baseline within five working days: minutes per item, volume, turnaround, or how often it's redone.

A baseline doesn't need special software. The studio's teachers each timed ten summaries with a phone stopwatch over two weeks and wrote the minutes in a shared sheet: the average was a little under 3 minutes, with the longest, about 6 minutes, for pupils preparing for exams. That spread matters later, because the exam pupils are where AI drafts are most likely to need rewriting.

Check it: ask where the number would come from. Practice summaries and make-up lessons passed, since both leave a trail in email and the diary. The chatbot failed a second time, because most enquiries came by phone and nobody counted them.

Test 3: it runs on tools you already have, plus one at most

A first project should teach you about AI, not about migrating systems. Pass if it runs on software you use now, plus at most one new subscription that you could cancel monthly. The studio runs on Google Workspace, whose plans include Gemini in Gmail from Business Starter upwards, and the owner already had a $20-a-month chat assistant.

Practice summaries, fee reminders and posts passed easily. The chatbot needed a new tool and changes to the website. Make-up lessons needed the booking system to expose free slots, which it couldn't. Both failed.

Test 4: it fits in one sentence

Write the job as: when [input arrives], produce [output], which [person] checks before [it goes where]. If you need "and" twice, it's two projects.

  • Pass: "When a teacher saves lesson notes, draft a five-line practice summary for the parent, which the teacher checks before sending."
  • Pass: "Two weeks after fees fall due, draft a polite reminder for each unpaid family, which the owner checks against the bank statement before sending."
  • Fail: "Answer whatever parents ask on the website, book trial lessons and take payments."

The failing sentence isn't wrong as an ambition. It's three projects, and the hardest one, answering anything, is buried in the middle.

Test 5: its worst likely mistake is cheap and caught early

Imagine the most likely bad output, not the most dramatic one, and ask what it would cost and who would catch it. Pass if a person checks before anything leaves the business and the mistake costs minutes, not money, trust or safety.

At the studio, a practice summary with a wrong piece named would be caught by the teacher; pass. A fee reminder sent to a family who had already paid is embarrassing, so the check against the bank statement is written into the sentence; pass. A chatbot quoting last year's prices or promising a slot that doesn't exist reaches the parent unchecked; fail. The general case for starting behind the scenes is set out in whether AI should go customer-facing or back-office first.

Here's how failing this test tends to show up. An illustrative pet shop chose a website chatbot as its first project because it looked the most modern. Within a fortnight it had given a customer a feeding amount for a puppy that the staff would never have recommended, and staff were spending more time checking the bot's conversations than they'd ever spent answering messages. The owner switched it off, and nobody wanted to hear about AI for months.

Test 6: a working version can be in daily use within six weeks

Six weeks of part-time effort is long enough to build and adjust something simple and short enough to keep attention. Pass if a first version could be in real, daily use within six weeks, with the owner or one person giving about two hours a week.

This is the test most often failed by good ideas that are too big. The fix is to cut the scope, not to extend the deadline. Make-up lessons failed as "reschedule cancellations automatically", but "draft the message offering the three free slots the teacher has picked" would pass.

Test 7: you can walk away at week six with nothing lost

You'll only judge a pilot honestly if stopping is cheap. Pass if there's no annual contract, your data and instructions stay in accounts you own, and customers haven't been promised anything that depends on the tool.

Watch the billing terms. Some AI products are annual-only; Microsoft 365 Copilot Business, for instance, requires an annual commitment. Annual versus monthly billing for AI tools explains when the discount is worth the lock-in, which for a first project is almost never. The pet shop's chatbot was on an annual plan, so the owner paid for eleven months of a tool that sat switched off: a failed test 7 on top of a failed test 5.

Five candidates put through all seven tests

TestPractice summariesFee remindersWebsite chatbotRecital postsMake-up lessons
1 Weekly, 2+ hoursPassPeak onlyFailFailPass
2 Measurable nowPassPassFailPassPass
3 Tools you havePassPassFailPassFail
4 One sentencePassPassFailPassFail
5 Cheap mistakesPassPassFailPassPass
6 Live in six weeksPassPassFailPassFail
7 Clean exitPassPassFailPassPass

Practice summaries pass all seven and become the first project. Fee reminders pass everything except the weekly volume and go second, timed for the next fee date. Recital posts pass six but save too little to be worth a pilot; the owner can just use the chat assistant for them. The chatbot fails everything that matters and goes on a "later, if ever" list. Make-up lessons need the booking system fixed first.

It's normal for only one candidate to pass everything. The tests are strict on purpose, because a first project has a second job: proving to you and your team that AI can be introduced without drama. Pick the candidate most likely to succeed, not the one with the biggest possible payoff.

Breaking a tie between two candidates that pass

If two candidates pass all seven, choose using these tie-breakers, in order:

  1. The person who does the job wants help with it. A willing teacher beats a sceptical one with twice the hours.
  2. The saving lands on your most stretched person. An hour freed for the owner is often worth more than an hour freed elsewhere.
  3. It leaves something reusable behind. A house-style guide, a set of example outputs or a shared project the next job can use.
  4. Its results are easy to show. A before-and-after anyone can read makes the second project easier to start.

Here's a tie in practice. An illustrative tutoring agency found two candidates that passed all seven tests. The first was turning tutors' session notes into short progress updates for parents, about 5 hours a week across two coordinators. The second was drafting reminders to tutors whose timesheets were late, about 2.5 hours a week, all the owner's.

  • Tie-breaker 1: both coordinators had asked for help with the updates; the owner was indifferent about timesheets. Point to the updates.
  • Tie-breaker 2: the owner was the most stretched person, and the timesheet reminders were entirely theirs. Point to the reminders.
  • Tie-breaker 3: the updates would leave behind a house style for writing to parents, which the agency's enquiry replies could reuse later. The reminders would leave a template, not much more. Point to the updates.
  • Tie-breaker 4: parents' replies to clearer updates would be visible to everyone; nobody would notice smoother timesheets. Point to the updates.

The updates went first. The owner handled the timesheet reminders with a plain email template in the meantime, which took ten minutes to write and needed no AI at all. Tie-breakers rarely produce a 4-0 result; the point is to make the reasoning visible, so nobody feels their job was passed over on a whim.

One hour on ten real cases before you commit

The tests say a candidate is sensible; they don't say AI does it well in your business. Before signing anything, spend an hour trying it by hand on ten real, recent cases with names removed. For the studio, the prompt was:

You write weekly practice summaries for parents of music pupils.
Use ONLY what is in the teacher's notes below. Five lines at most:
what we worked on, what to practise, how long, one encouragement.
Plain, warm, no jargon. If the notes don't say how long to
practise, write "[ASK TEACHER]" instead of guessing.

Teacher's notes: [paste notes, pupil's first name only]

An illustrative output for one pupil:

"This week Maya worked on the first page of her grade piece and her G major scale, hands together. Please practise the scale slowly every day, then the tricky bars 9 to 12 of the piece. Aim for 20 minutes a day. She's really improving her rhythm, well done!"

What you'd fix: the notes never mentioned 20 minutes, and the instruction to write [ASK TEACHER] was ignored, which is exactly the kind of invented detail to catch in a desk test rather than in a parent's inbox. The teacher added the real practice time and tightened the instruction to "never state a practice time the notes don't give". Across the ten cases, eight summaries needed only small edits, and the time per summary fell from about 3 minutes to about 1. That's enough to commit.

The commitment note: a page to sign before money moves

Write one page and have whoever owns the project sign it. It turns a good idea into a decision with limits. The studio's, filled in:

FIRST AI PROJECT - COMMITMENT NOTE
Job: weekly practice summaries for parents.
One sentence: when a teacher saves lesson notes, draft a five-line
  summary, which the teacher checks before sending.
Baseline: about 3 minutes per summary, 90 a week, about 4.5 hours
  a week across three teachers (timed over two weeks).
Target: about 1.5 minutes per summary including the check, with
  no more than 1 in 10 needing a rewrite.
Owner: studio owner, 2 hours a week for six weeks.
Tools: existing Workspace plan; owner's chat assistant ($20 a
  month). No new contracts.
Safety: first names only; teacher reads every summary.
Review: 15 minutes each Friday. Decision on [date]: keep,
  adjust or stop.
Exit: stop sending AI drafts; notes stay in our Drive.
Not in scope: chatbot, booking, payments.

A quick sum shows why the studio was happy to sign. Saving about 1.5 minutes on each of 90 summaries frees roughly 2.25 hours a week, or about 80 hours over a 36-week teaching year. At an illustrative $30 an hour for teaching staff, that's around $2,400 a year of time, against about $240 a year for the chat assistant. Even if the saving turned out at half the estimate, the project would pay for itself several times over, and the commitment note caps the downside at six weeks and one subscription.

The last line matters as much as the first. Writing down what the project won't do is how you stop it growing into the multi-part project that failed test 4. The target line is the start of your pass mark; setting pilot success criteria that hold up shows how to make it firm enough to judge the pilot fairly at week six.

When nothing on the list passes

Sometimes every candidate fails a test. That's useful information, and the failing tests tell you what to do:

  • Everything fails test 2: you're not measuring anything. Spend two weeks counting, then re-test.
  • Everything fails test 3: your core systems can't share data. The first project may be fixing that, with no AI involved.
  • Everything fails test 5: your list is all customer-facing. Look for the internal step behind each one, such as drafting the reply a person sends, rather than the reply itself.
  • Everything fails test 6: your candidates are too big. Cut each to its first step and test again.

The seven tests are a filter, not a verdict on whether AI suits your business. A list that fails today usually has a passing candidate hidden inside it, one step smaller and one step further from the customer.

Further reads

Sources: OpenAI, Anthropic and Google plan pages (individual and business plan prices); Google Workspace pricing page (Gemini in Starter and above); Microsoft 365 Copilot Business terms (annual commitment); Zapier plan limits.

Want a second opinion on your first AI project?

On a 1:1 call we'll run your shortlist through these tests together, pick the candidate most likely to succeed in your business, and write the commitment note before you spend anything.

Book a 1:1 call with me