Put each AI idea on one page and score it 1 to 5 on six things: how often the task happens, how long it takes, how cheap a mistake is to fix, whether the data is ready, whether one of your current tools can do it, and whether staff want it. Park anything that fails a knock-out question, then start with the highest score.
Below is a copyable one-page template, scoring anchors that make two people give the same idea the same score, a worked example from a veterinary practice, and a 45-minute routine for scoring a list of ideas with your team.
The one-page template
Copy this into a document or a shared spreadsheet tab. One idea per page. The person who suggests an idea fills in sections 1 and 2 before the scoring meeting, using real counts rather than impressions.
AI USE CASE: [short name, e.g. "Repeat prescription requests"]
Proposed by: [name] Date: [date]
1. THE JOB TODAY
What happens, in 3-5 steps:
Who does it:
How often (per week, counted not guessed):
Minutes each time (timed on 5 real examples):
Where the inputs come from (phone, email, web form, paper, system):
Where the result goes:
2. WHAT AI WOULD DO
The part AI drafts, sorts, extracts or answers:
The part a person still checks or decides:
Tool we'd try first: On our current plan? Y / N
3. KNOCK-OUT QUESTIONS (any YES = park it for now)
K1 Would AI make a clinical, legal or money decision with nobody signing it off?
K2 Would personal or client data go into a tool without a business agreement?
K3 Could a mistake reach a customer before anyone has a chance to catch it?
4. SCORES (1-5, using the anchors)
Volume ___ Time per go ___ Error cost ___
Data readiness ___ Tool fit ___ Team appetite ___
Weighted total ___ / 100 Hours per month ___
5. EVIDENCE
Where the volume figure came from:
Who was asked about appetite:
6. DECISION: Build now / Trial in shadow / Park / Drop
Owner: Review date:
Section 5 is the one people skip and the one that matters most. A score with no evidence behind it is an opinion with a number attached.
Here's a page filled in, as an illustration, for a six-person plumbing and heating firm. The scores use the anchors in the next section:
AI USE CASE: Turn engineers' voice notes into job reports
Proposed by: office manager Date: 2 Sep
1. THE JOB TODAY
Steps: engineer sends a voice note after each job; office
listens, types the report, adds parts used, emails it to
the customer (and the landlord, for rentals)
Who: office manager, part-time admin
How often: 118 and 124 jobs in the last two weeks (~60/week)
Minutes each: timed 9, 11, 12, 14, 20; median 12
Inputs: WhatsApp voice notes on 4 engineers' phones
Result goes: customer email; copy in the job system
2. WHAT AI WOULD DO
AI: transcribe the note, draft the report in our layout
Person: checks parts and prices, sends it
Tool: transcription plus our AI assistant; on our current
plan? N (needs a business-plan subscription)
3. KNOCK-OUTS K1 No K2 No, if business plan only K3 No
4. SCORES Volume 5 Time 3 Error cost 3 Data 2 Tool 3
Appetite 5 (engineers asked for it)
Weighted total 71/100 Hours per month ~52
5. EVIDENCE Counts from the job system; timings by office
manager; appetite from all 4 engineers
6. DECISION Trial in shadow for 4 weeks
Owner: office manager Review: 30 Oct
Data readiness scored 2 because the notes sit on four personal phones, which is also why the idea only passes K2 with a condition. The condition, "business plan only", is written into the knock-out line so nobody forgets it when the trial starts.
AI can help turn someone's rambling description into sections 1 and 2, but watch what it fills in. A prompt like this:
Here's how our receptionist describes a task. Turn it into
the "job today" and "what AI would do" sections of our use-case
template: 3-5 steps, who does it, inputs, where the result goes.
Leave "how often" and "minutes each" as [COUNT] and [TIME].
Description: [PASTE]
An illustrative response gets the steps and inputs right, then writes "How often: around 25 per week" and "Tool: our booking system's AI assistant (already included)", despite being told to leave the count blank and having no idea what your plan includes. Delete both. The two fields the AI is keenest to fill are the two that must come from a count and a check of your plan.
Scoring anchors so two people score the same idea the same way
Without anchors, one person's 4 is another person's 2. These definitions give every score a meaning you can check. The weights add up to 100, so the weighted total reads as a percentage.
| Criterion (weight) | Scores 1 | Scores 3 | Scores 5 | How to check it |
|---|---|---|---|---|
| Volume (20) | Under 5 times a month | 5-20 times a week | 50+ times a week | Count two normal weeks in the phone log, inbox or booking system |
| Time per go (15) | Under 2 minutes | 5-15 minutes | 30+ minutes | Time five real instances; use the median |
| Error cost (20), 5 = cheap | A slip could harm someone or breach a rule | Embarrassing, fixed within a day | Internal only, fixed in minutes | Ask: what happened the last time a person got this wrong? |
| Data readiness (15) | In people's heads or on paper | Digital but spread over 3+ places | In one system, filled in consistently | Open ten recent records and see how many are complete |
| Tool fit (15) | Needs a custom build | Needs a new subscription or a Zapier/Make link | A feature in a tool on your current plan | Check your plan tier, not just the product name |
| Team appetite (15) | The people doing it don't want it changed | Neutral | They asked for it | Ask the people who do the task, not their manager |
To get the total, multiply each score by its weight, add them up and divide by 5. In a spreadsheet with the six scores in columns B to G, the formula is:
=(B2*20 + C2*15 + D2*20 + E2*15 + F2*15 + G2*15) / 5
Add one more column that isn't scored: hours per month, which is volume per week times minutes each, times 4.3, divided by 60. It keeps the conversation honest. Here are both sums for idea B in the veterinary example below. Scores of 5, 2, 3, 4, 3 and 4 give (100 + 30 + 60 + 60 + 45 + 60) / 5 = 71. About 120 requests a week at 4 minutes each is 120 × 4 × 4.3 / 60, or roughly 34 hours a month. An idea can score well and still only free up 90 minutes a month, and if you're weighing a one-off chore, the tutorial on whether a weekly task is worth automating covers that edge case.
Why the knock-out questions come before the arithmetic
Weighted scores average things out, and that's exactly the problem with risk. An idea can score 5 on volume, time and appetite and still be one you shouldn't touch, because a single failure would be serious. Averaging lets three good scores hide one dangerous one. So the three knock-outs sit above the scoring and act as gates.
- K1, unsupervised decisions. AI can draft a treatment explanation, a refund decision or a quote. It shouldn't send one that commits you, with no named person approving it.
- K2, data without an agreement. Client or patient details only go into tools on a business plan with terms you've read. The tutorial on classifying business data before using AI helps you decide which data counts.
- K3, no chance to catch mistakes. If output would go straight to a customer, the idea isn't dead. Redesign it to start in shadow mode, where AI drafts and a person sends, and keep it there until you trust the error rate.
"Park" doesn't mean "never". Write on the page what would have to change for the idea to pass: a business licence, a human sign-off step, a cleaner data source. That note is what brings it back in six months. For example, a small home-care agency's idea "summarise carers' visit notes into a weekly update for families" fails K2, because the notes contain health details. Its parking note might read: "Parked, fails K2. Un-park when: the AI summary feature in our care-management software is on our plan and its data terms have been checked by our data-protection adviser. Not to be tried in any general chat tool, with or without names."
Worked example: five ideas from a veterinary practice
Say a two-site small-animal practice with four vets, six nurses and five receptionists collects five ideas in a staff meeting. The figures below are invented for illustration, but they show how the sheet behaves. PMS means the practice management system, the software that holds appointments and clinical records.
| Idea | Vol | Time | Error | Data | Tool | Appetite | Total | Hours/month | Knock-out |
|---|---|---|---|---|---|---|---|---|---|
| A. Draft discharge instructions from the vet's notes | 4 | 3 | 2 | 4 | 3 | 5 | 69 | 32 | Pass (vet signs off) |
| B. Take repeat prescription requests and queue them for a vet | 5 | 2 | 3 | 4 | 3 | 4 | 71 | 34 | Pass |
| C. Summarise referral histories from other practices | 2 | 5 | 3 | 2 | 4 | 4 | 65 | 10 | Pass |
| D. Personalise vaccination reminders | 5 | 1 | 4 | 5 | 5 | 3 | 78 | 11 | Pass |
| E. Answer "should I bring my dog in?" calls | 5 | 3 | 1 | 2 | 2 | 2 | 51 | 40 | Fails K1 |
A few things stand out, and they are typical.
- The top score isn't the biggest saving. Idea D wins on points because the PMS already sends reminders and only needs its templates tuned. About 150 reminders a week each get a minute of manual tidying today, so it frees roughly 11 hours a month. That makes it a one-week configuration job to do straight away, not the main project.
- The main project is B. Around 120 requests a week at four minutes each is roughly 34 hours a month of reception time. AI collects the request, checks when the animal was last seen, and puts it in a queue for a vet to authorise. The vet still makes the decision.
- A goes to shadow mode. The vets want it badly, but a wrong dose in a discharge sheet is serious, so for the first month the AI drafts and a vet edits every one.
- C waits for better data. Many referral histories arrive as scanned images. The parking note reads: "revisit when referrals arrive as text PDFs or email."
- E is parked, not scored away. It would save the most time of all, 40 hours a month, but it asks AI to make a clinical triage call. The note says what would un-park it: a vet-approved script that only books appointments and never gives advice.
Notice how close the totals are: 65 to 78 for the four ideas that pass. That's normal. The score puts the ideas in order for discussion; the hours column and the knock-outs make the decision.
Running the scoring session in 45 minutes
- A week before: ask each person with an idea to fill in sections 1 and 2, with a two-week count and five timed examples. Ideas without counts wait for the next round.
- Minutes 0-5: read each idea's knock-out answers aloud. Park any that fail and write the un-park condition.
- Minutes 5-30: for each remaining idea, everyone writes their six scores silently, then reveals them together. Only discuss criteria where scores differ by two or more, and discuss the evidence, not who is right.
- Minutes 30-40: sort by total, look at the hours column, and pick one idea to build and at most one to trial.
- Minutes 40-45: give each chosen idea a named owner and a review date no more than 90 days out.
Silent scoring matters more than it sounds. In a small team, if the owner says "I think this is a 5" first, everyone else anchors on it.
The step 3 discussion is where the anchors earn their place. At the veterinary practice, idea C's error-cost scores came back as a 4 from one vet and a 2 from a senior nurse. The vet's reasoning: a referral summary is internal and a vet reads the full history anyway. The nurse's: the last time a history was summarised by hand, a drug reaction noted on page six never made it to the summary, and the animal was nearly given the same drug again. The anchor for 3 reads "embarrassing, fixed within a day"; the nurse's case was worse than that, but it was caught because a vet re-read the original. They settled on 3, with a condition written on the page: the summary must list every medication and reaction mentioned, even if it seems minor.
The session also needs a plan for the idea everyone likes that scores badly. Often it's the owner's own. At the practice, the practice owner's idea of AI-written social posts about pet care scored 49: volume 2 (two posts a week), time 3 (about 20 minutes each), error cost 2 (wrong pet-care advice in public could harm an animal), data 4, tool 3, appetite 1, since nobody else wanted to do it. That's (40 + 45 + 40 + 60 + 45 + 15) / 5 = 49, and its hours column showed about 3 a month. The honest outcome was "drop from this round", with a note that the owner can try it on her own time with a free tool, as long as no client or patient details go in. Writing that down stops the idea coming back at every meeting with a fresh coat of enthusiasm.
Where scoring sheets flatter an idea
- Guessed volume. Annoying tasks feel more frequent than they are. People confidently say "we get dozens of those calls a day", then the log shows eleven. Count, always.
- Hours nobody can use. Saving six minutes, fifteen times a day, spread across five people, rarely turns into anything. The tutorial on why AI saves time but not money explains how to turn freed minutes into a real gain.
- Two ideas in one row. "AI for reception" isn't a use case. Split it into phone, email, forms and reminders, and score each one. A small hotel's "AI for guest emails" row scored a comfortable 70 until it was split: pre-arrival questions (parking, check-in times) scored 80; booking changes scored 64 because they touch the reservations system; and complaints failed K3, since a drafted reply to an angry guest must never go out unchecked. One average had been hiding a knock-out.
- Data scored by the wrong person. The manager scores data readiness 4; the receptionist who fills in the records knows half the fields are blank. The person closest to the data scores this one.
- Tool fit that ignores the plan tier. "Our booking system has AI" often means "the top tier has AI". Check which plan you're on before scoring 5.
- No baseline. If you don't record today's time and error rate before you start, you can't tell later whether the idea delivered. Set it now, using the method in how to set a baseline before you introduce AI.
Keeping the sheet useful after the first round
A scoring sheet that's filled in once and filed is wasted effort. Three habits keep it working.
Keep a parked tab. Every parked or dropped idea stays on the sheet with its reason and its un-park condition. When something changes, such as a new plan tier, a cleaner data feed or a signed data agreement, look down that tab first.
Record the result next to the prediction. After 90 days, add the actual hours saved and the actual error rate beside the scores. For the veterinary practice's idea B, that row might read: predicted 34 hours a month, actual about 22. The reason is in the notes column: roughly a third of repeat-prescription requests still arrived by phone, where nobody had changed anything, so reception still keyed them in by hand. The next move was a phone-line change, not a better prompt, and that only showed up because someone wrote the actual figure next to the prediction. After two or three rounds, you'll see where your team scores too high (usually volume and appetite) and you can adjust.
Re-score every quarter. Scores go stale. One of your existing tools gains a feature, so an idea's tool fit jumps from 3 to 5. A new hire takes over a task and appetite changes. Fifteen minutes a quarter keeps the list current.
If your idea list is thin, start from the customer's side instead: walking through your customer journey usually turns up three or four candidates nobody had written down.
Questions about scoring AI ideas
Should I change the weights in the template?
Yes, if your business has a clear constraint. A practice with a strict regulator might raise error cost to 30 and cut team appetite to 10. Change the weights once, before you score anything, and keep them for the whole round. Adjusting weights after seeing the totals is how a favourite idea gets pushed to the top.
How many ideas should we score in one session?
Between five and twelve works well in 45 minutes. Fewer than five and you are not really choosing. More than twelve and scoring gets rushed, so the last ideas get lazy middle scores. If you have thirty ideas, run the knock-out questions on all of them first, then score only the survivors.
What if every idea scores below 60?
That usually means the data is scattered or your current tools can't do the job yet. Look at which criterion drags the scores down. If it is data readiness, the first project is tidying one system, not buying AI. If it is volume, your business may simply not have a task big enough to automate yet, which is a useful answer too.
Further reads
- What Should a Small Business Automate First With AI? — Common first targets to add to your idea list.
- How to Calculate AI ROI for Your Business (Worked Example) — Turn the top-scoring idea into a money case.
- How to Run Your First AI Pilot Project in a Small Business — What to do once an idea wins the vote.
- How to Run a Pre-Mortem Before You Launch an AI Project — Stress-test the winner before you launch it.
- AI Use Cases by Department for Small Businesses (With Examples) — Examples to seed a thin idea list.
- A Simple AI Risk Register for Small Businesses (With Template) — Where parked risks from the knock-outs should live.
- How to Write a One-Page AI Strategy for Your Business — The seven boxes a one-page AI strategy needs, a filled-in garden centre example, and five tests that show whether your page will guide real decisions.
- How to Run a 60-Minute AI Workshop for Your Team — A minute-by-minute plan for a one-hour AI session: preparation, live demos on real tasks, paired practice, data rules and the week after.
- How to Map a Business Process Before You Automate It — A practical process-mapping method for automation: six columns per step, swimlanes drawn with AI help, an exception count and a label for every step.
- AI Tools and AI Development: The Complete 2026 Guide — the AI hub, including every tutorial in the AI-for-business series.