List every recurring job in the business, measure each for two weeks, then score them on weekly hours, how much of the work is reading or writing, and how costly a mistake would be. Desk-test the top three on ten real cases each and rank the survivors. Expect six to ten hours spread over three weeks.
An audit finds candidates, not answers. Expect the jobs to sort into piles: a few worth an AI pilot, several that need a simple rule or template rather than AI, and one or two where the process needs fixing before any tool will help. Knowing which pile each job belongs in, before you buy anything, is the whole point of doing it.
Stage 1: list every recurring job, role by role (about 90 minutes)
Start with a list, not ideas. Spend 15 minutes with each role, or each person in a very small team, and ask three questions:
- What do you do every week?
- What do you do every month, term or season?
- What do you redo, chase or fix after someone else?
Don't ask "what could AI do for you?" That question produces wish lists and science fiction. Asking what people do, and what they redo, produces the raw material.
The main example here is an illustrative language school: a director, two administrators, a director of studies and ten teachers, running group courses and private lessons for adults and teenagers, with around 220 students a term. Its 90 minutes of conversations produced 16 recurring jobs:
- Answering course enquiries by email and web form
- Marking the written part of placement tests
- Enrolment and invoicing
- Timetabling classes and rooms
- Arranging cover when a teacher is off
- Following up absences
- Writing end-of-term progress reports
- Writing feedback on students' homework
- Adapting lesson materials to a class's level
- Producing certificates of attendance
- Social media posts
- The re-enrolment newsletter
- Summarising end-of-course feedback surveys
- Handling complaints
- Ordering books from suppliers
- Preparing teachers' hours for payroll
Write each job as a verb and an object ("writing progress reports", not "reports"), with who does it. If two people describe the same job differently, keep both descriptions; the difference is often where the rework hides.
Stage 2: put hours on each job for two weeks
Estimates from memory are usually wrong in both directions: people overstate the jobs they dislike and forget the small ones that happen twenty times a day. So measure. Each person keeps a simple tally for ten working days, about ten minutes a day:
Name: __________ Week of: __________
Job Times done Minutes (total) Redone?
Answering course enquiries |||| 55 1
Placement test marking || 40 0
Homework feedback ... ... ...
Where records already exist, use them instead of tallying. Your email can count for you: in Gmail, a search such as subject:enquiry after:2026/09/01 before:2026/09/15 shows how many enquiry emails arrived in a fortnight, and from:me newer_than:14d shows everything you sent. Booking systems, order histories and invoicing software give you volumes; you only need to time a sample of ten to get minutes per item.
Two things to capture beyond the weekly average. First, peaks: the language school's progress reports took about 30 teacher-hours, all in the last fortnight of term, which a weekly average of 2.5 hours hides completely. Second, redo rates: a job that's fast but redone one time in five is a candidate for a checklist before it's a candidate for AI.
Stage 3: sort each job by the kind of work it is
AI is good at some kinds of work and pointless for others, so label each job with its main type:
| Type | What it looks like | Usual answer |
|---|---|---|
| Drafting | Writing replies, reports, posts or feedback from notes and known facts | Strong AI candidate, with a person checking |
| Extracting | Pulling details out of emails, forms, PDFs or scans into a list or system | Often a good AI candidate |
| Rules | Moving data between systems, sending the same message when X happens | Plain automation or a template; no AI needed |
| Scheduling | Fitting people, rooms and times together | Usually a scheduling tool or a better process |
| Judgement | Decisions about money, welfare, complaints, people | Keep with people; AI may prepare notes |
A chat assistant can do the first pass of this labelling. If your job list is in Google Sheets, the =AI() function can label each row in place, for example =AI("Label this task as drafting, extracting, rules, scheduling or judgement", A2); it generates up to 350 cells at a time and doesn't refresh on its own, so re-run it if you edit the list. Or paste the list into a chat:
Here are the recurring jobs in a small language school, with who
does each. Label each one as drafting, extracting, rules,
scheduling or judgement, and say in one line why. Mark any job where
the output reaches a student, parent or customer with [OUTSIDE].
[paste the list]
An illustrative reply, abridged:
"Answering course enquiries: drafting, since replies are written from known course details [OUTSIDE]. Marking placement tests: extracting, as the marker pulls answers and scores from the paper. Enrolment and invoicing: rules, since the same data moves from form to invoice. Handling complaints: judgement [OUTSIDE]…"
What you'd fix: placement-test marking isn't extraction. The written section needs a teacher's judgement of level, and the result decides which class a student joins, so it's judgement, with AI at most suggesting a level for the teacher to confirm. The assistant labels by what a task sounds like; you label by what goes wrong if it's done badly. Check every label before scoring.
Stage 4: score and rank with one simple formula
Give each job three numbers and multiply them:
- Hours: total hours per week across everyone who does it, averaged over the fortnight.
- Fit, 0 to 3: 3 for drafting from known facts, 2 for extracting or drafting that needs some judgement, 1 for rules or scheduling work, 0 for pure judgement.
- Safety, 0 to 1: 1 if a mistake is caught internally and costs minutes, 0.5 if the output reaches someone outside even after a check, 0 if it's a decision about money, welfare or someone's rights.
Here are the language school's scores for its main jobs:
| Job | Hours/week | Fit | Safety | Score | Pile |
|---|---|---|---|---|---|
| Homework feedback | 12 | 2 | 0.5 | 12 | AI candidate |
| Course enquiry replies | 6 | 3 | 0.5 | 9 | AI candidate |
| Placement test marking | 4 | 2 | 0.5 | 4 | AI candidate, teacher confirms |
| Progress reports | 2.5 (30 at term end) | 3 | 0.5 | 3.75 | AI candidate for the peak |
| Timetabling | 3 | 1 | 1 | 3 | Fix the process first |
| Enrolment and invoicing | 5 | 1 | 0.5 | 2.5 | Rules automation, no AI |
| Absence follow-up | 2.5 | 2 | 0.5 | 2.5 | Template; welfare cases to staff |
| Certificates | 1 | 1 | 1 | 1 | Mail merge template |
| Complaints | 1 | 0 | 0 | 0 | Keep with people |
| Payroll hours | 2 | 1 | 0 | 0 | Keep with people; tidy the timesheet |
The formula is deliberately crude. Its job is to stop the loudest complaint winning and to make the trade-off between hours and risk visible. Progress reports scored low on the weekly average but got a separate note because of their peak: 30 hours in a fortnight is a real problem even if 2.5 hours a week isn't.
Stage 5: desk-test the top three on ten real cases
A score is a guess. Before anything goes on the final list, try each top candidate by hand on ten real, recent, anonymised cases, using a business AI account. It takes about an hour per job and turns guesses into evidence.
Run each desk test the same way, so the results can be compared:
- Pick ten recent cases, including at least two awkward ones. Ten easy cases flatter every tool.
- Remove names and contact details, and anything else your AI rule says stays out.
- Give the assistant what a new member of staff would need: the facts, the house style and one good past example.
- Time the whole thing, including reading and correcting the draft, not just the drafting.
- Grade each output as usable as written, usable after light edits, or rewritten. Note why the rewrites failed.
The language school's results:
- Course enquiry replies: with the course list and term dates pasted in, 8 of 10 drafts were usable after light edits. Two gave the wrong start date for a course, because the term calendar had changed and the pasted version was old. Staff took about 12 minutes per enquiry; drafting and checking took about 5. At roughly 30 enquiries a week, that's around 3.5 hours saved a week, provided the course details are kept current.
- Homework feedback: teachers rated 6 of 10 AI drafts usable. The other four were generic ("good use of vocabulary") in a way students would notice. The saving was real but smaller than the score suggested, perhaps 2 minutes on a 6-minute task, and it depends on each teacher's willingness to adapt the drafts.
- Placement test marking: the written answers were on paper. Scanning them first added more time than the AI saved at the school's volume. (If you do scan paper, note that Microsoft Lens has been retired; the scan feature in the OneDrive or Google Drive mobile app does the same job.) Parked until tests move online.
The desk test also lets you put a value on the winner. At about 3.5 hours a week over 46 working weeks, the enquiry job frees roughly 160 hours a year. At an illustrative $22 an hour for administrator time, that's about $3,500 a year. The school already runs Google Workspace on a plan that includes Gemini, so the tool cost could be nothing extra; even two ChatGPT Business Standard seats on monthly billing would be $50 a month, or $600 a year. Either way the job clears its costs comfortably, which is what earns it the top line of the register.
Desk tests often reorder the list. Here, the job with the highest score fell to second because its saving depended on teacher buy-in, and the third candidate dropped out on a practical detail no score would have caught.
Stage 6: write the opportunity register
The deliverable is one table that anyone in the business can read in two minutes. The language school's, after desk tests:
| Rank | Job | Verdict | Estimated saving | Next step | Owner |
|---|---|---|---|---|---|
| 1 | Course enquiry replies | Pilot with AI drafting | About 3.5 hours a week | Put current course details in a shared project; four-week pilot | Senior administrator |
| 2 | Homework feedback | Pilot with two volunteer teachers | 1 to 4 hours a week, uncertain | Agree the student-work rule; pilot after the enquiry pilot | Director of studies |
| 3 | Enrolment and invoicing | Rules automation, no AI | About 2 hours a week | Price a form-to-invoice automation | Director |
| 4 | Progress reports | Revisit before term end | Peak relief, not weekly | Desk-test on last term's notes | Director of studies |
| 5 | Timetabling | Fix the process first | n/a | One room-booking sheet, not three | Director of studies |
| n/a | Complaints, payroll hours | Keep with people | n/a | None | n/a |
Note the verdicts that aren't AI at all. Enrolment and invoicing needs a plain automation and timetabling needs one source of truth. An audit that only lists AI candidates has usually skipped the sorting in stage 3. Worked examples of workflow automations help you picture the rules-based ones, and checking whether a process is ready to automate is the next step for any job marked "fix the process first".
A 40-minute version for a very small business
With one to four people, the full audit is overkill. An illustrative pet shop, the owner and two part-timers, ran a compressed version:
- Ten minutes: the owner listed 12 recurring jobs from memory and the diary.
- Two weeks of tallies on only the three that felt biggest: supplier reorders, answering product questions by message, and social posts.
- Fifteen minutes to score them: reorders at 3 hours a week came out on top, because the orders go to suppliers the owner knows, mistakes are caught on delivery, and the emails follow a pattern.
- Fifteen minutes of desk test: five past orders drafted from the stock sheet, four usable as written.
The result was a one-line register ("reorder emails first, product questions second") and a pilot starting the following Monday. Small businesses don't need the full table, but they do need the tallies and the desk test; those are what stop you automating the job you complain about rather than the one that takes the time.
Where DIY audits mislead: five traps and how they showed up
- Scoring the loudest job, not the longest. The language school's director disliked timetabling and assumed it was the biggest drain. Homework feedback took four times the hours but was spread thinly across ten teachers, so nobody saw its total until it was tallied.
- Assuming the whole job disappears. Every AI draft still needs reading. The enquiry saving was 7 minutes out of 12, not 12 out of 12. Estimate from the desk test, never from the full duration of the job.
- Averaging away the peaks. A job that's small on average but brutal at term end, month end or the start of a season can be worth more than its weekly score suggests.
- Ignoring where the data physically is. Paper forms, a phone's notes app or one person's inbox can sink a good candidate. Check the inputs are reachable before ranking a job highly.
- Treating the audit as the decision. The register says what's worth trying. Committing to one still needs its own checks, which is where mapping the chosen process in detail comes in.
What to do with the register once it's written
Share it with the team, including the jobs that stay with people, because "we're not automating complaints or payroll" is reassuring news. Pick the top candidate for a four to six-week pilot, with the owner named in the register running it. Park the rest with a date to revisit.
When the first pilot ends, re-score the next two candidates with fresh numbers rather than trusting the original tallies; volumes shift and the first pilot often changes how neighbouring jobs are done. If your business spans several sites or sensitive data, or you'd like the audit done independently, a paid AI audit covers the same ground with an outside view and usually a closer look at your systems and contracts. The register you've built yourself will make that cheaper and faster, because the counting is already done.
Running your own AI opportunity audit: common questions
Should I tell staff why I'm timing their work?
Yes, before you start. Say plainly that you're looking for the repetitive work that eats their week so tools can take some of it, and that the numbers are about tasks, not about judging individuals. Invite them to flag the jobs they'd most like help with. Staff who suspect a hidden agenda under-report the tedious work, which is exactly the work you're trying to find.
Can an AI assistant run the audit for me?
It can help with parts of it: turning interview notes into a job list, labelling each job by type, and summarising your timing sheet. It can't observe your business, so it doesn't know which jobs exist, how long they take or what a mistake costs you. Use it for the paperwork and keep the counting, the scoring and the desk tests with people.
How often should the audit be repeated?
Once a year is enough for most small businesses, or sooner after a big change such as a new booking system, a new service or a jump in volume. Keep the register as a living list in between: when a pilot finishes, re-score the next candidates with fresh numbers rather than relying on figures that are a year old.
What if the top-scoring job belongs to the owner?
Then start there, and treat it as an advantage. The owner can change their own process without negotiating with anyone, feels the time saving directly, and learns what the tool can and can't do before asking staff to use it. The one risk is that the owner is also the person with least time to run the pilot, so book the hours first.
Further reads
- How to Choose Your First AI Project: 7 Tests Before You Commit — Seven pass-or-fail tests for the job at the top of your register.
- How to Set a Baseline Before You Introduce AI — Firm up the timing for the job you choose to pilot.
- What Should a Small Business Automate First With AI? — The traits that make a job a good first candidate.
- AI Audit vs AI Readiness Assessment: What's the Difference? — What a paid audit adds, and how it differs from a readiness check.
- How Much Does It Cost to Automate One Workflow With AI? — Put a price on the automation your audit points to.
- How to Measure Time Saved After Rolling Out AI in a Small Firm — Check later that the estimated savings actually arrived.
- How to Document Your Processes Before Adding AI — How to write down a process so AI can follow it: capture methods, the seven elements to record, turning judgement into rules, and a template.
- What Happens in a 1:1 AI Implementation Consultation? — What to send beforehand, how a consultant maps your work and scores AI jobs, the tool check, and the one-page plan you should leave with.
- What Is an AI Readiness Assessment and What Does It Involve? — What an AI readiness assessment examines, how it runs, a scored example for a homeware brand, and what vendor readiness reports leave out.
- How Much Does an AI Readiness Assessment Cost? — Free, vendor-funded and paid readiness assessments compared, where the days go, your staff's share of the cost, and a way to compare three quotes fairly.
- How to Prepare for an AI Consultation and Leave With a Plan — The processes, volumes, timings and tool plans to gather before an AI consultation, a filled-in pre-read, the questions to ask and a written plan to leave with.
- Signs Your Business Isn't Ready for AI Yet and What to Fix First — Eleven signs a business isn't ready for AI yet, how each shows up day to day, and the cheap fix to make before paying for any tool or project.
- AI Implementation Roadmap for Small Businesses: 5 Phases in 90 Days — Groundwork, design, pilot, roll-out and review: a five-phase, 90-day AI implementation roadmap with the hours, costs and exit gate for each phase.
- AI Tools and AI Development: The Complete 2026 Guide — the AI hub, including every tutorial in the AI-for-business series.
Sources: Gmail help centre (search operators); Google Workspace help on the AI function in Sheets; Microsoft support notice on the retirement of Microsoft Lens.