Spend four to six weeks using AI on three of your own admin jobs, such as enquiry replies, quotes and invoice chasers, logging the time saved and every mistake it makes. Then turn what worked into a short playbook of prompts and rules, hand one task to one colleague, and widen from there once they get the same results.
Going first has a real advantage: you know what a good quote or a good reply looks like, so you're the best-placed person to spot when AI gets it subtly wrong. It has a real trap too. What works for the owner often relies on knowledge only the owner has, quietly applied while editing. The plan below is built to catch that before anyone else inherits it.
Choose three of your own jobs with a simple filter
List the admin you did last week, then score each job on four questions. You want jobs you do several times a week, that are mostly writing or reading, where you can judge the result in seconds, and where a mistake would be embarrassing rather than expensive. Here's the filter applied to an illustrative locksmith owner's week:
| Job | Times a week | Mostly text? | Can I judge it fast? | Cost if wrong | Verdict |
|---|---|---|---|---|---|
| Email replies to lock-change and upgrade enquiries | 20 | Yes | Yes | Low (I read before sending) | Pick |
| Written quotes for commercial lock upgrades | 5 | Yes | Yes | Medium (pricing) | Pick, with a pricing check |
| Chasing unpaid invoices | 15 | Yes | Yes | Low | Pick |
| Supplier re-order emails | 3 | Yes | Yes | Low | Later (too few to matter) |
| Master-key system planning for a client | 1 | Partly | No | High (security) | Never |
| Answering emergency lockout calls | 10 | No, phone | n/a | High | Not an owner-admin task |
The "never" row matters as much as the picks. Some information in every business is too sensitive for any AI tool. For a locksmith that's key codes, master-key charts and alarm details. Write your version of that row down now, because it goes straight into the playbook later.
The same filter gives a different answer in a different business. For the owner of an illustrative two-room physiotherapy clinic, the picks were replies to new-patient enquiries (about 15 a week), rebooking reminders for patients who hadn't booked their next session (about 10), and the monthly newsletter (once a month, but two hours each time). The "never" row was clinical notes and anything from a patient's assessment, and the "later" pile held insurer paperwork, which was frequent but had too many forms that differed by insurer to judge quickly. Each owner's list looks different; the four questions stay the same.
Set it up on a business account from day one
If you trial AI on a personal account, everything you build (saved prompts, projects, chat history) is stuck there when you want to share it. Start on the business account your team will eventually use. If you're on Google Workspace, Gemini is already in Gmail: Help me write drafts replies in the compose window and Summarize this email condenses long threads (see using Gemini in Gmail to clear your inbox). On Microsoft 365, Outlook's Copilot features such as Draft with Copilot and Summarize by Copilot do the same job, depending on your licence; Copilot in Outlook for email triage walks through them. A separate chat assistant on a business plan works too, at the cost of some copying and pasting.
Weeks one and two: do the jobs with AI, and keep a log
Use AI every time one of your three jobs comes up, and log each one. It takes 30 seconds per entry and is the most valuable thing you'll produce in the whole exercise, because it becomes the evidence for the rollout and the source of the rules.
OWNER AI LOG
Date | Task | Mins before (usual) | Mins with AI (incl. checking)
| What I changed in the draft | Error? (type) | Would a colleague
| | | have spotted it?
12 Oct | Enquiry reply | 5 | 2 | Removed "24/7" | Wrong claim | Maybe
12 Oct | Invoice chaser | 4 | 1 | Nothing | - | -
13 Oct | Commercial quote | 25 | 14 | Applied trade discount | Missing rule | No
The last column is the important one. "Would a colleague have spotted it?" separates errors anyone would catch (a typo, a wrong date) from errors only you caught because of something you know. The second kind is what the rollout has to solve.
Log the checking time honestly, because it's the number owners most often leave out. Consider an illustrative kitchen-fitting owner who logged commercial quotes at 25 minutes before and 8 with AI, a 17-minute win on paper. What he hadn't logged was the six minutes he spent afterwards re-checking unit prices against the supplier's list, because he didn't trust the draft's figures. The real saving was 11 minutes, so across four quotes a week it came to 44 minutes rather than 68. Still worth having, but the smaller figure is the one a colleague will actually experience, and it pointed at the real fix: paste the current price list in with every quote request so the check gets quicker.
Weeks three and four: turn the good results into prompts and rules
By now you'll have found a way of asking that works for each task. Save it as a reusable prompt, and fold in the corrections from your log so you stop making them by hand. For invoice chasers, for instance, a stable prompt might look like this:
Draft a payment reminder email for the invoice details below.
Tone depends on how overdue it is:
- 7 to 13 days: friendly nudge, assume it's been missed.
- 14 to 29 days: polite but firm, ask for a payment date.
- 30+ days: firm, state that further work is on hold until paid,
offer a phone call. Do not threaten legal action.
Always include: invoice number, amount, due date, payment link.
Never mention other customers or other invoices.
Sign off as the office, not as me.
Invoice details: [paste]
Here's what that prompt gives back for an invoice 32 days overdue, and what the owner still changes. The output is illustrative:
Subject: Invoice 2291 is now 32 days overdue
Dear Accounts team,
Our records show that invoice 2291 for $640.00, due on 15 September,
remains unpaid. As this is now over 30 days overdue, we have placed
further work on hold until payment is received. We may need to
consider further steps if this is not resolved.
You can pay here: [payment link]. If there's a problem with the
invoice, please call us on [number] and we'll sort it out.
Kind regards,
The office
Two fixes. "We may need to consider further steps" is a veiled threat the prompt told it to avoid, so it comes out. And the placeholders show the AI didn't have the payment link or number, which means the invoice details pasted in were incomplete. The owner adds both to the business context so every chaser has them. Everything else can go as it is, and it took 40 seconds to check.
Keep a short checking list next to each prompt: the three or four things you always verify before sending. If you want a quick, repeatable method for those checks, a five-minute fact-check routine for AI output is worth adapting.
Separate what worked from what only works for you
Go back through the log and pull out every entry where you answered "no" in the last column. These are your silent fixes: the trade discount you always apply, the customers you know pay late because their accounts department runs monthly, the phrase you always delete because it overpromises. Each one has to become one of three things:
- A rule in the prompt ("apply the trade discount for customers tagged Trade in the job system").
- A line in the business context the AI reads (your service hours, what you don't offer).
- A human decision that stays with you, written down as such ("any quote over $2,000 comes to me before sending").
The business-context route is the one owners underuse, so here it is filled in for the locksmith after four weeks of log entries:
BUSINESS CONTEXT (read before every draft)
- Emergency lockouts: 7am to 10pm, seven days. No overnight call-outs.
- We don't fit alarms, CCTV or smart doorbells. Refer those elsewhere.
- Trade customers are tagged "Trade" in the job system: 10% off parts.
- Never quote a price for a lockout by email; say "we'll confirm on
the phone once we know the lock type".
- Sign off as "The office". Phone: [number]. Payment link: [link].
The difference shows in the first line of a typical enquiry reply. Before the context existed, a draft for a late-evening lockout enquiry read "We offer a 24/7 emergency service and can usually be with you within the hour", which the owner deleted every time. After, the same request produced "We cover lockouts until 10pm; if it's later than that, call us from 7am and we'll book you the first slot." One line of context replaced a correction the owner had been making by hand about ten times a week.
This step is where owner-first trials either become useful to a team or quietly remain the owner's private trick. If a lot of what you know turns up here, using AI to capture what only you know takes the idea further.
Write a one-page playbook for each task
PLAYBOOK: [task] OWNER: [you] DATE:
WHEN TO USE IT:
TOOL AND ACCOUNT:
THE PROMPT: (paste the saved version)
WHAT TO PASTE IN / WHAT NEVER TO PASTE IN:
CHECK BEFORE SENDING:
1.
2.
3.
ESCALATE TO [you] IF:
TIME IT SHOULD TAKE: about ___ min (my average was ___)
KNOWN MISTAKES IT MAKES:
Filled in for the locksmith's invoice chasers (from the worked example below), it's short:
PLAYBOOK: Invoice chasers OWNER: owner DATE: 9 Nov
WHEN TO USE IT: Monday and Thursday mornings, from the overdue list
TOOL AND ACCOUNT: Gemini in Gmail, business account
THE PROMPT: "Invoice chasers v3" in the shared prompts doc
WHAT TO PASTE IN: invoice number, amount, due date, customer name
WHAT NEVER TO PASTE IN: job notes, key codes, other invoices
CHECK BEFORE SENDING:
1. Bank feed: has it been paid since the overdue list was run?
2. Amount and invoice number match the accounts system
3. No threats, no mention of other customers
ESCALATE TO owner IF: over 60 days, over $1,500, or the customer
disputes the work
TIME IT SHOULD TAKE: about 2 min (my average was 1.5)
KNOWN MISTAKES IT MAKES: hints at "further steps"; drops the
payment link if it isn't pasted in
The "time it should take" line comes straight from your log and gives the colleague a benchmark. The "known mistakes" line saves them rediscovering your errors one by one. These rules also feed your eventual AI usage policy, so nothing here is wasted.
Read the log for rollout decisions, not just savings
Before you hand anything over, your four weeks of notes can answer most of the questions a rollout raises. Spend half an hour going through them with these in mind:
- Which tool and plan? Did you hit usage limits, need to upload files, or wish the AI could see your calendar or job system? If the built-in assistant in your email did everything, you may not need another subscription at all.
- How many seats? Only the people who do your three tasks need access for this rollout. Everyone else can wait for a task that suits them.
- What are the rules? Your "never" row and your silent fixes are the first draft of your AI rules. You've found them from real work, which makes them far more specific than a template policy.
- How much training? Look at how long it took you to reach stable times. A colleague following your playbook should get there faster, but not instantly; plan for the first two weeks being slower.
- What does success look like? Your average times are the benchmark. A colleague within a minute or two of your numbers, with no new error types, means the task has transferred.
Hand over one task to one person
Pick the task with the fewest silent fixes and the person who either does it now or will do it. Then:
- Do two real instances together, with them driving and you watching.
- They do the next five on their own; you check every one before it's sent.
- Compare their times and errors with your log. If they're close to your numbers, the playbook works. If they're much slower or making errors you didn't, the playbook is missing something you know. Add it and repeat.
- Once five in a row are clean, move to sample checks, then to their normal review.
Expect the first comparison to look worse than it is, and find out why before touching the prompt. In the locksmith's handover, the administrator's first five chasers averaged four minutes each against the owner's 1.5. Watching her do one explained it: she was opening each invoice PDF and retyping the amount, because the overdue list the owner worked from was an export from the accounts system that her login didn't have permission to run. The fix was a permission change, not a better prompt, and her next five averaged two minutes.
A month after the handover, check that it held. Pull ten of the colleague's sent messages and score each against the playbook's checking list: amount right, invoice number right, no threats, payment link present. Nine or ten clean means the task has transferred; seven or fewer means going back to checking every one for a week. Then look at the result the task exists for. For invoice chasers that's the number of invoices more than 30 days overdue on each Monday's list; if it holds steady or falls after the handover, the chasers are still doing their job.
Only then start the second task, or the second person. Rolling out one task at a time feels slow, but each handover exposes gaps in your playbook that would otherwise surface as customer complaints. If staff are hesitant, getting your staff to actually use AI tools covers the people side.
Worked example: a three-van locksmith
Take an illustrative locksmith business: the owner, two locksmiths and a part-time administrator who works three days a week. The owner does about nine hours of admin a week, mostly in the evenings, and uses Google Workspace.
After two weeks of logging, the numbers were clear. Enquiry replies went from about 5 minutes to 2, across 20 a week. Invoice chasers from 4 minutes to around 1.5, across 15. Commercial quotes from 25 minutes to 12, across 5. Together that's roughly 160 minutes a week, about 2.7 hours back from the evenings.
The log also showed the silent fixes. Gemini kept adding "24/7 emergency service" to enquiry replies, which the business doesn't offer after 10pm; that became a line in the business context. The owner was applying trade discounts to commercial quotes from memory; that became a rule tied to a tag in the job system. And quotes involving master-key suites stayed with the owner entirely, with the charts themselves never going near an AI tool.
In week five the administrator took over invoice chasers, matched the owner's times within a week, and caught one thing the owner hadn't: a customer who had already paid by bank transfer, which the chaser prompt couldn't know. The playbook gained a line, "check the bank feed before sending". Enquiry replies followed in week seven. Commercial quotes stayed with the owner, which is a perfectly good outcome: not every task needs handing over.
The whole exercise cost the owner about three hours of logging and prompt refining spread over six weeks, plus two hours writing the playbooks and an hour sitting with the administrator. Against 2.7 hours a week recovered, it had paid for itself before the first handover.
When your own trial says "not yet"
Sometimes the log shows the saving is small or the error rate is uncomfortable. That's a cheap and useful result: a few weeks of your time rather than a failed team rollout. Before giving up, check two things. Was the task a good fit, or did it fail the filter in hindsight? And were you using a stable prompt, or asking differently each time? If both check out and it still isn't worth it, drop the task, pick the next one from your list, and log again.
A "not yet" looks like this in practice. The owner of an illustrative wholesale florist tried AI on her weekly supplier orders: turning a scribbled list of stems into a formatted order email. The log showed about three minutes saved per order, but in the second week it recorded two orders where stem counts had changed (a 50 became 5, and a mixed bunch was split into the wrong varieties). She caught one before sending; the supplier queried the other. Three minutes saved against a chance of a wedding short of roses was an easy call. She dropped the task and moved to customer enquiry replies, where she reads every word anyway and a slip costs an apology rather than a wedding's flowers.
The habit you've built, measuring before rolling out, is worth keeping whatever the result.
Questions from owners going first
I'm a sole trader. Is there anyone to roll it out to?
Your future self, and anyone you hire or subcontract to. The log and playbook still pay off: they stop you drifting back into slow habits, they make it quick to restart after a busy spell, and they become the induction material for a first hire or a virtual assistant. Sole traders can skip the handover stage and move straight to adding a fourth task once the first three are steady.
Should I tell staff I'm trying AI on my own work?
Yes, briefly. Say what you're testing and why, and that you'll share what works. Staff notice anyway, and silence feeds worry about what the owner is planning. Telling them early also invites useful input: the person who usually chases invoices may point out the customers who need a phone call rather than an email, which belongs in the playbook.
Which tool should I start with for my own admin?
The one built into the email and documents you already use. On Google Workspace, that's Gemini in Gmail and Docs; on Microsoft 365, Copilot in Outlook and Word. Starting there means no new logins and no copying between apps. Move to a separate assistant only if the built-in one can't handle a task your log shows is worth automating.
Further reads
- How to Learn AI as a Business Owner: A 30-Day Self-Study Plan — A structured 30-day plan to build your own skills.
- AI Quick Wins: 12 Things a Small Business Can Set Up This Week — More candidate tasks for your first three.
- ChatGPT Prompts for Small Business Owners: 50 Tested Examples — Tested prompts to start your playbook from.
- Is My Business Too Small for AI? What Works for Teams of 1 to 10 — What works for teams of one to ten.
- How to Choose and Support an AI Champion in a Small Team — Pick the colleague who takes over from you.
- How to Build a Shared Prompt Library for Your Team — Where your playbook prompts should live once shared.
- Is a 1:1 AI Consultation Worth It for a Small Business? — The real cost of a 1:1 AI consultation, how many hours it must save to break even, and three cafés that get three different answers.
- What Should a Small Business Automate First With AI? — The six traits of a good first AI automation, the jobs that fit most small firms, what to leave until later, and a one-sitting way to rank your own list.
- AI for Small Business Owners: A Plain-English Beginner's Guide — What an owner should know before starting with AI: how chat tools really work, the four kinds of product, data rules, costs and a first fortnight.
- AI Adoption Stages: Where Is Your Business Now, and What's Next? — Place your business on a five-stage AI adoption scale with a ten-question check, then see the one move that gets you to the next stage.
- Is AI Too Complicated for Non-Technical Business Owners? — Four levels of AI use, from typing into a chat box to custom builds, with honest learning times and five tests for when to bring in help.
- How Much of a Property Manager's Week Can AI Take Over? — A task-by-task breakdown of a 45-hour property management week: which hours AI can take, which stay human, and the order to claim them in.
- How to Choose Your First AI Project: 7 Tests Before You Commit — Put each candidate for your first AI project through seven pass-or-fail tests, then commit to the one that passes all seven, with a one-page note.
- AI Tools and AI Development: The Complete 2026 Guide — the AI hub, including every tutorial in the AI-for-business series.
Sources: Gmail Help, Collaborate with Gemini in Gmail and Summarize email threads; Microsoft Learn, Microsoft 365 Copilot usage report (Outlook features such as Draft with Copilot and Summarize by Copilot).