Measure a handful of named tasks, not "AI use". Time each task for two weeks before rollout, time the same tasks for two to four weeks after, and subtract the minutes spent checking and fixing AI output. Net minutes saved per task, multiplied by how often the task happens each month, gives a time-saved figure you can defend.
Two shortcuts overstate the result, and most firms take one of them. Asking staff "how much time does AI save you?" produces a memory of the best days, not an average. Vendor dashboards estimate rather than measure: Microsoft's "Copilot assisted hours", for example, is built from the actions people take in Copilot multiplied by time-saving factors from Microsoft's own research, not from your team's clocks. Both are useful signals of adoption. Neither tells you how many hours came back.
Pick five tasks and define where each starts and stops
Choose tasks that happen often enough to time, that AI is actually meant to change, and that have a clear start and finish. Five is plenty for a small firm. An illustrative set for a community interest company with seven staff running youth and community programmes:
| Task | Starts when | Stops when | Unit | Per month |
|---|---|---|---|---|
| Funder progress report | Template opened | Sent to manager for sign-off | Per report | 4 |
| Meeting minutes | Meeting ends | Minutes circulated | Per meeting | 6 |
| Grant application first draft | Funder guidance read | Complete first draft saved | Per application | 2 |
| Programme enquiry reply | Email opened | Reply sent | Per reply | 120 |
| Volunteer rota and reminders | Availability collected | Reminders sent | Per week | 4 |
The start and stop points are the whole trick. Without them, one person times the grant draft from "reading the guidance" and another from "starting to type", and the numbers can't be compared. Define them once and put the table where everyone recording can see it. If you haven't done this before any rollout, setting a baseline before you introduce AI covers the groundwork; this tutorial picks up at the point where you're measuring a rollout that's happening now.
The CIC's first draft of the list shows the kind of task that can't be measured. It read "emails", "admin", "writing" and "research". None has a start or a stop, and "emails" covers everything from a two-line confirmation to a complaint that takes an hour. Rewritten, "emails" became "programme enquiry reply", a single kind of email with a recognisable shape; "writing" split into the funder report and the grant draft; "admin" became the rota, because that was the admin AI was actually going to touch. "Research" was dropped altogether. Nobody could say when a piece of research was finished, and a task you can't finish is a task you can't time.
Two weeks of honest baseline timing
Record each instance of each task in a shared sheet: date, person, task, minutes, and a one-word quality note. A phone timer or the clock on the wall is fine; precision to the minute matters less than consistency. You don't need every instance of high-volume tasks. For enquiry replies, time ten a week per person rather than all 120 a month.
A few lines of an illustrative log:
date,person,task,minutes,phase,notes
03/03,PO,Funder report,195,baseline,waiting on attendance figures
04/03,KM,Minutes,70,baseline,board meeting
04/03,KM,Enquiry reply,9,baseline,
05/03,PO,Enquiry reply,6,baseline,
06/03,JS,Grant draft,510,baseline,new funder, long form
07/03,JS,Rota,85,baseline,
If a task happens only once in the baseline fortnight, note it as a rough figure and don't lean on it. Two weeks catches most weekly and monthly work; stretch to four if a key task is monthly.
Tell people why they're timing and that the numbers aren't about them. Being timed changes how people work, and staff who suspect the figures will be used to cut hours will, reasonably, slow down. The CIC's manager sent something close to this before the baseline began:
Subject: Timing five tasks for the next two weeks
We're trying out AI tools for reports, minutes, grant drafts,
enquiry replies and the rota. Before we do, I'd like to know how
long these take now, so we can tell later whether the tools help.
- Log each one in the shared sheet: date, your initials, task,
minutes. Rough minutes are fine.
- Start and stop points are on the first tab. Please stick to them.
- Busy days count. Please log the slow ones too.
- The figures are about the tasks, not about anyone's speed.
Nobody's hours will change because of this exercise.
I'll share the results with everyone at the end of May.
The line about busy days earns its place. Without it, people tend to skip logging on exactly the days a task drags, and the baseline comes out faster than reality.
After rollout: time the same tasks, including checking and fixing
Give people two weeks to settle in before timing starts again; the first days with any new tool are slower. Then log the same tasks with the same start and stop points, split into three columns:
- Working time with AI: prompting, reading, editing.
- Review time: whoever checks the output, including a manager's sign-off if that now takes longer.
- Rework: fixing anything that went out wrong or came back with questions.
Review and rework are where most measured savings shrink. A report that takes 70 minutes to draft instead of 180 but needs 30 minutes more manager checking hasn't saved 110 minutes. Checking AI is doing good work, not just fast work covers the quality side of this in more detail.
Automated tasks need their own treatment, because nobody is there to time them. At the CIC, suppose the volunteer reminders move to an automation that texts each volunteer two days before their shift. Working time for sending reminders drops to zero, but the task hasn't vanished: someone still spends about 15 minutes a week handling replies such as "can't make it, can anyone swap?", and the automation needs checking when the rota changes. Log those minutes under the same task name. Log failures too. If the automation stops one week because a spreadsheet column was renamed and the fix takes two hours, those two hours belong in the rework column for that month, not in a footnote.
The calculation, worked through
For each task: net minutes saved per instance = baseline minutes minus (working + review + rework). Multiply by instances per month. Then subtract the ongoing overhead (maintaining prompts, admin, fixing broken automations) and the set-up time spread over a year. For the illustrative CIC:
| Task | Baseline | After (work + review + rework) | Saved each | Per month | Saved per month |
|---|---|---|---|---|---|
| Funder report | 180 min | 70 + 20 + 10 = 100 | 80 min | 4 | 320 min |
| Minutes | 75 min | 25 + 10 + 5 = 40 | 35 min | 6 | 210 min |
| Grant first draft | 480 min | 200 + 60 + 40 = 300 | 180 min | 2 | 360 min |
| Enquiry reply | 8 min | 4 + 1 + 0.5 = 5.5 | 2.5 min | 120 | 300 min |
| Rota and reminders | 90 min | 60 + 5 + 0 = 65 | 25 min | 4 | 100 min |
That's 1,290 minutes, or 21.5 hours a month, gross. Subtract about 2 hours a month of ongoing overhead and 1 hour a month for the 12 hours of set-up spread over a year, and the defensible figure is roughly 18.5 hours a month across seven people: about two and a half hours each. That's less exciting than the "a day a week" people had guessed, and far more useful, because it's true. To turn hours into money, the method is in calculating AI ROI with a worked example.
Using AI to summarise the log, carefully
AI can tidy the log and draft the summary, but check its arithmetic, because averages across a table are where models slip. A prompt:
Attached is our timing log (CSV). For each task, give:
- number of baseline entries and after entries
- average baseline minutes
- average after minutes, as working + review + rework
- net minutes saved per instance
Exclude entries marked "settling-in". Show your working as a table
so I can check it. Do not estimate missing values.
A line from a sample reply (illustrative): "Funder report: baseline average 180 min (4 entries); after average 70 min; saving 110 min per report." The mistake is plain once you look: it used the working-time column only and ignored review and rework, so the saving is overstated by 30 minutes a report. The right figure is 80. Ask for the working as a table every time, and recalculate one task by hand or with a spreadsheet formula before you trust the rest.
If the log's columns run date, person, task, minutes, phase, notes (A to F), with the three after-rollout components logged as separate rows marked "after-work", "after-review" and "after-rework", two formulas settle the funder report in seconds:
Baseline average:
=AVERAGEIFS(D:D, C:C, "Funder report", E:E, "baseline")
After total per report (work + review + rework, 4 reports logged):
=SUMIFS(D:D, C:C, "Funder report", E:E, "after*") / 4
If the spreadsheet says 100 and the AI summary says 70, the summary missed two of the three columns. That one disagreement is enough reason to recheck every other line it produced.
Checks before you believe the number
- Did quality hold? A realistic example: in the CIC, enquiry replies fell from 8 minutes to 5.5, but the share of enquirers who wrote back with a follow-up question rose from about 1 in 10 to 1 in 5, because the AI drafts left out session times. Each follow-up cost another reply. Once the prompt required times and a booking link, follow-ups dropped back and the saving held.
- Did the task mix change? If the after period had two short grant applications and the baseline had two long ones, the saving is partly luck. Note the size or type of each instance.
- Did work move rather than shrink? If drafting got faster but a manager now spends longer reviewing, count the manager's time. It's in the review column for this reason.
- Is it still the learning curve? Re-time once more at three months. Savings often grow as prompts improve, and occasionally shrink as the novelty wears off and checking gets skipped.
- Did volume change? More enquiries in the after period make the total look bigger without the per-reply saving changing. Report per instance as well as per month.
- Are the after entries complete? Count entries per person per week in each phase. In an illustrative third week after rollout, enquiry-reply entries fell from about ten per person to six. The missing ones were the awkward enquiries, the ones that took longest and that people didn't stop to log. The average looked better than it was. Chase the gaps, or compare only weeks where entry counts match the baseline.
The task-mix point deserves a quick sum, because it can reverse a result. The baseline grant draft took 510 minutes for a long-form application of about 3,000 words: 170 minutes per 1,000 words. Say the after period's application was a short form of 1,500 words drafted in 300 minutes, including review and rework: 200 minutes per 1,000 words. On raw minutes, AI saved over three hours. Per 1,000 words, the new process was slower. For tasks whose size varies this much, record a size measure alongside the minutes (word count, number of funder questions, pages of minutes) and compare the rate, not the total.
What vendor dashboards can and can't tell you
If you're on Microsoft 365, the admin centre's Copilot usage report shows how many people used Copilot over the last 7, 28, 90 or 180 days, and in which apps. That answers "are people using it?", which is worth knowing: a tool nobody opens saves nobody time.
The Copilot Dashboard in Viva Insights goes further, with an estimate called Copilot assisted hours and a value calculator that multiplies those hours by an hourly rate (it starts at a default of $72 an hour, which you can change). Microsoft describes the assisted-hours figure as an estimate built from Copilot actions and research-based multipliers; Microsoft's Copilot Analytics documentation explains the tools. Two cautions for small firms: which features and metrics you see depends on how many licences you have, and minimum group-size privacy settings can hide figures for small teams. Treat any "hours saved" number from a vendor as a claim to test against your own log, not a result. Measuring whether Copilot is paying for itself goes into the Copilot-specific version.
Other business chat plans have admin usage views too. The same rule applies: usage is evidence of adoption, and your timing log is the evidence of time saved.
Where the saved time actually went
An owner's real question is usually not "how many hours?" but "what did we get for them?" Saved minutes scattered across a week often disappear into email. Ask each person, at the three-month review, what they did with the time, and look for signs in the business: a backlog cleared, fewer late evenings, more enquiries answered the same day, a second funding bid that wouldn't otherwise have been written. Why AI often saves time but not money explains why this step decides whether a rollout pays off, and AI KPIs worth tracking suggests outcome measures to pair with the hours.
Illustrative answers from the CIC's three-month review show why the question is worth asking person by person:
- Programme lead: "Started the second funding bid in April instead of July." Traceable: the application exists, with a date on it.
- Office coordinator: "Enquiries now get a same-day reply most days." Checkable: the shared inbox shows reply times, and the backlog of unanswered enquiries that used to build up on Mondays has gone.
- Youth worker: "Not sure, probably email." Honest, and typical. The half hour a week saved on minutes has been absorbed. Nothing is wrong, but it isn't a gain the business can point to yet.
The third answer is the prompt for a decision rather than a criticism: either name a job the freed time should go to, or accept that this saving is mainly a quality-of-life improvement and weigh the tool's cost with that in mind.
Don't reduce anyone's hours on the strength of an estimate. If the figures hold for three months and the work is genuinely smaller, that's a conversation to have openly, with the log in front of you.
Scaled down: a church office with one administrator
The method scales down. In one illustrative church office, the part-time administrator measured three tasks: the weekly notice sheet (90 minutes before, 35 after including the vicar's check), hall-booking replies (about 40 a month, 6 minutes each before and 3 after), and the monthly magazine's first layout of submitted items (4 hours before, 2.5 after). Net of about 30 minutes a month spent keeping the prompts current, that's roughly 6.5 hours a month, most of which went into the volunteer coordination that had been falling behind.
The report to the church council fitted on one page:
AI time review: March to May
Tasks measured: notice sheet, hall bookings, magazine layout
Hours saved per month (net of checking and upkeep): about 6.5
Quality: no increase in corrections to notices; two booking replies
needed a follow-up in May (prompt fixed)
Tool cost: included in our existing office suite subscription
Time redeployed to: volunteer rota and new-member welcome letters
Recommendation: continue; re-measure in September
That's the whole deliverable: a small number, honestly measured, with a note of where the time went and a date to check again.
Further reads
- How to Review an AI Tool After 90 Days: Keep, Fix or Cancel — Use your time figures to keep, fix or cancel each tool.
- How Much Time Can AI Save a Solo Consultant Each Week? — The same question answered for a one-person business.
- Turning Hours Saved by AI Into Advisory Work Clients Pay For — How time savings look in advisory and client work.
- How to Survey Your Staff Before an AI Rollout (With Questions) — Pair timing with what staff think before and after.
- How to Roll Out Microsoft 365 Copilot to a Small Team — A rollout plan that builds measurement in from day one.
- How Soon Should AI Pay for Itself? Payback Periods by Project — Turn hours saved into a payback period you can plan around.
- Should You Charge Clients Less When AI Speeds Up Your Work? — What to do about your prices when AI cuts the time a job takes: the rules for hourly billing, when fixed fees are fair, and how to tell clients.
- How Much Time Can AI Really Save a Small Business Each Week? — Realistic weekly time savings from AI: what research found, per-task estimates, a seven-person accountancy practice added up, and a two-week measuring method.
- Best AI Tools for Etsy Sellers, Ranked by Time Saved — Seven kinds of AI tool ranked by the hours they save a small Etsy shop each week, with prices, a worked example for each, and the tools not worth paying for.
- Which Event Planning Tasks Should You Hand to AI First? — Twelve event planning tasks scored for AI, the five to hand over first with worked examples, and the ones to keep even though AI will attempt them.
- How Much Time Can AI Save a Barber Each Day? — A minute-by-minute look at a barber's day: which admin AI shortens, which it can't touch, and a two-day tally to measure your own saving.
- What an AI Implementation Looks Like in a Small Accounting Firm — A seven-person practice followed through 12 weeks of AI implementation: time audit, tools switched on, records chasing, costs and what went wrong.
- AI Implementation Plan for a Small Law Firm: The First 90 Days — A day-by-day 90-day AI plan for a six-person law firm: policy first, two scored pilots, shadow testing, an error log and a go/no-go decision.
- AI Rollout Plan for a Ten-Person Marketing Agency — A ten-person agency's 12-week AI rollout, week by week: platform choice, client rules, two pilots, the numbers after a quarter and what went wrong.
- Use the AI Already in Your Practice Software Before Buying More — Five stages to find and test the AI in your practice management, accounting and office software, then buy only for the gaps it leaves.
- How to Chase Missing Client Records With AI Before Deadlines — Build a missing-items list, send AI-written chasers on a ladder counted back from each deadline, and let AI sort the replies for you.
- AI Tools and AI Development: The Complete 2026 Guide — the AI hub, including every tutorial in the AI-for-business series.
Sources: Microsoft Learn (Copilot Analytics introduction; Copilot Dashboard assisted hours and value calculator); Microsoft 365 admin centre Copilot usage report documentation.