Before switching on AI, measure the current process for two to four normal weeks: volume, time per item, response time, how often work is redone, and cost. Use records you already have where possible, time a sample of at least 30 items, and write the figures down with dates and definitions so you can compare like with like later.
A baseline is the "before" photograph. Without it, every claim about AI, whether faster, better or cheaper, is a feeling, and feelings about new tools are unreliable in both directions. Which weeks you measure matters as much as what you measure: a baseline taken in a quiet fortnight and compared with a busy month later will make any tool look worse than it is, so skip holiday weeks and note anything unusual.
The effort is modest. For one workflow, expect about an hour to set up and ten minutes a day per person while it runs. That's far less than the time you'd spend arguing later about whether the AI made any difference.
Choose three to five numbers, plus one that mustn't get worse
Measure only what the AI change is meant to move, plus a guardrail. More than five numbers and the measuring becomes the job.
| Number | Definition to write down | Example for an enquiry inbox |
|---|---|---|
| Volume | What counts as one item | Each new customer email thread; follow-ups in the same thread don't count |
| Effort per item | Minutes of hands-on work | From opening the email to pressing send, excluding interruptions |
| Turnaround | Time the customer waits | Working hours from arrival to first real reply (auto-replies don't count) |
| Redo rate | Items needing a second go | Customer had to write back because a question was missed or answered wrongly |
| Cost | Staff time valued at loaded cost | Hours a week multiplied by pay plus employment costs |
| Guardrail | Something that must not get worse | Complaints about wrong information given by email |
Effort and turnaround are different numbers and both matter. AI usually cuts effort. Turnaround only improves if the saved effort changes how often someone clears the inbox. Measuring both shows you which one actually moved.
The guardrail is the number people forget. It's what stops you celebrating a 60% time saving that came with twice as many complaints. For a longer menu of options, see our list of 12 AI KPIs worth tracking.
The same five slots work for jobs that have nothing to do with an inbox; only the definitions change. For an illustrative bookkeeping practice planning to use AI to read client receipts and suggest how to code them, the set might be:
| Number | Definition for receipt processing |
|---|---|
| Volume | Receipts and invoices entered per client per month; a multi-page invoice counts once |
| Effort per item | Minutes from opening the image to saving the coded entry |
| Turnaround | Working days from the client's upload to the entry being posted |
| Redo rate | Entries recoded at the month-end review, as a share of all entries |
| Guardrail | Queries sent back to clients because an entry was wrong |
Notice the redo rate here comes from the month-end review the practice already does. That's the pattern to look for: a check that happens anyway, turned into a count.
Where most of the numbers already live
You rarely need a stopwatch for everything. Most small-business systems record more than people realise.
| Source | What it gives you | Watch-out |
|---|---|---|
| Email timestamps | Turnaround: when a message arrived and when you first replied | Batch replying distorts averages; use the median |
| Till or ticketing system | Volumes, reissued tickets, refunds, express orders | Codes used inconsistently by different staff |
| Booking system | Bookings, cancellations, no-shows, lead time | Online and phone bookings may be recorded differently |
| Phone call log | Call volume, missed calls, call length | Personal mobiles used for work won't appear |
| Review platforms | Rating trend and complaint themes | Too few reviews a month to show short-term change |
| Payroll and rotas | Hours worked and loaded cost | Unpaid owner time is missing; log it separately |
| Complaints book or notes | Errors that reached customers | Only as good as the habit of writing things down |
Pull what you can from these first, then time only the gaps. If the process itself isn't written down, do that before measuring; documenting your processes before adding AI shows a quick way, and it tells you where each step starts and ends, which is what you need to time it.
Timing the work that leaves no trace
Effort per item is usually invisible in any system, so you have to sample it. A simple log, on paper or in a shared spreadsheet, is enough:
Date | Item ref | Type | Start | Finish | Mins | Redo? | Notes
3 Jun | E-041 | Stain query | 09:12 | 09:17 | 5 | No |
3 Jun | E-042 | Collection | 09:18 | 09:20 | 2 | No |
3 Jun | E-043 | Alteration | 09:31 | 09:53 | 22 | No | needed tailor's input
3 Jun | E-044 | Complaint | 10:05 | 10:19 | 14 | Yes | missed the refund question
- Sample, don't time everything. Ten items a day per person is plenty, spread across the day rather than the first ten.
- Record interruptions in the notes, and subtract them only if they're long (more than a couple of minutes).
- Expect a speed-up at first. People work faster when they know they're being timed. Run the log for at least two weeks and consider ignoring the first two days.
- Say clearly what it's for. Tell staff the log measures the process, not them, and that nobody's figures will be compared. Otherwise the numbers get quietly improved.
Baselining quality, not only time
Most baselines stop at speed, which leaves you unable to answer the question that matters later: is the AI-assisted work as good as before? Staff work isn't perfect either, and you need to know how good it was to judge whether AI made it better or worse.
During the baseline weeks, pick 20 finished items at random, such as sent replies, completed records or product descriptions, and score each against three or four plain questions:
- Did it answer everything the customer asked?
- Were all the facts correct (prices, dates, times, policies)?
- Did it follow your house rules, such as always giving a ticket number or never promising a same-day service?
- Would you be happy for it to be read out in front of the customer?
Record the share that pass all four, and a few words on each failure, because the reasons matter as much as the count. For the dry cleaner in the worked example below, the three failures in its sample of 20 read:
- "Asked about both a wool coat and a silk scarf; reply only covered the coat." (Missed a question.)
- "Said collection from 8am; the second branch opens at 8.30." (Wrong fact.)
- "Offered same-day service on a wedding dress." (Broke a house rule: wedding dresses are never same-day.)
Those three tell you what the AI's instructions must cover from day one: both branches' hours, the no-same-day list, and a check that every question in the email has an answer. If 17 of 20 pass, your quality baseline is 85%, and that's the figure the AI-assisted process has to match or beat, scored the same way by the same person. It's often a humbling exercise. Many owners find their own team's replies miss a question more often than they'd assumed, which is useful to know before blaming, or crediting, the AI for anything.
How long to measure, and how many items
A few working rules keep the baseline trustworthy without dragging on:
- At least two full weeks, so every day of the week appears twice, including your busiest.
- At least 30 timed items per job type, and 50 or more if the job varies a lot, such as quotes that range from simple to complicated.
- Note anything unusual: staff absence, a promotion, a public holiday, a one-off bulk order. Don't delete those weeks; label them.
- Mind the season. If your business has a busy season, compare like with like later. For volumes, last year's records for the same weeks are a useful cross-check.
- Use the median for time, meaning the middle value when you line them all up. One 40-minute monster email can drag an average up a lot, but it barely moves the median.
The median point is worth seeing in numbers. Say ten timed replies took 2, 3, 3, 4, 4, 5, 5, 6, 7 and 40 minutes. The average is 7.9 minutes; the median is 4.5. If the 40-minute email was a one-off dispute, the average claims every reply takes nearly eight minutes, and any AI tool will later "save" three minutes a reply that never existed. Record both figures, but set targets against the median, and keep the long ones in the log with a note: they're often the emails you'd never hand to AI anyway.
Seasonal businesses need one more adjustment. An illustrative wedding florist gets most of its enquiries in the first three months of the year, but wants to set up AI drafting in the autumn. An autumn baseline would show a handful of enquiries and a quick turnaround that tells you nothing about January. The practical split: take effort per enquiry and quality from the autumn sample, since those change little with the season, and take volume and turnaround from last year's inbox for the same weeks you'll compare against. Write on the sheet which figure came from where.
A worked example: a dry cleaner's enquiry inbox
Take an illustrative dry cleaner with two branches and one shared email inbox. It plans to use AI to sort incoming emails and draft replies to routine ones. Before switching anything on, the manager ran a three-week baseline from 2 to 20 June.
- Volume: 112 new email threads, about 37 a week, pulled from the inbox.
- Effort: 45 timed replies. Median 5 minutes; range 1 to 22 minutes. Alteration quotes were the slowest because they often needed the tailor.
- Turnaround: median 19 working hours from arrival to first reply, taken from email timestamps. 18% took longer than two working days.
- Redo rate: 6 of 112 threads (5%) needed a second reply because the first missed a question.
- Guardrail: 2 complaints about wrong information by email in three weeks, from the complaints notebook.
- Cost: 37 emails × 5 minutes is about 3.1 hours a week; at a loaded $22 an hour, roughly $68 a week.
Two notes went on the sheet: one member of staff was off for part of week two, and a single bulk enquiry about a uniform-cleaning contract was excluded as a one-off. Those notes are what make the comparison fair eight weeks later, when the same definitions are applied to the AI-assisted process.
Because the business has two branches, it can also do something smaller firms can't: start the AI at one branch first and keep the other as a comparison for a month. That separates the effect of AI from a generally quieter or busier period.
Here's why that helps. Suppose that, a month later, the branch using AI shows a median turnaround of 7 working hours against its baseline of 19. Impressive, until you see that the comparison branch, with no AI at all, has also dropped from 19 to 14, because July was quieter and both managers had more time. The fair reading is that AI is responsible for roughly the gap between the two branches (about 7 hours), not the whole 12-hour improvement. Without the second branch, the business would have credited the tool with a quiet month.
Recording the baseline so it survives
Numbers in someone's head or in a scratch file disappear. Put them on one sheet, then save a copy that nobody edits:
BASELINE RECORD
Process: Customer email enquiries (both branches)
Dates measured: 2-20 June (3 weeks)
Measured by: Branch manager; timing log kept by 2 staff
Definitions: Item = new thread. Effort = open to send.
Turnaround = working hours to first real reply.
Redo = customer had to write again for the same issue.
Results: Volume 37/week | Effort median 5 min (range 1-22)
Turnaround median 19 working hrs | Redo 5%
Complaints about wrong info: 2 in 3 weeks
Cost about $68/week at $22/hr loaded
Unusual events: Staff absence week 2; uniform contract enquiry excluded
Raw data: Shared drive > AI > Baseline June (read-only)
Agreed by: Owner, 21 June
The definitions line is the most important part. If "turnaround" later gets measured from a different starting point, the comparison is meaningless, however careful the rest was. An illustrative physiotherapy clinic found this out the awkward way. Its receptionist measured the baseline turnaround to the first reply of any kind, which included the automatic "thanks, we'll be in touch" message, so the median was a few minutes. Eight weeks after AI drafting went in, the practice manager measured to the first real answer and got four working hours. On paper, AI had made replies dozens of times slower. Nothing had changed except the definition, and the only fix was to rebuild the baseline from old email timestamps using the right one. When you come to compare, the same sheet feeds straight into an AI ROI calculation.
If you've already started without a baseline
It's common, and it's recoverable. In rough order of reliability:
- Reconstruct from records. Email timestamps, tickets and bookings from before the switch still exist. You can usually rebuild volume and turnaround, if not effort.
- Compare two groups. If some staff or branches use AI and others don't yet, compare them over the same weeks.
- Run a switch-off week. Do the job the old way for one week and time it. Only do this where going back is safe and staff won't resent it.
- Ask for estimates, and discount them. Staff memories of "how long it used to take" tend to run high. Treat them as an upper limit, not a figure.
Reconstructing turnaround is less work than it sounds. Search the inbox for the month before AI arrived, then open every fifth new thread until you have 40. For each, note the time the customer's first message arrived and the time of your first real reply, and convert the gap into working hours in a spreadsheet column. Forty threads take about an hour. You won't get effort per item this way, because nobody timed the writing, but you'll have a genuine turnaround and redo figure, and redo is easy to spot: it's any thread where the customer had to write again about the same thing.
Whichever you use, write down how the baseline was obtained. A reconstructed baseline is fine if everyone knows it's reconstructed. For more on the time side specifically, see how to measure time saved after rolling out AI, and if the AI will touch customers directly, add one customer-facing measure using how to measure customer reaction after introducing AI. The baseline also forms stage three of a first AI pilot project, if you're running one.
Further reads
- Did Your AI Pilot Work? How to Set Success Criteria That Hold Up — Turn your baseline into pass and fail thresholds.
- How to Map a Business Process Before You Automate It — Map the steps so you know what you're timing.
- How to Check AI Is Doing Good Work, Not Just Fast Work — Keeps the after picture honest about quality.
- How to Build a KPI Dashboard With AI When You Have No Data Team — Put baseline and current numbers on one screen.
- How to Review an AI Tool After 90 Days: Keep, Fix or Cancel — Where the before and after comparison gets used.
- What Should a Small Business Automate First With AI? — Choose the job to baseline in the first place.
- Is a 1:1 AI Consultation Worth It for a Small Business? — The real cost of a 1:1 AI consultation, how many hours it must save to break even, and three cafés that get three different answers.
- How to Get Started With AI in Your Small Business: First 7 Steps — Seven steps that take about a month, from a five-day time log to a keep-or-drop decision, with a yoga studio's numbers at each stage.
- Example AI Roadmap for a 12-Person Business, Month by Month — A 12-person driving school's first AI year, month by month: five workflows, about $30 a month in new tools, one idea postponed and why.
- Is AI Worth It for a Small Business? How to Work Out Your Answer — A three-number test for whether AI pays in your business, a break-even table, the costs the quick maths misses, and a two-week trial to settle it.
- AI Use Case Template: Score Every Idea on One Page — A one-page template with scoring anchors, knock-out questions and a worked veterinary example for ranking AI ideas before you spend anything.
- How to Find AI Savings Line by Line in Your Profit and Loss — A line-by-line worksheet for finding real AI savings in your P&L, separating cash you can bank from hours that quietly evaporate.
- What to Do Before You Buy Any AI Tool: A 10-Point Checklist — Ten checks to run before paying for any AI tool, from pricing the job it replaces to reading the exit terms, with a kitchen-fitter worked example.
- How Much Time Can AI Really Save a Small Business Each Week? — Realistic weekly time savings from AI: what research found, per-task estimates, a seven-person accountancy practice added up, and a two-week measuring method.
- What AI Implementation Actually Involves for a Small Business — The six pieces of work behind a small-business AI implementation, with rough hours for each and one workflow followed from baseline to handover.
- Why AI Projects Fail in Small Businesses (It's Rarely the Tech) — The seven organisational reasons small-business AI projects fail, what each looks like by week three, and an eight-question check to run before you start.
- AI Tools and AI Development: The Complete 2026 Guide — the AI hub, including every tutorial in the AI-for-business series.