How to Set a Baseline Before You Introduce AI

Coding Liquids tutorial cover featuring Sagnik Bhattacharya for How to Set a Baseline Before You Introduce AI.
Coding Liquids tutorial cover featuring Sagnik Bhattacharya for How to Set a Baseline Before You Introduce AI.

Before switching on AI, measure the current process for two to four normal weeks: volume, time per item, response time, how often work is redone, and cost. Use records you already have where possible, time a sample of at least 30 items, and write the figures down with dates and definitions so you can compare like with like later.

A baseline is the "before" photograph. Without it, every claim about AI, whether faster, better or cheaper, is a feeling, and feelings about new tools are unreliable in both directions. Which weeks you measure matters as much as what you measure: a baseline taken in a quiet fortnight and compared with a busy month later will make any tool look worse than it is, so skip holiday weeks and note anything unusual.

Follow me on Instagram@sagnikteaches

The effort is modest. For one workflow, expect about an hour to set up and ten minutes a day per person while it runs. That's far less than the time you'd spend arguing later about whether the AI made any difference.

Connect on LinkedInSagnik Bhattacharya

Choose three to five numbers, plus one that mustn't get worse

Measure only what the AI change is meant to move, plus a guardrail. More than five numbers and the measuring becomes the job.

Subscribe on YouTube@codingliquids
NumberDefinition to write downExample for an enquiry inbox
VolumeWhat counts as one itemEach new customer email thread; follow-ups in the same thread don't count
Effort per itemMinutes of hands-on workFrom opening the email to pressing send, excluding interruptions
TurnaroundTime the customer waitsWorking hours from arrival to first real reply (auto-replies don't count)
Redo rateItems needing a second goCustomer had to write back because a question was missed or answered wrongly
CostStaff time valued at loaded costHours a week multiplied by pay plus employment costs
GuardrailSomething that must not get worseComplaints about wrong information given by email

Effort and turnaround are different numbers and both matter. AI usually cuts effort. Turnaround only improves if the saved effort changes how often someone clears the inbox. Measuring both shows you which one actually moved.

The guardrail is the number people forget. It's what stops you celebrating a 60% time saving that came with twice as many complaints. For a longer menu of options, see our list of 12 AI KPIs worth tracking.

The same five slots work for jobs that have nothing to do with an inbox; only the definitions change. For an illustrative bookkeeping practice planning to use AI to read client receipts and suggest how to code them, the set might be:

NumberDefinition for receipt processing
VolumeReceipts and invoices entered per client per month; a multi-page invoice counts once
Effort per itemMinutes from opening the image to saving the coded entry
TurnaroundWorking days from the client's upload to the entry being posted
Redo rateEntries recoded at the month-end review, as a share of all entries
GuardrailQueries sent back to clients because an entry was wrong

Notice the redo rate here comes from the month-end review the practice already does. That's the pattern to look for: a check that happens anyway, turned into a count.

Where most of the numbers already live

You rarely need a stopwatch for everything. Most small-business systems record more than people realise.

SourceWhat it gives youWatch-out
Email timestampsTurnaround: when a message arrived and when you first repliedBatch replying distorts averages; use the median
Till or ticketing systemVolumes, reissued tickets, refunds, express ordersCodes used inconsistently by different staff
Booking systemBookings, cancellations, no-shows, lead timeOnline and phone bookings may be recorded differently
Phone call logCall volume, missed calls, call lengthPersonal mobiles used for work won't appear
Review platformsRating trend and complaint themesToo few reviews a month to show short-term change
Payroll and rotasHours worked and loaded costUnpaid owner time is missing; log it separately
Complaints book or notesErrors that reached customersOnly as good as the habit of writing things down

Pull what you can from these first, then time only the gaps. If the process itself isn't written down, do that before measuring; documenting your processes before adding AI shows a quick way, and it tells you where each step starts and ends, which is what you need to time it.

Timing the work that leaves no trace

Effort per item is usually invisible in any system, so you have to sample it. A simple log, on paper or in a shared spreadsheet, is enough:

Date  | Item ref | Type          | Start | Finish | Mins | Redo? | Notes
3 Jun | E-041    | Stain query   | 09:12 | 09:17  | 5    | No    |
3 Jun | E-042    | Collection    | 09:18 | 09:20  | 2    | No    |
3 Jun | E-043    | Alteration    | 09:31 | 09:53  | 22   | No    | needed tailor's input
3 Jun | E-044    | Complaint     | 10:05 | 10:19  | 14   | Yes   | missed the refund question
  • Sample, don't time everything. Ten items a day per person is plenty, spread across the day rather than the first ten.
  • Record interruptions in the notes, and subtract them only if they're long (more than a couple of minutes).
  • Expect a speed-up at first. People work faster when they know they're being timed. Run the log for at least two weeks and consider ignoring the first two days.
  • Say clearly what it's for. Tell staff the log measures the process, not them, and that nobody's figures will be compared. Otherwise the numbers get quietly improved.

Baselining quality, not only time

Most baselines stop at speed, which leaves you unable to answer the question that matters later: is the AI-assisted work as good as before? Staff work isn't perfect either, and you need to know how good it was to judge whether AI made it better or worse.

During the baseline weeks, pick 20 finished items at random, such as sent replies, completed records or product descriptions, and score each against three or four plain questions:

  • Did it answer everything the customer asked?
  • Were all the facts correct (prices, dates, times, policies)?
  • Did it follow your house rules, such as always giving a ticket number or never promising a same-day service?
  • Would you be happy for it to be read out in front of the customer?

Record the share that pass all four, and a few words on each failure, because the reasons matter as much as the count. For the dry cleaner in the worked example below, the three failures in its sample of 20 read:

  • "Asked about both a wool coat and a silk scarf; reply only covered the coat." (Missed a question.)
  • "Said collection from 8am; the second branch opens at 8.30." (Wrong fact.)
  • "Offered same-day service on a wedding dress." (Broke a house rule: wedding dresses are never same-day.)

Those three tell you what the AI's instructions must cover from day one: both branches' hours, the no-same-day list, and a check that every question in the email has an answer. If 17 of 20 pass, your quality baseline is 85%, and that's the figure the AI-assisted process has to match or beat, scored the same way by the same person. It's often a humbling exercise. Many owners find their own team's replies miss a question more often than they'd assumed, which is useful to know before blaming, or crediting, the AI for anything.

How long to measure, and how many items

A few working rules keep the baseline trustworthy without dragging on:

  • At least two full weeks, so every day of the week appears twice, including your busiest.
  • At least 30 timed items per job type, and 50 or more if the job varies a lot, such as quotes that range from simple to complicated.
  • Note anything unusual: staff absence, a promotion, a public holiday, a one-off bulk order. Don't delete those weeks; label them.
  • Mind the season. If your business has a busy season, compare like with like later. For volumes, last year's records for the same weeks are a useful cross-check.
  • Use the median for time, meaning the middle value when you line them all up. One 40-minute monster email can drag an average up a lot, but it barely moves the median.

The median point is worth seeing in numbers. Say ten timed replies took 2, 3, 3, 4, 4, 5, 5, 6, 7 and 40 minutes. The average is 7.9 minutes; the median is 4.5. If the 40-minute email was a one-off dispute, the average claims every reply takes nearly eight minutes, and any AI tool will later "save" three minutes a reply that never existed. Record both figures, but set targets against the median, and keep the long ones in the log with a note: they're often the emails you'd never hand to AI anyway.

Seasonal businesses need one more adjustment. An illustrative wedding florist gets most of its enquiries in the first three months of the year, but wants to set up AI drafting in the autumn. An autumn baseline would show a handful of enquiries and a quick turnaround that tells you nothing about January. The practical split: take effort per enquiry and quality from the autumn sample, since those change little with the season, and take volume and turnaround from last year's inbox for the same weeks you'll compare against. Write on the sheet which figure came from where.

A worked example: a dry cleaner's enquiry inbox

Take an illustrative dry cleaner with two branches and one shared email inbox. It plans to use AI to sort incoming emails and draft replies to routine ones. Before switching anything on, the manager ran a three-week baseline from 2 to 20 June.

  • Volume: 112 new email threads, about 37 a week, pulled from the inbox.
  • Effort: 45 timed replies. Median 5 minutes; range 1 to 22 minutes. Alteration quotes were the slowest because they often needed the tailor.
  • Turnaround: median 19 working hours from arrival to first reply, taken from email timestamps. 18% took longer than two working days.
  • Redo rate: 6 of 112 threads (5%) needed a second reply because the first missed a question.
  • Guardrail: 2 complaints about wrong information by email in three weeks, from the complaints notebook.
  • Cost: 37 emails × 5 minutes is about 3.1 hours a week; at a loaded $22 an hour, roughly $68 a week.

Two notes went on the sheet: one member of staff was off for part of week two, and a single bulk enquiry about a uniform-cleaning contract was excluded as a one-off. Those notes are what make the comparison fair eight weeks later, when the same definitions are applied to the AI-assisted process.

Because the business has two branches, it can also do something smaller firms can't: start the AI at one branch first and keep the other as a comparison for a month. That separates the effect of AI from a generally quieter or busier period.

Here's why that helps. Suppose that, a month later, the branch using AI shows a median turnaround of 7 working hours against its baseline of 19. Impressive, until you see that the comparison branch, with no AI at all, has also dropped from 19 to 14, because July was quieter and both managers had more time. The fair reading is that AI is responsible for roughly the gap between the two branches (about 7 hours), not the whole 12-hour improvement. Without the second branch, the business would have credited the tool with a quiet month.

Recording the baseline so it survives

Numbers in someone's head or in a scratch file disappear. Put them on one sheet, then save a copy that nobody edits:

BASELINE RECORD
Process:             Customer email enquiries (both branches)
Dates measured:      2-20 June (3 weeks)
Measured by:         Branch manager; timing log kept by 2 staff
Definitions:         Item = new thread. Effort = open to send.
                     Turnaround = working hours to first real reply.
                     Redo = customer had to write again for the same issue.
Results:             Volume 37/week | Effort median 5 min (range 1-22)
                     Turnaround median 19 working hrs | Redo 5%
                     Complaints about wrong info: 2 in 3 weeks
                     Cost about $68/week at $22/hr loaded
Unusual events:      Staff absence week 2; uniform contract enquiry excluded
Raw data:            Shared drive > AI > Baseline June (read-only)
Agreed by:           Owner, 21 June

The definitions line is the most important part. If "turnaround" later gets measured from a different starting point, the comparison is meaningless, however careful the rest was. An illustrative physiotherapy clinic found this out the awkward way. Its receptionist measured the baseline turnaround to the first reply of any kind, which included the automatic "thanks, we'll be in touch" message, so the median was a few minutes. Eight weeks after AI drafting went in, the practice manager measured to the first real answer and got four working hours. On paper, AI had made replies dozens of times slower. Nothing had changed except the definition, and the only fix was to rebuild the baseline from old email timestamps using the right one. When you come to compare, the same sheet feeds straight into an AI ROI calculation.

If you've already started without a baseline

It's common, and it's recoverable. In rough order of reliability:

  1. Reconstruct from records. Email timestamps, tickets and bookings from before the switch still exist. You can usually rebuild volume and turnaround, if not effort.
  2. Compare two groups. If some staff or branches use AI and others don't yet, compare them over the same weeks.
  3. Run a switch-off week. Do the job the old way for one week and time it. Only do this where going back is safe and staff won't resent it.
  4. Ask for estimates, and discount them. Staff memories of "how long it used to take" tend to run high. Treat them as an upper limit, not a figure.

Reconstructing turnaround is less work than it sounds. Search the inbox for the month before AI arrived, then open every fifth new thread until you have 40. For each, note the time the customer's first message arrived and the time of your first real reply, and convert the gap into working hours in a spreadsheet column. Forty threads take about an hour. You won't get effort per item this way, because nobody timed the writing, but you'll have a genuine turnaround and redo figure, and redo is easy to spot: it's any thread where the customer had to write again about the same thing.

Whichever you use, write down how the baseline was obtained. A reconstructed baseline is fine if everyone knows it's reconstructed. For more on the time side specifically, see how to measure time saved after rolling out AI, and if the AI will touch customers directly, add one customer-facing measure using how to measure customer reaction after introducing AI. The baseline also forms stage three of a first AI pilot project, if you're running one.

Further reads

Want help measuring before the AI goes in?

On a 1:1 call we'll pick the numbers that matter for your workflow, find where your systems already record them, and set up a two-week baseline your team can run without fuss.

Book a 1:1 call with me