Measure it against a baseline you record before go-live: total calls, missed calls, callers who rang back anyway, and staff time on the phone. Then track eight numbers weekly, count only bookings from calls that would otherwise have been missed, and at day 90 compare the margin they earned plus staff time saved, minus errors, with what the receptionist costs.
The step most owners skip is the baseline, and without it any result is a guess. The 90 days fall into four stages, each with the numbers to collect, followed by a worked example for a butcher's shop that takes phone orders and a check on how far the answer moves if one assumption is wrong.
Days −14 to 0: record the baseline you'll be judged against
Spend the fortnight before go-live collecting these figures. If the receptionist is already live, use the same fortnight from last month or last year and accept the result is rougher.
- Total inbound calls and unanswered calls, from your phone provider's call log or online dashboard. Most business phone systems show missed calls by hour.
- Callback share. Of the numbers that went unanswered, how many called again within 24 hours? Sort the log by number to count them. This is the most important baseline figure, because those callers weren't lost; they were delayed.
- Bookings or orders taken by phone, from your booking system or order book, and their average value.
- Staff time on the phone. A tally sheet by the phone for a week is enough: each call, who took it, rough minutes.
- Your gross margin on a typical phone booking or order. Revenue overstates the gain; margin is what you keep.
Counting the callback share is fiddly the first time, so here is an illustrative extract from a call log sorted by number, as the butcher's shop below would see it:
| Caller (last 4 digits) | Time | Result | Counts as |
|---|---|---|---|
| …4471 | Tue 10:02, then 10:40 | Unanswered, then answered | Callback (rang back within 24 hours) |
| …0918 | Tue 12:15 | Unanswered | Lost (never rang back) |
| …3302 | Wed 17:55, then Fri 09:10 | Unanswered, then answered | Lost for this purpose (rang back after more than 24 hours) |
| …7765 | Thu 11:20, 11:21, 11:23 | Unanswered three times, then answered 11:30 | One callback, not three missed calls |
The last row is the trap. Three missed attempts in a few minutes are one impatient customer, not three lost ones, so count unique numbers, not missed-call rows. Doing that on the baseline fortnight is what turns "110 unanswered" into a figure you can use.
Also note anything unusual about the fortnight: a holiday, a promotion, a staff absence. If your trade is seasonal, write down how the same weeks usually compare with the rest of the year.
The scorecard: eight numbers and where each comes from
| Number | Where to get it | What it tells you |
|---|---|---|
| 1. Total calls | Phone provider log | Whether demand changed, which affects every other figure |
| 2. Answer rate (staff plus AI) | Phone log and AI dashboard | Whether fewer callers now reach nobody |
| 3. Calls the AI handled start to finish | AI dashboard | How much work it's really doing |
| 4. Bookings or orders the AI created | Booking system source tag, or the AI's booking log | The raw output, before attribution |
| 5. Value of those bookings | Booking system | The size of the prize |
| 6. Transfers and callbacks requested | AI dashboard | Calls it passed to people; too many means its instructions are thin |
| 7. Errors | Your own error log | Wrong bookings, wrong answers, missed handovers |
| 8. Staff phone time | Tally sheet, one week a month | The time saving, measured rather than assumed |
Add one warning signal: calls that end within 20 seconds of the AI answering. A rising share means callers are hanging up on it, often because the greeting is too long or it isn't clear how to reach a person.
The greeting is usually the culprit. An illustrative before and after from a shop whose early-hang-up share was running at about 14% of AI-answered calls:
BEFORE (about 18 seconds read aloud)
"Thank you for calling. You're through to our automated assistant,
which can help with placing orders for collection, meat box
subscriptions, opening hours, our Christmas order form and general
questions. Please tell me in a few words how I can help you today."
AFTER (about 6 seconds)
"Hi, you're through to the shop's automated assistant. I can take
your order, or say 'person' and I'll put you through."
Over the next three weeks the early-hang-up share in that example fell to about 6%. The shorter version works because it says what the caller most wants to know, that a person is available, before they have decided to hang up.
Days 1 to 30: shadow the receptionist and fix what breaks
Treat the first month as training, not measurement. Results will be worse than they'll settle at, and that's normal.
- Read every call summary or transcript for the first week, then a daily sample of ten.
- Check every booking or order the AI creates against what the caller asked for, until errors are rare.
- Log each error with the date, what happened, the cause and the fix. Most causes are missing rules or unclear information, which you fix in the AI's instructions.
- Tune the handover rules. Our tutorial on call scripts and escalation rules has templates for when the AI should pass a caller on.
Here is what a first month's error log looks like for the butcher's shop worked through below, with illustrative entries. The pattern to look for is not zero errors but fewer each week and no cause appearing twice:
| Date | What happened | Cause | Fix |
|---|---|---|---|
| Week 1, Tue | "A leg of lamb" ordered with no size; shop sells half and whole | Product list had no size options | Sizes added; AI must ask "half or whole?" |
| Week 1, Fri | Collection booked for Sunday; shop is closed | Collection hours missing from instructions | Hours added, with public-holiday closures |
| Week 2, Wed | Meat-box customer asked to skip a week; AI cancelled the subscription | No rule separating "skip" from "cancel" | Skips allowed; cancellations transfer to staff |
| Week 3, Mon | Asked if sausages are gluten-free; AI said yes from an old description | Recipe changed, description didn't | Description updated; all allergen questions now hand over |
Errors in that month ran 5, 3, 2 and 2 a week, which is why the worked example below uses 2 a week as its steady-state figure.
If errors reach a customer, deal with them as you would a staff mistake; what to do when an AI receptionist gets a booking wrong sets out the recovery steps. And if the setup itself was rushed, the checklist in setting up an AI receptionist without losing callers is worth revisiting now.
Days 31 to 60: count only what the AI caused
Not every booking the AI takes is extra business. Some callers would have reached a person anyway, and some who hit voicemail would have rung back. Counting all AI bookings as gains is the most common way owners overstate the return.
Two rules keep the count honest:
- Count only bookings from calls that would otherwise have been missed. If the AI only receives calls your team didn't answer within a set number of rings, or calls outside opening hours, this is automatic. If the AI answers every call, you can't separate them cleanly; use the baseline missed-call rate as the share that counts.
- Discount by the callback share. If 36% of missed callers used to ring back within a day, only 64% of the AI's bookings from missed calls are true gains.
Keep the error log running and add a monthly tally-sheet week for staff phone time.
Days 61 to 90: the full calculation, worked through
Here's an illustrative butcher's shop that takes phone orders for collection and weekly meat boxes. The AI receptionist gets only calls unanswered after four rings, and all calls outside opening hours.
Baseline fortnight: 420 calls, 110 unanswered, 40 of those rang back within a day (a 36% callback share). Phone orders average $45 at a 35% gross margin. Staff spent about 9 hours a week on the phone.
Steady state, weeks 5 to 12: the AI takes an average of 28 orders a week. Staff phone time falls to about 5 hours a week, because they no longer return voicemails or field repeat calls from customers who couldn't get through. The error log shows about 2 wrong orders a week, each costing roughly $15 to put right. The receptionist costs $150 a month (about $35 a week), and checking summaries takes an hour a week.
| Line | Working | Per week |
|---|---|---|
| Extra orders caused by the AI | 28 × (1 − 0.36) ≈ 18 orders | |
| Margin on those orders | 18 × $45 × 35% | $283 |
| Staff time saved | 4 hours × $20 loaded hourly cost | $80 |
| Cost of errors | 2 × $15 | −$30 |
| Receptionist subscription | $150 a month | −$35 |
| Owner review time | 1 hour × $20 | −$20 |
| Net gain | $278 |
Return on cost = net gain ÷ (subscription + review time) = $278 ÷ $55, roughly five times. Over the eight steady weeks that's about $2,200 of net gain. The first month, with its heavier checking and more errors, should be reported separately, not blended in to flatter or sink the figure.
Dashboard numbers that look good but mislead
Every AI receptionist comes with a dashboard, and most lead with the figures that flatter the product. Read them, but don't put them in your calculation without translating them first.
- "Calls handled." A call the AI answered and then transferred, or one where the caller hung up after ten seconds, often counts as handled. Use calls handled start to finish instead.
- "Hours saved." Usually calculated as calls handled multiplied by an assumed call length. Your tally sheet is the real figure; the dashboard's is an estimate built to look generous.
- "Revenue booked." The value of every booking the AI touched, including ones from callers who'd have reached your team anyway. Apply the attribution rules above before using it.
- "Containment" or "resolution rate." The share of calls that didn't reach a person. High is good only if those callers got what they needed; a caller who gave up also never reached a person.
Translated for the butcher's shop in a typical steady month, the difference is large. Illustrative dashboard figures on the left, what goes into the calculation on the right:
| Dashboard says | After translation |
|---|---|
| 212 calls handled | 133 handled start to finish; 61 were transferred and 18 hung up early |
| 21 hours saved (212 calls × an assumed 6 minutes) | About 17 hours, from the tally sheet (4 hours a week) |
| $5,400 revenue booked (120 orders × $45) | 77 extra orders after the 36% callback discount: $3,465 revenue, about $1,210 margin |
| 71% containment (151 calls never reached a person) | Includes the 18 callers who hung up; the real figure is nearer 63% |
The receptionist still pays well in this example. But a decision made on the left-hand column would assume it earns about four times what it does, and the next decision, such as upgrading to a bigger plan, would be made on the wrong number.
Seasonal businesses need one more adjustment. A butcher's order calls in the run-up to a holiday, or a catering company's in wedding season, can double without the AI doing anything. Compare each week with the same week last year where you can, or at least note the seasonal effect next to your figures so the day-90 decision isn't made on a spike.
How much the answer moves if an assumption is wrong
Before trusting a result, change the shakiest input and see whether the conclusion survives. In the butcher's example, the callback share is the one to test, because it came from only one fortnight.
- If the true callback share were 60% rather than 36%, extra orders fall to about 11 a week, margin to $173, and net gain to about $168 a week. Still clearly positive.
- If 90% of missed callers would have rung back, extra orders fall to about 3 a week and the receptionist loses money on order margin alone; only the staff-time saving keeps it positive, at roughly $42 a week.
If your conclusion flips within a plausible range of an input, you need a longer baseline or a cleaner test, such as switching the AI off on alternate weeks for a month and comparing.
That alternate-week test settles the callback question directly. An illustrative month at the butcher's: with the AI on in weeks 1 and 3, the shop took 131 and 127 phone orders; with it off in weeks 2 and 4, when missed calls went back to voicemail, it took 112 and 115. That is about 15 or 16 extra orders a week with the AI on, a little below the 18 the baseline estimated, which suggests the true callback share is nearer 45% than 36%. Rerun the table at 15 extra orders and the net gain is about $231 a week instead of $278. Still a clear keep, now on a figure measured rather than assumed. Choose four ordinary weeks for this test; one public-holiday week in the "off" pair would make the AI look far better than it is.
A tracking sheet to copy
Week | Total calls | Unanswered (no AI) | AI-handled | AI orders/bookings
| Value of AI bookings | Transfers | Errors | Cost of errors
| Hang-ups <20s | Staff phone hours (tally weeks) | Notes
Baseline (days -14 to 0):
Calls ____ Unanswered ____ Rang back within 24h ____ Callback share ____%
Average booking value $____ Gross margin ____% Staff phone hrs/wk ____
Day 90 summary (steady weeks only):
Extra bookings = AI bookings from missed calls x (1 - callback share)
Margin gained = extra bookings x average value x margin %
+ staff hours saved x loaded hourly cost
- error costs
- subscription and overage
- review hours x hourly cost
= Net gain. Return = net gain / (subscription + review cost)
Day 90: keep, change or cancel
- Keep if the net gain is clearly positive, errors fell over the three months, and hang-ups are stable or falling.
- Change if the gain is positive but errors plateaued or transfers stay high. Rewrite its instructions, narrow the call types it handles, or try a different forwarding rule for another 30 days.
- Cancel if extra bookings don't cover the cost even with generous assumptions, or if errors are reaching customers faster than you can fix the causes. A simpler tool, such as a missed-call text-back, may be the better answer.
Before renewing on an annual plan, compare your real minutes or calls against your tier. Many businesses find they're paying for the wrong allowance once real usage is known; what an AI receptionist costs a small business explains the pricing models. If you're weighing whether a human service would do better for your call mix, see AI receptionist or answering service.
Further reads
- How to Measure Whether Your AI Chatbot Is Actually Working — The same measuring discipline for website chatbots.
- AI Call Summaries: Log Every Phone Enquiry in Your CRM — Log every call automatically so the numbers collect themselves.
- How to Set Up Missed Call Text Back for a Service Business — A cheaper benchmark to compare the receptionist against.
- Best AI Receptionists for Salons, Compared by Price — Receptionist prices compared, if you're still choosing.
- Is Paying for AI Help Worth It? Your Time vs an Expert's Fee — How to put a value on your own time in the sums.
- How to Calculate AI ROI for Your Business (Worked Example) — The AI ROI formula worked through line by line for an illustrative nail salon, with payback period, a pessimistic case and a spreadsheet layout to copy.
- AI Phone Answering for Restaurants: What It Costs and Pays Back — Published prices for restaurant AI phone answering, the costs that never make the pricing page, and a payback calculation built from your own missed calls.
- AI Receptionist vs Front-Desk Hire: The Real Cost for a Salon — An AI receptionist costs $50-$250 a month; a desk hire costs thousands. But the desk does more than phones. Here's how to compare them fairly.
- AI Ordering vs Delivery Apps: What a Takeaway Actually Pays — Delivery app commission against AI phone and web ordering, line by line, with a break-even worksheet a takeaway can fill in with its own numbers.
- How to Cut Salon No-Shows With AI Reminders and Deposits — A deposit rule by booking risk, policy wording, reminder timings and the booking-system settings that cut salon no-shows without upsetting regulars.
- How Dry Cleaners and Laundries Use AI for Orders and Collections — Where AI fits between the counter and collection: phone agents that book pickups, ready texts, an uncollected-garment ladder and calmer claim replies.
- Should a Hair Salon Use an AI Receptionist? — A decision test for salon owners: count your missed calls, check your software, and set the colour and patch-test rules before any AI answers the phone.
- What Is an AI Receptionist and How Does It Handle Bookings? — A plain-English look at what happens inside an AI receptionist during a booking call, and the three very different things vendors mean by 'handles bookings'.
- AI or Human Answering Service: Which Suits a Property Manager? — AI and human answering services priced for a 300-unit portfolio, with emergency routing rules and ten test calls to make before you sign.
- Can an AI Chatbot Book Valuations for Estate Agents Out of Hours? — What an out-of-hours valuation bot needs: live diary access, eight qualifying questions, written diary rules and a list of things it must never say.
- AI Tools and AI Development: The Complete 2026 Guide — the AI hub, including every tutorial in the AI-for-business series.
Sources: Method and worked example are illustrative. Receptionist price used in the example is within the range published on vendor pricing pages checked 27 September 2026.