What an AI Implementation Looks Like in a Small Accounting Firm

Coding Liquids tutorial cover featuring Sagnik Bhattacharya for What an AI Implementation Looks Like in a Small Accounting Firm.
Coding Liquids tutorial cover featuring Sagnik Bhattacharya for What an AI Implementation Looks Like in a Small Accounting Firm.

In a small accounting firm, AI implementation is usually a 10–12 week sequence of small changes rather than one new system: switch on the AI inside your ledger and practice software, give staff one business-grade assistant with a client-data rule, then rebuild records chasing and routine client emails around AI drafts, with a person approving everything that goes out.

What trips most practices up is the order. Buying a tool first and hunting for uses afterwards is the usual way AI stalls by month two. Logging where the hours go, then testing features inside the software the firm already runs, costs very little and shows quickly whether the bigger changes are worth making.

Follow me on Instagram@sagnikteaches

The practice before anything changed

Take a seven-person practice as the illustration: two partners, three accountants, a payroll and bookkeeping clerk and a practice manager, looking after about 220 clients. Most are small companies and sole traders on Xero. The firm runs its workflow in Karbon and its email and files in Microsoft 365 Business Standard. Nobody on the team is technical.

Connect on LinkedInSagnik Bhattacharya

Before choosing anything, the practice manager asked everyone to log one ordinary week against six headings. The team totals looked like this:

Subscribe on YouTube@codingliquids
TaskHours per week (whole team)Who does most of it
Chasing missing records, receipts and signatures9Accountants, practice manager
Reading, sorting and forwarding the shared inbox7Practice manager
Routine client emails (deadlines, reminders, simple queries)6Everyone
Reviewing bank coding and reconciliations14Clerk, accountants
Building year-end query lists4Accountants
Writing up client meeting and call notes3Partners
Total43

Forty-three hours is more than one full-time person. Just as useful was what the log ruled out. The partners had assumed tax-return preparation would be the big win, but it barely registered, because their tax software already did most of the heavy lifting. Without the log they would have spent the first month on the wrong problem.

Weeks 1–2: one decision about tools, one rule about client data

The partners made two decisions and wrote both on a single page.

The first was the assistant. Because the team lives in Outlook and Excel, Microsoft 365 Copilot made more sense than a separate chat app. The practice moved from Business Standard at $14 per user a month to the Business Standard + Copilot bundle at $23.50 per user a month on annual billing, which costs less than adding the $21 Copilot Business licence on top of the existing plan. For seven people, the extra spend came to $66.50 a month. A firm standardised on Google Workspace or ChatGPT Business would make the equivalent choice there; the reasoning is the same.

The second was the client-data rule: client names, figures and documents go only into the approved tools (Copilot signed in with a work account, Karbon and Xero), never into personal apps or free chatbots, on laptops or phones. Whether an accountant can safely use ChatGPT with client data depends on the plan and its settings, and a one-line rule saves a lot of case-by-case judgement calls.

They also added a sentence to their engagement letter template saying the firm uses AI tools within its software to prepare drafts and that a qualified member of staff reviews all work. Check wording like this against your professional body's guidance and your insurer's requirements before you use it. Total time for these two weeks: about ten hours, mostly partner time.

The finished page was short enough to pin above every desk. Filled in, it read roughly like this (illustrative):

  • Our assistant: Copilot, signed in with your work account. Nothing else for client work.
  • Client data goes only into: Copilot, Karbon, Xero. Not personal apps, not free chatbots, not on your own phone unless Copilot is installed and signed in.
  • Everything AI drafts is read before it leaves the firm. No exceptions for "just a reminder".
  • Never ask AI for a figure you haven't checked. Tax amounts, deadlines and balances come from the ledger or the tax software.
  • Clients who object: flag them in Karbon; their emails are written by hand.
  • Questions or a slip-up: tell the practice manager the same day. Nobody gets into trouble for reporting.

The fifth line mattered sooner than expected. In week four, a long-standing client replied to the updated engagement letter saying they didn't want AI used on their affairs. Because the rule already existed, the answer was simple: a "no AI drafting" tag on the client in Karbon, visible on every work item, and chasers for that client written the old way. It cost about ten minutes a month. Without a written opt-out route, the likely outcome is that one person remembers and the others don't.

Weeks 3–4: switching on what the software already had

Before buying anything else, the practice manager went through the settings of the three systems the firm already used. It took two afternoons and produced most of the early gains.

  • Karbon. Karbon made its AI features available to all customers in March 2026: email and thread summaries, client summaries that pull together emails, notes and work items, editable quick replies, and emails drafted from tasks. Each person tried summaries on their five longest client threads.
  • Xero. Xero's JAX assistant reads bills and receipts through Smart Document Capture. Xero has announced automatic chasing of missing receipts for October 2026, and its automatic bank reconciliation is still rolling out. The clerk turned on document capture for the twenty clients with the most paper.
  • Copilot in Outlook. Everyone tried thread summaries and draft replies on the shared inbox for a week, under one instruction: no draft goes out unedited.

Nobody changed how work was done yet. These two weeks were for finding out which features were accurate enough to build on. Karbon's summaries passed easily. Copilot's drafts were usable roughly two times in three and needed the firm's own phrasing the rest of the time. Document capture read typed invoices well and handwritten receipts badly, which told the clerk which clients to move first.

The routine-email test was the most instructive. A typical before and after, from a client asking about payroll (illustrative):

  • The client's email: "Hi, we've got a new starter on the 3rd, what do you need from me and when does it need to be in for this month's run?"
  • Copilot's first draft: a friendly reply asking for "their details and any relevant documents by the usual cut-off", ending "let me know if you have any other questions".
  • What the clerk sent: a list of the four items the firm actually needs for a new starter (full name, start date, pay rate and hours, and the starter form from the payroll software), the firm's real cut-off date for that month, and a link to the secure upload page.

The draft was polite and useless, because "the usual cut-off" and "relevant documents" lived in the clerk's head, not in the email thread. That one example changed how the team used drafts: for anything with firm-specific facts, they pasted a short standard list into the prompt, and saved the most-used lists as snippets. The tone came from Copilot and the facts came from the firm.

Weeks 5–8: rebuilding records chasing around AI drafts

Chasing was the biggest block of time that didn't need professional judgement, so it became the first real workflow. The old process relied on an accountant noticing a gap, digging out the last email, writing a chaser from scratch and remembering to follow up. The new one runs from Karbon work items:

  1. Every year-end or quarterly job carries a checklist of required records in Karbon.
  2. If an item is still missing seven days before the internal deadline, the accountant opens the work item and asks for a chaser drafted from the task, not from the email thread.
  3. The draft lists exactly what's missing, the period it covers and the date the firm needs it, in the firm's tone.
  4. The accountant reads it, corrects it and sends it. A second chaser follows three days later, then a phone call at the deadline.

For chasers written in Copilot rather than Karbon, the team shared one prompt so the emails read the same whoever wrote them:

Draft a short, friendly email to a client asking for missing records.
Client type: [sole trader / limited company]
Missing items: [list each item and the period it covers]
We need them by: [date], because: [e.g. to finish the accounts in time to file]
How to send them: [upload link / reply with attachments]
Tone: plain English, no jargon, no blame. Under 120 words.
End with one sentence offering a quick call if anything is unclear.
Do not add any figures, dates or items that are not listed above.

Here's the kind of draft that prompt produced for a landscaping client (illustrative), with the prompt's gaps filled in:

Subject: A few records we still need for your accounts

Hi [first name],

We're nearly ready to finish your accounts for the year to 31 August,
but we're still missing three things:

- Statements for the business savings account, June to August
- The invoice for the ride-on mower bought in July
- Your signed approval of the draft figures we sent on 2 October

Could you send these by Friday 14 November, so we have time to finish
before the filing date? You can upload them through your client portal
or reply to this email with attachments.

If anything's unclear, I'm happy to have a quick call.

Two fixes before it went out. The model paired the date with the wrong weekday: 14 November 2026 is a Saturday, and weekday slips like this are one of the commonest errors in AI-drafted chasers. It also said "client portal" when the firm calls it the "secure upload page", which confuses clients who have used it before. Both took ten seconds to correct, and both would have caused a reply asking what was meant.

Chasing time fell steadily over the four weeks, mostly because nobody had to reconstruct what was outstanding. The practice manager also tracked one outcome number alongside the hours: of the jobs that reached the seven-day chaser, how many had every record in by the internal deadline. In an illustration like this, that might move from 26 of 40 jobs in the month before the change to 33 of 40 in week eight. The hours tell you the work got cheaper; the completion count tells you the chasers still work, which is the thing clients and deadlines care about. There's more on sequencing and escalation in chasing missing client records with AI before deadlines.

Weeks 9–12: query lists and meeting notes

With chasing settled, the accountants tried Copilot in Excel on year-end work. They exported each trial balance with last year's comparatives and asked for a first-draft query list:

This sheet shows a client's trial balance for this year and last year.
List every account where the change is more than 20% AND more than 1,000.
For each one, write a plain-English question we could ask the client
about why it changed. Do not guess the reason.
Group the questions under income, costs, assets and liabilities.
If you can't interpret an account, flag it instead of assuming.

For the same landscaping client, part of the reply looked like this (illustrative):

COSTS
- Equipment hire: down 48% (12,400 to 6,450).
  Question: Did the business hire less equipment this year? If so, why?
- Fuel: up 31% (5,200 to 6,810).
  Question: Did the number of vehicles or jobs increase this year?
- Bank charges: up 25% (400 to 500).
  Question: Have your bank's fees or payment methods changed?
ASSETS
- Plant and machinery: up 9,800.
  Question: Please send invoices for any equipment bought this year.
FLAGGED
- Account 4905 "Sundry": can't interpret without more detail.

The accountant made three changes. Bank charges moved by only 100, so they failed the "more than 1,000" test and shouldn't have been listed; models often apply one half of an AND condition, so check thresholds against the figures. The fall in equipment hire and the rise in plant and machinery were one story, the purchase of the mower, so she merged them into a single question. And the sundry account became a request for the ledger detail rather than a question the client couldn't answer. The output was always a starting point: the accountant still added the questions only someone who knew the client would think to ask. Even so, building a query list dropped from about an hour to about 35 minutes per client.

For online client meetings, the partners began using Teams meeting recaps, telling clients at the start and offering to switch recording off. For face-to-face meetings they kept handwritten notes and dictated a summary into Word afterwards, which Copilot tidied into the firm's file-note format.

The tidying needed one habit to be safe. A partner's dictation after a meeting with a café owner included "said we'd look at whether moving to a limited company makes sense, no promises". Copilot's file note turned that into an action: "The firm will advise the client on incorporating." That's a commitment the partner hadn't made, recorded in the firm's own file. The fix was a line added to the tidy-up prompt, "keep any hedging exactly as dictated; don't turn possibilities into actions", and a rule that the partner reads the note against the dictation before it's saved. Five minutes of review protects the record that matters most if a client later disputes what was said.

What the first quarter cost

ItemOne-offMonthlyNotes
Copilot bundle upgrade, 7 users–$66.50$23.50 bundle vs $14 before, annual commitment
Karbon and Xero AI features–Nothing extra in this caseDepends on your Karbon plan and each client's Xero plan; check before assuming
Internal time: time log, policy, settingsAbout 30 hours–Partners and practice manager
Internal time: staff trying features, learning promptsAbout 21 hours–Three hours each
Internal time: week-12 measurementAbout 8 hours–Repeating the log and comparing
Outside helpOptional–Not used in this illustration

At an internal cost of, say, $60 an hour, the 59 hours of staff time is worth about $3,540, far more than a quarter's software. That's normal. Most of what an implementation costs is the team's own time, which is exactly why the baseline log matters: it tells you whether you'll earn that time back. A fuller breakdown by firm size is in how much AI costs a small accounting firm.

Three problems the review step caught

A chaser for the wrong year-end. Two clients with near-identical company names had been discussed in the same email thread. A draft built from that thread asked the wrong client for the wrong period. The accountant spotted it only because of the rule that nothing is sent unread. The lasting fix was the one already described: draft from the Karbon work item, which belongs to exactly one client and one job.

A coding error repeated with confidence. The clerk's first monthly spot-check found that suggested bank matches for a removals firm had coded its van lease payments as fuel, every month, because an early wrong match had been accepted and then learnt. Automation that learns from past decisions learns past mistakes too. The practice now samples 20 automatically matched transactions per client each month and corrects the underlying rule, not just the transaction.

Personal apps creeping back. In week six, a staff member mentioned using a free chatbot on their phone to rephrase a client email that included the client's name and a tax figure. Nobody was reprimanded. Instead the partners made sure Copilot was installed and signed in on every phone, because people reach for whatever is quickest. The written policy followed; an AI acceptable use policy for a small professional firm covers what it should say.

What the partners measured at week 12

They repeated the same one-week log against the same six headings. Week 12 fell in a quieter month than the baseline, so they compared hours per job rather than raw weekly totals, then converted back to a typical week.

TaskBaseline (hours/week)Week 12 (hours/week)
Chasing records95.5
Shared inbox75
Routine client emails63.5
Bank coding review1412.5
Year-end query lists42.5
Meeting and call notes31.5
Total4330.5

Twelve and a half hours a week, spread across seven people, is real but easy to lose: it vanishes into the day unless someone decides what it's for. The partners chose two uses: finishing year-end work earlier in the season, and a short monthly management-accounts call for their fifteen largest clients, priced as a new service.

Bank coding barely moved. Automatic matching saved time on clean files and added review time on messy ones, so the net gain was small. Expect that pattern: the review burden grows as automation takes on more, and it has to be counted. The logging method they used is set out in measuring time saved after rolling out AI in a small firm.

What they would do in a different order

  • Write the client-data rule before anyone opens a tool. The phone incident happened because the rule was a conversation, not a document.
  • Spot-check automation from the first month. The van-lease error ran for three months. A 20-transaction sample in month one would have caught it in the first batch.
  • Give one person an hour a week to own it. The practice manager became the person colleagues asked, and progress slowed noticeably whenever she was on leave.
  • Park the ambitious ideas until the basics work. One partner wanted a website chatbot answering tax questions. Leaving it for later was right: one confident wrong answer in public would cost more trust than all the hours saved.

Questions practices ask before starting

How long before AI saves a small accounting firm any time?

In the illustration, the first measurable savings came in weeks three and four, from features already inside the practice and ledger software, and the main records-chasing workflow paid back its setup time within the quarter. Firms that skip the baseline time log often can't tell whether anything changed, which is one of the main reasons AI use fades after the first month.

Should we start an AI implementation during busy season?

It's better not to. Busy season distorts the baseline, leaves no time for careful review, and staff under pressure skip checks. Start the time log and the client-data rule straight after a deadline, when the pain points are fresh but the pressure is off, and pilot the new workflows on routine jobs before the next peak arrives.

Do clients need to be told the firm uses AI?

Telling them is sensible and is usually well received when it's framed around review: the firm uses AI tools inside its software to prepare drafts, and a qualified person checks all work. Add a line to engagement letters, check your professional body's guidance and your insurer's terms, and always ask before recording or transcribing a client meeting.

Further reads

Sources: Microsoft 365 business plan and Microsoft 365 Copilot Business pricing pages; Karbon release notes (31 March 2026); Xero announcements from Xerocon 2026 on JAX features and availability.

Want your practice's first 12 weeks mapped out?

On a 1:1 call we'll look at where your team's hours actually go, check which AI features your ledger and practice software already include, and pick the first workflow worth rebuilding.

Book a 1:1 call with me