Rolling Out AI in a Bookkeeping Practice Without Losing Control

Coding Liquids tutorial cover featuring Sagnik Bhattacharya for Rolling Out AI in a Bookkeeping Practice Without Losing Control.
Coding Liquids tutorial cover featuring Sagnik Bhattacharya for Rolling Out AI in a Bookkeeping Practice Without Losing Control.

Roll it out in stages with the controls built in first. Decide which outputs always need a person's sign-off, pilot one workflow on five to eight clients, set user permissions and review queues in your ledger software, check a fixed sample of AI-coded work every month, and expand only when error rates stay low for two month-ends running.

The complication is that AI is arriving in your software whether or not you plan a rollout. Intuit's help pages for QuickBooks Online say there currently isn't a way to turn off its AI features individually. Xero currently charges nothing extra for JAX's chat, and what JAX can do is set by the pricing plan and each user's role. So "without losing control" is less about deciding whether to switch AI on and more about deciding who approves what, and keeping evidence that they did.

Follow me on Instagram@sagnikteaches

Stage 1: Write down your control points (half a day)

Before touching any setting, list the practice's recurring work and mark where AI may prepare and where a person must approve. This is the document everything else hangs on. A filled-in example for a practice doing monthly bookkeeping:

Connect on LinkedInSagnik Bhattacharya
ActivityAI may preparePerson must approveEvidence kept
Bank transaction codingSuggested account and tax codeAll items over $500, all new payees, all during the pilotReviewer's initials in the review log; software audit history
Bank reconciliationSuggested matchesEvery match before reconciling; month-end reconciliation sign-offSigned month-end checklist
Receipt and bill captureExtracted supplier, date, amount, line itemsAnything flagged as low confidence; totals that differ from the bank lineCapture tool's review queue history
Client chasers for missing documentsDraft email listing missing itemsFirst send to each client; any wording changeSent items
Client queries and explanationsDraft reply in plain EnglishEvery reply, alwaysSent items
Management reports and commentaryFirst draft of commentaryEvery figure and sentenceReviewed version saved with the report
Filings and returnsWorking papers, reconciliationsAlways, by the responsible personYour normal filing sign-off

Notice what the table does not say: it never says "AI does this on its own". Some items may earn that later, for specific clients, on evidence. None starts there.

Subscribe on YouTube@codingliquids

Stage 2: Take stock of the AI already in your stack (two hours)

Most practices are surprised by this list. Ask each team member what they use, and check each product's current AI features. An illustrative inventory for a practice with clients split between Xero and QuickBooks Online:

AI INVENTORY - completed 2 October
Tool                    AI feature in use          Who    Status
Xero                    JAX suggestions, bank rec  All    In use, unreviewed
QuickBooks Online       Accounting AI suggestions  2      Can't switch off per
                        (categorisation)                  feature - needs review
Dext                    Extraction, supplier rules All    In use; rules last
                                                          reviewed 2024
Hubdoc (Xero clients)   Document extraction        1      In use
Microsoft 365           Copilot Chat (web)         All    Personal use, no rules
ChatGPT (free)          Drafting client emails     1      Pasting client names -
                                                          STOP, move to business plan
Meeting note-taker      Auto-joins Zoom calls      1      Records clients without
                                                          asking - switch off auto-join

The last three lines are the ones that cost practices control, because they involve client data leaving the practice's systems without anyone having decided it should. Whether it's safe to use ChatGPT with client data covers the plan types and settings; the short version is that business plans (ChatGPT Business, Claude Team, Microsoft 365 Copilot) don't train on business content by default, while consumer plans need the model-training switch turned off in privacy settings, and client identifiers are best removed either way.

Check the ledger products carefully too. In its recent Xerocon announcements, Xero has described JAX abilities such as emailing clients for missing documents, sending reminders and matching documents to transactions for review and approval. Features like that change who is talking to your clients, so decide whether they are on, for which clients, and who approves the wording.

Stage 3: Pick one workflow and a handful of clients (one hour)

Choose a pilot that is frequent, easy to check and low in consequence if a mistake slips through for a day. Bank coding for clients with clean, repetitive transactions is the usual first choice; missing-document chasers are a good second. Avoid starting with anything that produces a filing.

Pick five to eight clients who between them represent your book: a couple of simple sole traders, a retailer with card takings, a client with lots of small suppliers. Leave out your messiest client and anyone mid-dispute. Tell the pilot clients if the change will be visible to them.

The first six rows of an illustrative shortlist, drawn up in ten minutes from the client list:

ClientTransactions a monthIn or outReason
Sole-trader decorator60InFew suppliers, same pattern every month
Sole-trader copywriter35InSimple; checks how AI handles low volume
Gift shop with card takings240InTests card payouts and fees
Café with many small suppliers310InTests new-payee handling
Letting agent holding client money400OutClient-money accounts need a person on every item
Contractor mid-dispute with a customer90OutFigures likely to be scrutinised line by line

The two "out" rows are as useful as the four "in" rows. Writing down why a client is excluded stops someone adding them to the pilot in month two because the first month went well.

Stage 4: Set permissions and review settings (half a day)

  • User roles. In each ledger product, check who can approve suggestions, reconcile and publish. Junior staff can prepare; a named senior approves. In Xero, JAX's abilities follow the user's role and permissions, so the role settings are where control actually lives.
  • Bank rules versus AI suggestions. A bank rule you wrote is explicit and auditable. An AI suggestion is a prediction. Where a pattern is stable, turn the reviewed suggestion into a rule, so you know exactly why it codes that way.
  • Capture tool rules. Review supplier rules in Dext or Hubdoc before the pilot. Old rules quietly apply the wrong account for years.
  • Anything you can't switch off. Where a feature can't be disabled, as with QuickBooks' AI features, compensate with review: make "accepted AI suggestions" part of the monthly sample.

For the setup of capture and bank feeds in detail, see AI receipt capture and bank categorisation: a bookkeeper's setup.

Stage 5: Review and sample, with fixed rates

This stage is where control is kept or lost. Fix the sampling rates in advance so nobody relaxes them because the first month looked fine.

  1. Weeks 1 to 4 (first month-end): review 100% of AI-suggested codings and matches for pilot clients.
  2. Month 2: review every item over your threshold (say $500), every new payee, every item the software marks as low confidence, plus a random 20% of the rest.
  3. Month 3 onwards, if the error rate stays under your target: the same mandatory items plus a random 10%.
  4. Any month the error rate jumps: go back to 100% for that client until you know why.

Put numbers on those rates before month 2 starts, so the reviewer knows the size of the job. For the pilot clients described in Stage 8 below, about 1,020 of the month's 1,200 transactions arrive with an AI suggestion. Suppose 140 of those are mandatory checks (over $500, new payees, low confidence). Month 2 is then 140 plus 20% of the other 880, which is 316 items; month 3 is 140 plus 10%, or 228. At roughly half a minute per item, that's about 2.6 hours and 1.9 hours of review. If the reviewer reports 40 minutes, the sample isn't being done.

"Random" needs a method too, because reviewers under pressure pick the easy-looking lines. Export the month's accepted suggestions to a spreadsheet, add a column with =RAND(), sort by it, and review the top 20% or 10%. It takes two minutes and removes the choice from the reviewer.

Finally, decide what counts as an error and when a single error overrides the rate. Count any item that needed changing, whether the account, the tax code or a missing split. Then add an override: any single error above an amount you set, say $5,000, sends that client back to 100% review whatever the monthly rate says. A 1.2% error rate looks healthy until you find that one of the errors was a $12,000 loan receipt coded as sales, which would have overstated the client's profit by exactly that amount.

Log every error, not just the count. An excerpt from an illustrative review log:

Client  Date   Item                      AI suggested        Correct             Cause
C-14    03/10  Transfer 2,000 to savings  Sales income        Transfer (balance)  Own-account transfer
                                                                                 read as receipt
C-09    05/10  Fuel card statement        Motor expenses      Split: motor + staff Multi-line statement
                                                              welfare (snacks)
C-22    07/10  Refund from supplier       Sales               Purchases (credit)  Credit mistaken for sale
C-14    12/10  Card takings payout        Sales, net of fees  Sales gross + fees  Fees netted off

The cause column is what makes the log useful. Three of those four point to a fix that stops them recurring: a bank rule for the client's own-account transfers, a rule that supplier credits go to purchases, and a split for card payouts so fees are recorded separately. The fuel-card one simply needs a person each month.

A realistic slip to watch for: a capture tool reads a restaurant receipt's total including a handwritten tip, while the bank shows a different amount; the AI "helpfully" matches it anyway because date and supplier line up. Make "amount differs from bank line" an automatic review item.

Stage 6: Tell clients, and update the paperwork

Clients should hear about AI from you, in plain words, before they notice it. Add a short clause to engagement letters and a paragraph to your onboarding pack. An illustrative clause to adapt with your own adviser's input:

Use of automation and AI tools
We use automation and AI features within our accounting software
and approved business tools to help prepare your records, such
as suggesting transaction categories and extracting details from
receipts. These tools prepare work; a member of our team reviews
it before it is finalised. We use only tools whose terms protect
the confidentiality of your data, and we do not use your data to
train public AI models. Ask us if you would like more detail.

Some clients will ask directly, often after reading about AI in the news. An illustrative reply to "Is a robot doing my books now?": "Not on its own. Our software suggests how each transaction should be categorised, which saves time on the routine ones, and one of our team checks the suggestions before anything is finalised. Anything unusual, anything large and anything new is always looked at by a person. Your figures are signed off the same way they always have been." It's specific about what the software does and who is accountable, which is what the question is really asking.

Check that clause against what you actually do; a promise you don't keep is worse than none. For the wider question of which client data may go into which tools, an AI acceptable use policy for a small professional firm gives a policy you can adapt, and your professional body may publish guidance on AI and confidentiality worth reading alongside it.

Stage 7: Give the team rules and prompts they can use

A one-page policy covers most of it: which tools are approved, what data may go into each, what always needs review, and who to ask. Pair it with a small shared set of prompts, so people aren't improvising with client data. An illustrative prompt for a missing-documents chaser, run on a business plan with identifiers removed:

Draft a short, friendly email to a bookkeeping client asking for
missing items before month-end. Plain English, no jargon, no
blame. List the items as bullet points. Ask for them by [date].
Missing: 3 supplier invoices (dates 4, 11 and 19 Sept),
bank statement for the savings account, receipt for a card
payment of 312.40 on 22 Sept.

Illustrative output: a clear four-bullet email with a polite deadline, and one line added unprompted: "If these aren't received, we may be unable to complete your return on time." That sentence is the part to fix. It raises a filing deadline the prompt never mentioned and sounds like a threat to a client who is simply behind on receipts. Delete it, and add "Don't mention deadlines, penalties or returns unless I include them" to the shared prompt. ChatGPT prompts for bookkeepers has more to build the shared set from.

Warning signs that control is slipping

Control rarely fails in one dramatic moment. It erodes as people get busy and the suggestions look right most of the time. Watch for these, and treat any one of them as a reason to go back to a full review for the client concerned:

  • Review time falls faster than the sample rate. If the reviewer's time drops to almost nothing while the sample is meant to be 20%, the sample isn't being done properly. Check the log for initials and dates.
  • Suggestions get accepted in bulk. Accepting a page of suggestions in one go is sometimes fine for a reviewed, stable client. For a pilot client, it means nobody looked. Ask reviewers how they work through the queue, and spot-check a few accepted items yourself.
  • The same error appears twice. The log should turn every recurring error into a rule or a checklist item. A repeat means the loop from log to fix is broken.
  • Clients start asking "why is this in that category?" Client queries about coding are the most honest error measure you have, because clients notice what samples miss.
  • Nobody can say which features are on. If a team member can't answer "is JAX emailing this client directly?" the inventory from stage 2 is out of date. Refresh it every quarter, because vendors add AI features with every release.

None of these means the rollout has failed. They mean the controls need attention, which is the point of having them.

Stage 8: Measure, then expand

Here is how the numbers might look for an illustrative four-person practice with 70 clients, piloting AI-assisted bank coding on 8 of them (about 1,200 transactions a month between them):

  • Before: about 9 hours a month to code and reconcile those 8 clients.
  • Month 1 (100% review): 7.5 hours. The AI suggested codes for 85% of items; 4% of suggestions were wrong, mostly transfers and card payouts.
  • Month 2 (20% sample, new bank rules in place): 5.5 hours, error rate in the sample 1.5%.
  • Month 3 (10% sample): 5 hours, error rate 1.2%. Two consecutive months under a 2% target: ready to expand.

Expand in groups of ten to fifteen clients, each starting again at 100% review for its first month-end, because every client's transactions are different. Track four numbers per client each month: time spent, error rate in the sample, the number of client queries about their figures, and corrections made after month-end. That last one is the true test: if corrections after sign-off rise, control has slipped, whatever the time savings say.

After bank coding, the natural next pilots are document chasers, receipt capture rules and first-draft management commentary, each run through the same stages. If you are choosing between ledger platforms partly on their AI, Xero vs QuickBooks AI compares what each saves in practice, and using Xero's JAX for invoices and cash-flow questions goes deeper on that assistant. If you'd rather map the rollout with someone who has done it before, that is exactly what my AI implementation consultation covers.

Bookkeeping AI rollout: questions practices ask

Do I need client permission to use AI on their books?

Using AI features built into the ledger software the client already uses is usually covered by your normal terms, but tell clients how you use AI and check your engagement letter covers it. Uploading client data to a separate AI tool is different: check that tool's data terms, and if in doubt ask your data-protection adviser or professional body before doing it.

How long should the pilot run before expanding?

At least two full month-end cycles for the pilot clients. One month shows whether the workflow runs; the second shows whether error rates hold once the novelty wears off and the team starts trusting the suggestions. Expand only when two consecutive samples come in under your error threshold.

What if a team member is already using ChatGPT on client work?

Don't start with a ban. Find out what they use it for, because it often points at your best pilot. Then move that use onto a business plan that doesn't train on your data, strip or anonymise client identifiers, and bring the task into your written policy with a review step.

Should AI-suggested transactions ever be approved without review?

Only for narrow, repeating items you have already reviewed many times, such as the same supplier, same amount range and same account every month, and only after your sampling shows a very low error rate for that client. Even then, keep sampling a share of them, because suppliers and client habits change.

Further reads

Sources: Intuit QuickBooks help (overview of AI agents in QuickBooks Online), Xero Central (About JAX), Xero Xerocon 2026 announcements as reported by Accounting Today, Dext Help Centre (plans for accountants and bookkeepers). Practice sizes, volumes and error rates in the examples are illustrative.

Want an AI rollout plan your practice can trust?

On a 1:1 call we'll list the AI already running in your software, choose the pilot workflow and clients, and set the sign-off points and sampling routine your team will follow.

Book a 1:1 call with me