Payroll bureaus get the most from AI at the two ends of the pay cycle, where errors start and queries arrive. Use it to read messy client input into your import template, compare each run with the last and explain unusual changes, and draft plain-English answers to payslip questions for staff to check. Gross-to-net calculations stay in your payroll software.
That boundary matters. Payroll software is built and updated to apply statutory rules; a general AI model is not, and it will happily produce a confident net pay figure that's wrong. Its value lies before the calculation (turning a photographed timesheet into clean rows) and after it (explaining why someone's pay moved). Most employee queries come from changes nobody explained, so a short "what changed" note after each run often cuts more queries than any clever reply.
Where bureau errors and queries start
| Source | Typical example | AI's role | Stays manual |
|---|---|---|---|
| Late or incomplete input | Overtime arrives after cut-off | Draft the chaser and a missing-items list | Deciding whether to run without it |
| Input in odd formats | Email body, photo of a paper timesheet, a manager's own spreadsheet | Extract into your import template | Checking flagged rows |
| Starters and leavers | Missing bank or tax details; leaver still on the run | Flag incomplete records | Collecting the details |
| One-off payments | Bonus typed as hours | Flag values outside the normal range | Confirming with the client |
| Deduction changes | Pension rate change, new court-ordered deduction | Explain the change in plain words | Setting it up |
| Unexplained changes | "Why is my pay lower?" | Draft the answer from two payslips | Checking and sending |
Stage one: messy client input into your import template
Every bureau has clients who send hours in the body of an email at 4pm on cut-off day. AI extraction, done carefully, turns that into rows your software can import. The rules that make it safe: extract into a fixed template with employee IDs, never let the model fill a gap, and flag anything uncertain.
Here is what arrives from an illustrative toy shop client:
Hi, hours for this week. [Employee A] did her usual plus Sat 7:30
(covering). [Employee B] 32.5. New starter [Employee C] started Weds,
3 days x 6hrs. [Employee D] off sick Mon-Tue, rest normal. Can you add the $150 Christmas bonus for
everyone who's been here since Sept? Thanks!
The prompt asks for rows in the bureau's template, using the client's employee list (IDs and contracted hours) pasted in, and a separate "needs checking" list. What comes back (illustrative):
employee_id, basic_hours, overtime_hours, sick_days, one_off_payment
TS-004 (A), 37.5, 7.30, 0, 150.00
TS-007 (B), 32.5, 0, 0, 150.00
NEW (C), 18.0, 0, 0, 0
TS-002 (D), 22.5, 0, 2, 150.00
Needs checking:
- Employee C is not on the employee list: starter details needed.
- "Here since Sept": confirm which employees qualify for the bonus.
- Employee D's sick days: confirm whether sick pay applies.
Two fixes before import. Employee A's Saturday "7:30" became 7.30 decimal hours; the client meant seven and a half hours, so it should be 7.5. That hh:mm-versus-decimal slip is the most common extraction error in payroll and worth a dedicated rule in the prompt ("convert all times to decimal hours; 7:30 means 7.5"). And the model awarded the bonus to three people without knowing their start dates; the "needs checking" list rightly raises it, but the $150 figures shouldn't be in the rows until the client confirms. For handwritten sheets the extraction step is covered in turning handwritten forms into spreadsheet data.
Managers' own spreadsheets bring a different trap. An illustrative café client sends a rota with shifts written as start and finish times, "08:00-16:30", and a line at the top saying shifts over six hours include a 30-minute unpaid break. Asked for totals, the model returns 8.5 hours for that shift; the payable figure is 8. Two changes stop it. Put each client's break rule into the prompt as its own line ("deduct 30 minutes from any shift longer than 6 hours"), and ask for the start, finish, deduction and total in separate columns so the checker sees the working rather than a single number. A hidden error becomes one you can spot in seconds.
Stage two: a variance check before every run
Most payroll software offers a variance or comparison report. If yours does, use it and let AI write the explanations; if it doesn't, a spreadsheet comparing this period with the last one does the same job. Set thresholds per client, then flag:
| Check | Illustrative threshold | Why |
|---|---|---|
| Net pay change | More than 15% vs last period, same frequency | Catches misplaced decimals and missing input |
| Gross pay spike | More than twice the three-period average | Bonus typed as hours, duplicated lines |
| Hours | Over 60 in a week | Keying errors, possible working-time issues |
| Pay rate | Any change without a pay-rise note | Wrong rate applied or rate not updated |
| Bank details | Any change since last run | Payroll diversion fraud |
| Starters and leavers | Missing details; paid after leaving date | Overpayments that are hard to recover |
| Net pay | Zero or negative | Deductions exceeding pay |
The bank details row deserves its own rule. Fraudsters email bureaus and employers pretending to be an employee and asking for pay to go to a new account. Any change requested by email gets confirmed by phone on a number you already hold, never one given in the email. No AI step should be able to update bank details.
Thresholds need setting per client, and seasonal clients show why. Take an illustrative garden centre with 30 staff whose hours roughly double between March and May. With a flat 15% net-pay threshold, the first April run flagged 19 of the 30. The checker, facing a long list of obviously seasonal increases, waved them all through, and the one real error went with them: a weekly-paid employee keyed on a monthly salary figure. The fix was to compare that client with the same period last year as well as the last period, raise its net-pay threshold to 40% for spring, and keep the pay-rate and bank-detail checks at zero tolerance all year. A useful rule of thumb: if a client's list regularly flags more than a fifth of its staff, the thresholds are wrong for that client, because nobody reads a flag list that long properly.
AI's job at this stage is the words: turning the flag list into a short query email to the client ("Three things to confirm before we run Friday's payroll...") and a note on your file of what was checked. For the timesheet side of this check, AI payroll checks for timesheet errors goes deeper.
For a charity shop client with 14 staff, three flags became this illustrative email, drafted in under a minute and checked in two:
Subject: 3 quick checks before Friday's payroll
Hi [contact name],
Before we run this week's payroll, could you confirm:
1. Employee CS-09: 64 hours this week (usually 20). Is that right,
or should it be 6.4 or 24?
2. Employee CS-03: hourly rate is now $14.20 (was $13.40). Has a pay
rise been agreed, and from what date?
3. CS-11 left on 12 September but is on this week's input. Should
this be final pay only?
If we don't hear by 2pm Thursday, we'll hold these three and pay
everyone else as normal.
The last line matters as much as the questions: it tells the client exactly what happens if they don't answer, so nobody is paid on a guess.
Stage three: answering "why is my pay different?"
Sort incoming queries first: net pay lower than expected, missing overtime or bonus, holiday or sick pay, pension changes, access to payslips, year-end tax forms. The shared-inbox sorting itself is covered in AI email triage for professional firms. For "why has my pay changed?", which is usually the biggest category, give the AI both payslips (with the name replaced by an ID) and ask for an explanation line by line.
Count before you write any templates. Say the bureau logs 160 queries in a month: 58 "why has my pay changed", 34 missing overtime or bonus, 27 payslip portal access, 22 holiday or sick pay, and 19 year-end forms and everything else. The overtime category looked like a reply-writing problem until someone sorted it by client. Twenty-six of the 34 came from two clients who sent overtime after cut-off, so it was paid a month late and every affected employee asked where it was. No reply template fixes that; an earlier chaser to those two clients does, and it removes the queries rather than answering them faster.
The query: "My pay this month is about $180 less than last month and I did the same hours. Can someone explain?" An illustrative draft reply:
Thanks for getting in touch. Comparing your June and July payslips,
your gross pay was the same, $2,640.00. Two deductions changed:
- Pension: your contribution rose from 3% to 5%, as required by law,
an increase of $52.80.
- Tax: $126.40 more, because your tax code changed on 1 July.
Together these account for $179.20 of the difference.
The arithmetic checks out, and the structure is exactly right. Two phrases need fixing. "As required by law" is invented: this employee chose to raise their pension contribution, which the pension record shows, so the reply should say "following your request in June". And "because your tax code changed" is fine, but the bureau can't say why it changed, since the tax authority sets codes; the reply should say so and tell the employee who to contact if they think the code is wrong. A reviewer who reads with the payslips open catches both in a minute.
Both errors came from the model filling a gap with a plausible reason, so the prompt now forbids exactly that:
Draft a reply to an employee's payslip query for a payroll bureau.
Attached: payslip A (last period), payslip B (this period), and any
notes from the payroll record.
1. List every line that differs between A and B, with both figures.
2. Explain each difference ONLY from the notes provided. If no note
explains it, write [reason to confirm] instead of a reason.
3. Never say a change was required by law unless the notes say so.
4. For a tax code change, say the tax authority sets codes and the
employee should contact it if they think theirs is wrong.
5. Add up the differences and check they equal the change in net
pay. If they don't, say so at the very top of the draft.
Plain words, under 150 words.
Rule 5 earns its place. When the differences don't add up to the change in net pay, there is usually a third line nobody noticed, such as a one-off deduction or an overpayment being recovered, and that is precisely the case where the employee's query is justified and needs a person rather than a template.
Stage four: the "what changed" note that heads off queries
Many queries are really "nobody told me". A short summary sent to the client contact after each run, which they can pass on to staff, answers them in advance. AI drafts it from the variance report. A filled-in example for a small charity client:
July payroll: 23 employees paid on 28 July. Total net pay $41,380.
What changed this month:
- 2 starters (Finance Assistant, Shop Supervisor), 1 leaver (final
pay includes 4 days' unused holiday).
- 3 employees have new tax codes issued by the tax authority; if
anyone thinks theirs is wrong, they should contact the tax
authority directly.
- 1 employee's pension contribution increased at their request.
- Summer events overtime paid for 6 employees at time and a half.
Nothing else changed. Payslips are available in the portal now.
Check it against the variance report before sending, particularly the counts, which is where a model most often slips.
The counts slip in a predictable way. In one illustrative run the note said "2 starters" when the variance report listed three: a returning seasonal worker had been re-entered under their old employee ID, so the model treated them as an existing employee whose pay had risen from zero. Nobody was paid wrongly, but the client's manager noticed the missing starter and emailed to ask whether the new person's details had arrived, which is the very doubt the note exists to remove. The check takes two minutes: tick each count in the note against the report, and read "Nothing else changed" last, because that is the one claim the model can't verify for itself.
Payroll data is the wrong place to be casual
A payroll file holds bank details, pay, identification numbers, sickness records and sometimes court orders: some of the most sensitive data a small firm processes, belonging to other companies' employees. The rules that follow from that:
- Use AI only on business plans that don't train on your content by default (ChatGPT Business, Claude Team, or Copilot if you're on Microsoft 365), never a staff member's personal account.
- Paste the minimum. The extraction step needs names and hours, not bank details or identification numbers. The query step needs two payslips with an ID, not the employee's name.
- Never give an AI tool write access to the payroll system.
- Check your client contracts and your privacy notice mention the software you use. The pseudonymising routine in anonymising client data before you paste it into AI works for payslips too.
For a single payslip query, the difference is easy to see. Before: a staff member pasted both payslip PDFs whole, carrying the employee's full name, home address, identification number, bank account details and year-to-date totals. After: they paste two short blocks headed "EMP-117", each holding only basic pay, overtime, tax, pension, other deductions and net pay for that period. The draft reply comes out the same, because none of the removed fields had anything to do with the explanation.
The mistakes AI makes with payroll input
- Two people with the same first name. In one illustrative case, a members' club had two bar staff with the same first name. The email said "[first name] 18 hrs"; the model matched it to the first of the two on the list, who hadn't worked that week. Fix: make the prompt refuse to match on first name alone when there's more than one, and put it in "needs checking".
- Times read as decimals, as with Employee A's 7:30 above.
- Filling in the missing. Asked for a complete table, a model may insert contracted hours for someone the client didn't mention. The template should allow a blank, and a blank should trigger a question.
- Double-counting. A client sends corrected hours as a reply to their own email; the model extracts both versions. Paste only the latest message, or tell it to use the most recent figure for each person.
- Staff trusting the summary. After a few clean weeks, checkers skim. Keep a spot-check: one employee per client per run compared by hand against the source.
What a 70-client bureau can expect
Run the numbers for an illustrative four-person bureau with 70 clients and about 1,900 payslips a month. Before AI, the team logs around 160 employee and client queries a month, about 8 per 100 payslips, and makes 12 corrections after payday, each needing a supplementary payment or a recovery letter. Rekeying odd-format input takes about 25 hours a month.
Three months after adding extraction, the variance explanations and the "what changed" note, a realistic target is input handling cut to roughly 12 hours, queries down to 5 or 6 per 100 payslips, and post-payday corrections halved. The note does most of the work on queries; the variance check does most of the work on corrections. Four seats on a business chat plan cost about $100 a month on monthly billing.
Converting those targets into hours makes them easy to sanity-check. Input handling falling from 25 to 12 hours saves 13. Queries dropping from about 160 to about 105 a month (5.5 per 100 payslips), at roughly ten minutes each including checking, saves about 9. Six fewer post-payday corrections, at around 45 minutes each for the supplementary payment and the letter, saves about 4.5. That is roughly 26 hours a month against about $100 in seats. If your own sum comes out far lower, the checking step is probably eating the gain, and that is the part to trim first.
Measure the same three numbers monthly: queries per 100 payslips, corrections after payday, and hours spent on input. If queries fall but corrections don't, your variance thresholds are too loose. If corrections fall but staff are working later, the checking step has grown and needs trimming. And if a client still sends hours at 4pm on cut-off day, chasing missing client records with AI covers the reminder sequence that moves them earlier.
Further reads
- AI Document Processing: Turn PDFs and Scans Into Usable Data — The wider options for reading PDFs and scans into data.
- How to Set Up Human Review for AI Work Without Slowing Down — Keep sign-off quick without letting unchecked drafts through.
- AI Error Log: Track Mistakes and Stop Them Happening Again — Log each payroll slip so the same one doesn't recur.
- How to Build a Company Knowledge Base AI Can Answer From — Build the answer library your query replies draw on.
- How to Classify Business Data Before Using AI Tools — Decide which payroll fields may go near an AI tool at all.
- How to Set a Baseline Before You Introduce AI — Measure queries and corrections before you change anything.
- AI Tools and AI Development: The Complete 2026 Guide — the AI hub, including every tutorial in the AI-for-business series.
Sources: ChatGPT Business, Claude Team and Microsoft 365 Copilot Business plan pages (pricing, business-data defaults), checked September 2026.