Why Is AI Bad at Maths? What to Check in Quotes and Invoices

Coding Liquids tutorial cover featuring Sagnik Bhattacharya for Why Is AI Bad at Maths? What to Check in Quotes and Invoices.
Coding Liquids tutorial cover featuring Sagnik Bhattacharya for Why Is AI Bad at Maths? What to Check in Quotes and Invoices.

Chat assistants predict the next piece of text rather than calculate, and they see numbers as chunks of characters, not quantities. So an assistant can explain a pricing method perfectly and still multiply the figures wrongly. It becomes reliable only when it runs real code or a spreadsheet does the sums. For quotes and invoices, never send unchecked AI arithmetic.

The good news is that the errors are predictable. Once you know why they happen and which eight kinds reach real quotes and invoices, a few minutes' checking catches them, and a small change to how you quote stops AI doing the sums at all.

Follow me on Instagram@sagnikteaches

What's happening inside: tokens, not numbers

A language model reads and writes text in small chunks called tokens. A word might be one token; a number like 2,440.00 might be split into several pieces, such as "2", ",", "440" and ".00". The model doesn't hold 2,440 as a quantity. It holds a sequence of fragments it has seen in similar places before.

Connect on LinkedInSagnik Bhattacharya

From billions of examples, it has learned that "7 × 8 =" is usually followed by "56". That's pattern-matching, and for small, common sums it works well. Longer calculations are different. Multiplying 8.9304 by 48 means carrying digits across several steps, and the model predicts each digit of the answer in turn. One wrong digit early on flows into everything after it, and the answer still looks like a tidy number with two decimal places. The format is right; the value isn't.

Subscribe on YouTube@codingliquids

Two things have improved this. Reasoning models work through a problem step by step before answering, which cuts arithmetic errors considerably but doesn't remove them. And some assistants can run code. ChatGPT, for example, can write and run Python for data analysis, and OpenAI's help page on data analysis advises reviewing the generated code and assumptions before relying on the result. When the code runs, the computer does the arithmetic, and the error that's left is in what the model told the code to calculate.

That's the same root cause as invented facts, which AI hallucinations explained for business owners covers in more depth: the model produces what's likely, not what's checked.

Eight errors that show up in real quotes and invoices

Only the first of these is arithmetic in the narrow sense. The rest are errors of method, and they're more common.

ErrorWhat it looks likeHow to spot it
Arithmetic slip3 × $428.66 written as $1,258.98 (it's $1,285.98)Recalculate every line total
Right method, wrong inputTwo fitters in the notes, one in the quoteCompare quantities with the job sheet
Margin confused with markup"30% margin" on a $100 cost priced at $130; that's a 30% markup. A 30% margin needs $142.86Check which one your pricing uses and test one line
Discount on the wrong baseDiscount taken off everything instead of materials only, or after tax instead of beforeRead your terms, then check which lines the discount touches
Unit mix-upsMillimetres multiplied as if they were metres; price per square metre applied to linear metresCheck units on every line and area calculation
Inconsistent roundingArea rounded to 2.98 m² before multiplying, so every line drifts by a few centsRound at line totals only, the same way every time
Date arithmetic"30 days" counted from the wrong date, or end-of-month terms ignoredCount the due date yourself for one invoice
Invented or remembered figuresA rate from last month's chat, or a plausible "market rate" that isn't yoursEvery rate must match your current price list

The margin error is worth seeing in full, because the arithmetic in it is flawless. Say a small furniture restorer asks: "Price these three jobs at a 30% margin. Costs: $100, $240 and $65." An illustrative reply:

Job 1: $100 + 30% = $130.00
Job 2: $240 + 30% = $312.00
Job 3: $65 + 30% = $84.50
All three are priced at a 30% margin.

Every sum is right and every price is wrong. Adding 30% to cost is a 30% markup, which gives a margin of about 23%. A 30% margin means the profit is 30% of the selling price, so the price is cost ÷ 0.7: $142.86, $342.86 and $92.86. On those three jobs the restorer would leave $52.08 on the table, and the reply's last line confidently says otherwise. In a spreadsheet, =B2/(1-0.3) settles it, and the question to ask yourself first is which of the two your pricing actually uses.

Taking tax out of a price is the same trap in reverse. Ask for "the price before tax on $110, which includes 10% tax" and a common reply is $99, reached by knocking 10% off $110. But the tax was added to the pre-tax price, not taken from the total, so the right answer is $110 ÷ 1.1 = $100. A $1 slip on one line becomes a real gap across a month of receipts, and the reply reads just as confidently either way.

Dates go wrong in a similar, quiet way. An invoice dated 17 March on terms of "30 days from the end of the month" is due on 30 April: the end of March, plus 30 days. Drafts often give 16 April instead, counting 30 days from the invoice date because that's the more common phrase. The customer who pays on 30 April is then chased two weeks early by an automated reminder, for an invoice that wasn't late.

Worked example: a sign maker's AI-drafted quote

An illustrative sign maker pastes rough job notes into a chat assistant and asks for a quote table. The notes say: three printed aluminium composite panels, each 2,440 × 1,220 mm, at $48 per square metre; 14 linear metres of cut vinyl lettering at $6.50 a metre; two fitters for 3.5 hours at $55 per fitter per hour; 10% discount on materials only for a repeat customer; tax at 10% (an illustrative rate).

Here's what came back, next to what it should have been:

LineAI draftCorrect
Panels8.94 m² × $48 = $429.128.9304 m² × $48 = $428.66
Vinyl lettering14 m × $6.50 = $91.0014 m × $6.50 = $91.00
Fitting3.5 h × $55 = $192.502 fitters × 3.5 h × $55 = $385.00
Subtotal$712.62$904.66
Discount10% of everything: −$71.2610% of materials ($519.66): −$51.97
Tax at 10%$71.26, worked out on the pre-discount subtotal$85.27, on the discounted $852.69
Total$721.62$937.96

Five problems, and the draft looks entirely professional. The area was rounded too early. The second fitter disappeared. The discount was applied to labour as well as materials. Tax was worked out on the wrong base. And the stated total, $721.62, doesn't even match the draft's own lines, which come to $712.62. Sent as written, the quote undercharges by $216.34, about 23% of the job, and the customer has it in writing.

A 10-point check for any quote or invoice AI has touched

  1. Quantities match the job sheet, order or enquiry, line by line.
  2. Rates match your current price list, not a figure from memory or an earlier chat.
  3. Units are consistent: measurements converted to the unit the rate uses, and areas worked out in square metres, not square millimetres.
  4. Each line total equals quantity × rate. Recalculate them, don't eyeball them.
  5. The lines add up to the subtotal.
  6. Discounts apply only to the lines they should, in the order your terms say.
  7. Tax is calculated on the right base at the right rate.
  8. Rounding happens once, at line totals, the same way on every document.
  9. Dates and payment terms are counted correctly from the right starting date.
  10. The total in the email or summary matches the total in the table, digit for digit.

The fastest way to run points 4 to 8 is to paste the AI's lines into your usual quote spreadsheet and let the formulas recalculate. Any difference between the spreadsheet's total and the AI's is a flag. For proposals where the figures are claims rather than calculations, how to catch made-up figures in AI-drafted proposals covers the other half of the problem.

Invoices coming in: when AI reads your supplier bills

The same weaknesses apply in the other direction. If an AI tool reads supplier invoices into your accounts, it's extracting numbers from a document, and extraction errors look just like the arithmetic errors above: a 7 read as a 1 on a scanned bill, a unit price taken as a line total, a tax figure picked up from the wrong column, a credit note treated as an invoice.

Three checks catch most of them. First, have the system confirm that the extracted lines add up to the invoice's own printed total, and flag any mismatch rather than accepting either figure. Second, compare each bill with the purchase order or the supplier's usual amount, and look at anything more than 10 to 20% out of pattern. Third, sample a handful of processed invoices each week against the original PDFs, especially from suppliers whose invoices are scanned or handwritten.

Here's how the first check earns its keep, in an illustrative café. A dairy invoice has the line "24 × oat milk 1L @ $2.15 = $51.60". The extraction tool records the line total as $2.15, taking the unit price from the wrong column. On its own the bill now looks $49.45 cheaper than it is, and if nobody compared totals, the café would underpay and receive a chaser a month later. Because the tool checks the extracted lines against the invoice's printed total, the mismatch is flagged the moment the bill comes in, and someone spends 30 seconds fixing one field.

None of this takes long, and it's much cheaper than finding the error when the supplier chases a short payment.

Making an assistant calculate properly

If you do use a chat assistant for a quote, tell it to use code and to show its working in a form you can check:

Use code to do every calculation in this task. Don't do any
arithmetic in your head.

Produce a table with these columns: item, quantity, unit,
rate, formula used, line total.

Rules:
- Convert all measurements to metres before working out areas.
- Apply the discount only to lines I've marked [DISCOUNT],
  before tax.
- Tax rate: [X]%, applied to the discounted subtotal.
- Round to 2 decimal places at line totals only.

After the table, list every assumption you made and any figure
you couldn't find in my notes. Don't fill gaps with estimates.

Job notes:
[paste notes here]

Run on the sign maker's notes from the worked example, an illustrative extract of what comes back:

item     | qty    | unit | rate  | formula used            | line total
Panels   | 8.9304 | m2   | 48.00 | 3 x 2.44 x 1.22 x 48.00 | 428.66
Vinyl    | 14     | m    | 6.50  | 14 x 6.50               | 91.00
Fitting  | 3.5    | h    | 55.00 | 3.5 x 55.00             | 192.50

Assumptions:
- "two fitters for 3.5 hours" read as 3.5 hours in total.
- No delivery or travel charge found in the notes.
- Discount applied to Panels and Vinyl only, as marked.

The code did every sum correctly, and the panels line is now right to the cent. The fitting line is still wrong, because the model decided "two fitters for 3.5 hours" meant 3.5 hours between them. The difference is that this time the mistake is written down in the assumptions list, where it takes 20 seconds to spot and one word to correct ("each"). That's why the prompt asks for assumptions: code removes the arithmetic errors, and the list exposes the method errors.

In ChatGPT you can usually expand the analysis step to see the code it ran. If there's no code step at all, it didn't calculate; it predicted. Not every assistant or plan can run code, so check what yours does before relying on this prompt. Even with code, run the 10-point check: code does the sums correctly, but it calculates whatever the model decided the sums were.

Better still: keep the sums out of the AI entirely

The most reliable arrangement splits the work by what each tool is good at:

  1. AI reads the messy input. It turns an email, voice note or site notes into structured fields: item, quantity, dimensions, finish. This is language work, and AI is good at it. A customer email like "Need 2 signs for the shop front, about 2.4 by 1.2, plus lettering on the window, maybe 14 metres? Could you fit next week?" becomes panel_qty: 2, width_mm: 2400 (approx), height_mm: 1200 (approx), vinyl_m: 14 (customer estimate), fitting: yes, requested: next week, plus a flag: "dimensions approximate, measure on site before final quote".
  2. Your spreadsheet or quoting software calculates. Rates come from your price list, formulas do the maths, and the same inputs always give the same total. How to build a job costing sheet in Excel with AI help shows how to set one up, using AI to write the formulas rather than the numbers.
  3. AI writes the words around the result. The covering email refers to the total your software produced; it never retypes the figures.
  4. A person approves before sending. How to set up an approval step for AI-written quotes makes that step quick enough to keep.

When AI works inside the spreadsheet, as with Copilot in Excel or Gemini in Google Sheets, the same principle holds: have it write formulas you can inspect, not paste in values. What Copilot in Excel can and can't do with your numbers goes through where that works well. For the full enquiry-to-quote flow, see how to create quotes and estimates in minutes with AI.

Before trusting the new set-up, test it on answers you already know. Take three past quotes whose correct totals you've checked by hand, ideally one simple job, one with a discount and one with odd measurements, and run their original notes through the whole chain. All three totals must match to the cent. For the sign maker, the job that catches problems is the worked example above: if the extracted fields say one fitter, or the spreadsheet discounts the fitting line, the total won't come to $937.96 and you know which step to fix. Repeat the test whenever you change the price list, the prompt or the spreadsheet.

Where AI is genuinely useful with numbers

None of this means keeping AI away from figures altogether. It's good at explaining how a calculation works, suggesting a formula, spotting the one line in a long list that looks out of pattern, sense-checking whether a total is in the right range, and asking what you forgot to price.

Spotting the out-of-pattern line is the most useful of these, because it plays to pattern-matching rather than arithmetic. Say a small cleaning company pastes in last month's 60 supplier invoice lines with the prompt "Which of these look unusual compared with the rest? Give the line and your reason. Don't recalculate anything." An illustrative reply:

Line 23: "Cleaning supplies, $4,500.00". The other 11 lines from
this supplier are between $380 and $520. Possible decimal error.
Line 41: "Window cleaning poles x 2, $0.00". Zero-value line;
may be a missing price or a free replacement.
Line 57: Same invoice number as line 12, same amount ($212.40).
Possible duplicate.

All three are worth a look, and none needs the model to calculate anything: it's comparing amounts and spotting a repeated invoice number. You then check each one against the original invoice, which is where the figure is either confirmed or corrected. Use it as a sharp-eyed assistant who can't be trusted with a calculator, and it earns its place in your quoting.

More questions about AI and numbers

Are the newer reasoning models good enough to trust with sums?

They're much better than older models, because they work through a problem step by step before answering. They still slip on long calculations with decimals, units and several rules at once, and the answer looks equally confident either way. For anything a customer will pay, have the sums done by code or a spreadsheet, and check the result.

Can Copilot in Excel or Gemini in Sheets get calculations wrong too?

The spreadsheet itself calculates correctly. What the AI can get wrong is the formula it writes: the wrong range, a missing row, an absolute reference that should be relative. The advantage is that a formula is visible and checkable. Click into the cell, read the formula, and test it on a row you can work out by hand.

Why did it give a different total when I asked the same question again?

Chat assistants generate each answer fresh, with some built-in variation, so an arithmetic slip can appear in one answer and not the next. Two different totals are a clear sign that neither was calculated properly. Ask it to use code for the calculation, or move the sums into a spreadsheet where the same inputs always give the same result.

Further reads

Sources: OpenAI help centre article on data analysis with ChatGPT (checked September 2026). The worked example uses illustrative prices and an illustrative tax rate.

Want quotes where AI never touches the arithmetic?

On a 1:1 call we'll trace how a quote moves from enquiry to invoice in your business, and set it up so AI handles the wording while your spreadsheet or software does every sum.

Book a 1:1 call with me