Use traditional OCR when you need searchable text from clean, typed pages, or when the layout never changes. Use AI document extraction when you need specific fields (totals, dates, names) from documents whose layouts vary, such as payslips from dozens of employers. Most small businesses need AI extraction plus a validation check, with OCR underneath.
The difference matters because the two fail in opposite ways. OCR gets characters wrong but never invents a value: a smudged 8 becomes a 3, and you can see the garbled text. AI extraction reads the page for meaning, so it can pick the right figure from an unfamiliar layout, but it can also return a confident, well-formatted number that isn't on the page. Your choice is really about which failure you can catch more cheaply.
Three technologies hiding behind the word "OCR"
Vendors blur these, so it helps to separate them before comparing anything.
- Traditional OCR (optical character recognition) turns an image of a page into characters, with their positions. Tesseract, the open-source engine, and the scan-to-text features in PDF editors work this way. Output: a text layer. It doesn't know which number is the total.
- Template or zonal OCR adds rules on top: "the invoice number is always in the box 40mm from the top right". Excellent when every document comes from the same form; it breaks the day a supplier redesigns their invoice.
- AI extraction comes in two flavours. Pre-trained document models (such as Amazon Textract's expense and lending analysis, or the prebuilt invoice and receipt models in Azure's and Google's document services) are trained for specific document types and return fields with confidence scores. General AI models (Claude, GPT, Gemini reading a PDF or photo) can extract anything you describe in a prompt, but don't give reliable confidence scores.
Most small-business tools you'll meet, including receipt capture in accounting software, are some mix of all three.
Same payslip, two outputs
Here's what each approach hands you from one payslip. Traditional OCR, illustrative output:
PAY ADVICE Period 06 Pay date 28/09/2026
Employee J. SAMPLE Emp No 00417
Basic Pay 160.00 18.75 3,000.00
Overtime 12.00 28.13 337.56
Gross Pay 3,337.56
Tax 512.40
Pension 5% 166.88
Net Pay 2,658.28
Year to date Gross 20,025.36 Tax 3,074.40
Everything is there, but it's a block of text. To use it you still need a person, or rules, to decide that 2,658.28 is net pay and that 20,025.36 is a year-to-date figure, not this month's.
AI extraction with a prompt asking for named fields, illustrative output:
{
"employee_name": "J. Sample",
"pay_date": "2026-09-28",
"pay_frequency": "monthly",
"basic_pay": 3000.00,
"overtime": 337.56,
"gross_pay_period": 3337.56,
"net_pay": 2658.28,
"gross_pay_ytd": 20025.36,
"source_text_net_pay": "Net Pay 2,658.28"
}
That's ready to drop into a spreadsheet or case file. But notice "pay_frequency: monthly": the payslip says "Period 06", and the model inferred monthly. It's probably right, but it wasn't read, it was deduced. That is the kind of field you either forbid ("only report what's written") or verify.
Eight criteria side by side
| Criterion | Traditional OCR | Pre-trained document AI | General AI model |
|---|---|---|---|
| Varied layouts | Poor unless each gets a template | Good for the types it was trained on | Good for almost anything |
| Invents values | No | Rarely | Sometimes, confidently |
| Confidence scores per field | Per character or word | Yes, per field | No reliable score |
| Handwriting | Weak | Moderate | Often better, still check |
| Tables across pages | Weak | Good | Good, may merge rows |
| Cost per page | Near zero | About $0.01-$0.07 on Textract | Fractions of a cent to a few cents |
| Consistency run to run | Identical | Identical | Can vary slightly |
| Set-up effort | Low (text) or high (templates) | Moderate, usually via an app | Low for a prompt, moderate to automate |
Read the table as a trade: general AI models give you flexibility and low set-up effort, and charge you for it in checking. Pre-trained models sit in the middle and are often the best choice when your documents are standard types like invoices, receipts and ID.
What a thousand pages a month costs each way
Illustrative sums at list prices, before any staff time:
- Tesseract or your PDF editor's OCR: $0 in usage. The cost is the rules or the person who turns text into fields.
- Amazon Textract, first-tier list prices on its pricing page: plain text detection is $0.0015 a page ($1.50 per 1,000); expense analysis for invoices and receipts $0.01 a page ($10); the lending document analysis $0.07 a page ($70); forms, tables and queries together $0.07 a page. New accounts get a three-month free tier, including 1,000 pages a month of text detection and 100 of expense analysis.
- Azure Document Intelligence has a free tier of 500 pages a month for testing; paid rates vary by region, so check its calculator.
- A general AI model through an API: Anthropic's PDF documentation says a page typically uses 1,500-3,000 text tokens, plus image tokens because each page is also read as an image. Call it about 4,500 input tokens and 400 output tokens a page. On Claude Haiku 4.5 ($1 per million input tokens, $5 per million output) that's roughly $0.0065 a page, or about $6.50 per 1,000 pages. The Batch API halves it if you can wait for results.
At small-business volumes the usage cost is rarely the deciding factor; a few dollars a month either way. What decides it is how much checking each approach needs, because staff time at even $25 an hour dwarfs any per-page fee. The tutorial on what automated invoice processing costs per invoice adds up the full picture including review time.
How a three-adviser mortgage practice handles case documents
Consider an illustrative mortgage brokerage with three advisers and a case administrator. It runs about 20 cases a month, each arriving with around 25 documents: three months of payslips, three months of bank statements, photo ID, proof of address, details of existing credit commitments and sometimes an employer's letter. Roughly 1,400 pages a month.
Before: the administrator keyed income, regular outgoings and commitments into the fact-find, about 50 minutes a case, or 17 hours a month. Errors surfaced when a lender's underwriter queried a figure, which delayed offers.
What they tried first: template OCR from their scanning software. It worked on the two big employers' payslips and failed on everything else. Payslip layouts differ by payroll software, and the templates needed constant fixing.
What they settled on:
- Every PDF runs through OCR first so it has a searchable text layer for the archive.
- Payslips and bank statements go to a general AI model through an automation, with a prompt that asks for named fields, a quote of the source text for each, and null when a field isn't printed.
- A validation step compares documents against each other: the net pay on each payslip must match a salary credit on the bank statement for the same month within $1; three pay dates must be consecutive; year-to-date gross must rise by roughly one month's gross each time.
- Any field that fails a check, or came back null, is highlighted for the administrator. Everything else is pre-filled.
Illustrative result: about 80% of cases pass all checks and take 10-15 minutes to review; the rest take 25-30 minutes. Monthly keying time drops from about 17 hours to about 6. API usage runs to around $10 a month. The practice kept one rule: the adviser still reads the income figures against the documents before anything goes to a lender, because the adviser is responsible for the recommendation, not the software.
The cross-document check is the part worth copying. A single extracted figure is a claim; two documents agreeing is evidence.
Field definitions and checks that make AI extraction safe
Most extraction errors trace back to a vague field name. "Amount" on an invoice could mean net, tax, gross or the balance still owed. Write each field the way you'd brief a new temp, and the model follows it far more reliably:
Extract these fields from the attached payslip. Return JSON only.
- pay_date: the date labelled "Pay date" or "Payment date", as YYYY-MM-DD.
If the day and month could be read either way, return null and explain.
- gross_pay_period: gross pay for THIS period only, never year-to-date.
- net_pay: the amount labelled Net Pay / Take-home / Amount paid.
- overtime_amount: the money value for overtime, not the hours.
- For every field include source_text: the exact characters you read.
- If a field is not printed on the page, return null. Never estimate.
Then add checks that don't depend on the model at all. Five kinds cover most business documents:
- Arithmetic: line items add up to the subtotal; net plus tax equals gross; gross minus deductions equals net.
- Cross-document: the payslip's net pay appears as a credit on the bank statement; the invoice total matches the purchase order.
- Range: a monthly salary between $500 and $30,000; an invoice date no more than 90 days old and not in the future.
- Format: dates parse, account numbers have the right number of digits, a supplier name exists in your contact list.
- Presence: required fields aren't null; every value has source text.
A spreadsheet can run all five with ordinary formulas, and a failed check sends the document to a person. That's the step that turns "the AI is usually right" into "we know which ones to look at".
Where AI extraction goes wrong and OCR doesn't
These are realistic failures, and how each tends to show up:
- A plausible value that isn't there. Asked for an "employer reference", the model returned a neat alphanumeric code that appeared nowhere on the payslip. It was caught because the prompt demanded source text and the source field was empty. Fix: require a quote for every value and reject values without one.
- The wrong figure from the right area. Year-to-date gross returned as this period's gross on a payslip with an unusual layout. The consistency check (period gross roughly one-sixth of YTD in period 6) flagged it.
- Two documents merged into one. A PDF containing two months of payslips came back as a single record with the later month's figures. Fix: split multi-page PDFs first, or ask the model to return one record per pay date and count them.
- Helpful normalisation. "Overtime 12.00 hrs" came back as overtime of $12.00 because the prompt asked for amounts. Tighter field definitions fixed it: "overtime_amount is the money column, not hours".
- Different answers on a re-run. Running the same statement twice gave slightly different lists of regular payments. For anything you file, run once, store the output, and review that version.
For a wider view of when generative models are the wrong tool entirely, see generative AI vs traditional AI.
Where traditional OCR goes wrong and AI doesn't
- A supplier changes their invoice design. Template OCR silently reads the invoice number from where it used to be, which is now the date. AI extraction doesn't care where fields sit.
- Tables that wrap across pages. A bank statement's transaction table continues on page 3 with no header. Plain OCR gives you text in reading order with the columns jumbled; AI and pre-trained models usually reassemble the rows.
- Handwritten additions. A client's handwritten note on a form ("salary increases to 3,400 from October") is read badly or skipped by OCR. General AI models often read it correctly, which is useful, and risky if nobody notices it came from handwriting. The tutorial on turning handwritten forms into spreadsheet data covers that case in detail.
- Photos taken at an angle. A crumpled receipt photographed in a van. OCR character accuracy falls sharply; AI models are more forgiving, though you should still reject images below a basic quality bar.
The video production company that needed both
A different illustrative business shows the hybrid in a lighter form. A video production company handles about 150 documents a month: kit-hire invoices, freelancer invoices, location permits and receipts from shoots. It doesn't need an API at all.
- Supplier invoices go to the accounting software's own document capture, which uses pre-trained extraction and shows its guess for review. Standard document types, so the pre-trained route is ideal.
- Shoot receipts are photographed into the same capture feature. The bookkeeper checks every receipt over $100.
- Signed location permits and release forms only need to be searchable, not turned into data. Plain OCR in the PDF editor, then filed. Using AI here would add cost and risk for no benefit.
- Once a quarter, the producer asks an AI assistant on the company's business plan to list every permit expiring in the next 90 days from the OCR'd files, and checks each against the document.
Three document types, three different answers. That's normal.
Testing both on 30 of your own documents
Don't choose from a demo. Run this test in an afternoon:
- Pick 30 documents that represent your real mix, including five bad ones: a skewed photo, a multi-page table, a handwritten note, an unusual layout, a two-in-one PDF.
- List 5-8 fields you actually need from each type.
- Type the correct answers yourself into a sheet. This answer key is the whole test.
- Run each approach and paste its results next to the key.
- Score every field as correct, wrong but visibly wrong (garbled, blank), or wrong and plausible. The last category is the dangerous one.
An illustrative filled-in summary from a test like this:
Documents: 30 (12 invoices, 10 receipts, 8 payslips) Fields scored: 186
Correct Visibly wrong Plausibly wrong
Plain OCR + person n/a (text only - person extracted fields)
Pre-trained expense model 152 27 7
General AI model + quotes 171 9 6
General AI, no quote rule 169 4 13
Notes: 5 of the 6 plausible errors in the quoted run were on the two-in-one
PDF and the handwritten note. Without the quote rule, plausible errors doubled.
The pattern in that last line is common: asking for source text doesn't make the model much more accurate, but it makes its mistakes far easier to catch. Choose the approach with the fewest plausibly wrong fields that you can't catch with a validation rule, not the one with the highest raw score.
Picking the right approach for each document type
Summarising the choice for a small business:
- Need the document searchable or archived, not turned into data: traditional OCR. Cheap, exact, no invention.
- One fixed form you control (your own intake form): template OCR, or better, replace the paper form with an online one.
- Standard types from many senders (invoices, receipts, ID): a pre-trained document model, usually inside a tool you already use.
- Varied or unusual documents, or fields no pre-trained model offers (payslip line items, clauses, handwritten notes): a general AI model with quotes, null-when-absent rules and a validation step.
- Anything sensitive: whichever approach you use, process it on a business plan or API that doesn't train on your data, keep the retention period short, and don't let extracted figures reach a client, lender or regulator without a person checking them.
For taking the extracted fields onward into your CRM or accounts, AI document processing covers the end-to-end workflow, and AI invoice processing for supplier bills focuses on the accounts-payable case.
Further reads
- AI vs Rule-Based Automation: Which Does Your Task Need? — The same rules-versus-AI choice applied to whole workflows.
- Why Is AI Bad at Maths? What to Check in Quotes and Invoices — Why extracted totals still need an arithmetic check.
- Can AI Sort and File Incoming Documents for Your Business? — Sort documents by type before you extract anything.
- How to Redact Personal Data From Documents With AI Before Sharing — Strip personal data from documents before they go anywhere else.
- How to Stop Retyping Data Between Apps With AI Automation — Send the extracted fields straight into your other apps.
- How Accurate Is ChatGPT? What Owners Should Expect by Task — Realistic accuracy expectations for general AI assistants.
- Go Paperless Before You Add AI: A Step-by-Step Plan — Which paper to kill, which to scan and which to leave in the cabinet, so AI tools have clean files to work from. Six steps, about eight weeks.
- How to Automate Vaccination Checks for a Grooming Salon — Three layers of vaccine-check automation: software that watches expiry dates, AI-assisted certificate reading, and a full pipeline with a privacy catch.
- AI Submission Intake: Stop Re-Keying Data Into Insurer Portals — Key once, check once: extract client documents into a master record, validate it with simple rules, then feed insurers by API, rater or copy-ready blocks.
- Paraplanning With AI: What to Automate and What to Keep Human — Twelve paraplanning tasks sorted into automate, assist and keep human, with worked examples of data extraction, chasers and the checks that keep it safe.
- What Mortgage Brokers Can Automate With AI, and What Stays Advice — The line between admin and advice across a mortgage case, with wording that stays on the right side of it and three automations that quietly cross it.
- Can AI Read Emailed Orders Into a Wholesaler's System? — When AI can reliably turn emailed orders into sales orders for a wholesaler, what it depends on, and three different wholesalers' answers.
- Fleet Admin for Small Businesses: Automate Services, Checks and Fuel — Run a small fleet's services, daily checks and fuel records from one register, with reminders, phone forms and AI that reads receipts and flags odd fuel use.
- Compliance Calendar: Track Licences, Insurance and Certificates — Track every licence, policy and certificate in one register, let AI pull the dates from new documents, and send reminders early enough to act on.
- Can AI Read Delivery Notes and Update Stock Automatically? — How AI reads delivery notes into goods-in records, when stock can update without a person, the mapping table that matters most, and the tests to run first.
- AI Tools and AI Development: The Complete 2026 Guide — the AI hub, including every tutorial in the AI-for-business series.
Sources: Amazon Textract pricing page; Azure AI Document Intelligence pricing page; Anthropic PDF support documentation; Anthropic and OpenAI API pricing pages. Checked September 2026.