AI Document Processing: Turn PDFs and Scans Into Usable Data

Coding Liquids tutorial cover featuring Sagnik Bhattacharya for AI Document Processing: Turn PDFs and Scans Into Usable Data.
Coding Liquids tutorial cover featuring Sagnik Bhattacharya for AI Document Processing: Turn PDFs and Scans Into Usable Data.

AI document processing reads PDFs, scans and photos, works out what each document is, pulls the fields you need into a structured table and checks them before they reach your spreadsheet, CRM or accounting system. For a few dozen documents a month, a chat assistant with a saved prompt is enough; for hundreds, use an extraction tool that exports automatically.

The split that decides accuracy is the input. Digital PDFs, which contain their text, are easy. Scans, phone photos, handwriting and complex tables are where errors appear, so scan quality and a validation step matter more than which AI you pick. Most failed projects chose a good tool and skipped the checks.

Follow me on Instagram@sagnikteaches

Digital PDFs, scans and photos behave differently

A quick test tells you which you have: open the PDF and try to select a word with your cursor. If you can, the text is inside the file. If the whole page highlights as one block, it's an image, and the tool must first run optical character recognition (OCR, turning a picture of text into text) before any AI can read it.

Connect on LinkedInSagnik Bhattacharya
InputWhat the tool has to doTypical trouble
Digital PDF (exported from software)Read the text and layoutTables split across pages
Scanned PDFOCR, then readSkewed pages, faint print, stamps over text
Phone photoCrop, straighten, OCR, then readShadows, angles, curled pages
Handwritten formHandwriting recognition, then readSimilar digits (1/7, 0/6), crossings-out
Word or spreadsheet fileRead directlyMerged cells, hidden sheets

Older OCR systems needed a template per layout; modern AI extraction understands layouts it hasn't seen, which is why a new supplier or a changed form no longer breaks everything. AI extraction against traditional OCR explains the trade-offs if you're choosing between them.

Subscribe on YouTube@codingliquids

The six stages of a document pipeline

Whether it's a chat prompt or a paid tool, the same stages happen. Knowing them tells you where to look when a row comes out wrong.

  1. Capture. Documents arrive by email, upload, scanner or phone. Give them one entry point, such as a dedicated email address or folder, so nothing is processed twice or missed. For paper, the scanner in the OneDrive or Google Drive app works well; Microsoft Lens, which many people used, was retired in early 2026.
  2. Clean-up. Straighten, crop, improve contrast, split multi-document scans into separate files. Good tools do this automatically; bad input still defeats them.
  3. Classify. Decide what each document is (an application form, a certificate, a contract) because each needs different fields.
  4. Extract. Pull the defined fields into a structured record, often with a confidence score per field.
  5. Validate. Test the record against rules: formats, totals, allowed values, cross-checks with data you already hold. Send anything that fails, or has low confidence, to a person.
  6. Export. Write the checked record to its destination: a spreadsheet, CRM, accounting system or database, with a link back to the original file.

Define the output before choosing a tool: the field spec

The most useful hour in any document project is writing down exactly which fields you want, in what format, and what counts as valid. Tools differ far less than people expect once this exists. A filled-in spec for an illustrative estate agency's seller property questionnaire:

FieldFormatValid ifIf missing
property_addressTextMatches an address in the CRMStop: human review
bedroomsWhole number1-10Flag
heating_typeOne of: gas, electric, oil, heat pump, otherIn listFlag
boiler_install_yearYYYY1980 to this yearLeave blank
parkingOne of: none, street, allocated, garage, drivewayIn listFlag
annual_service_chargeNumber, no symbol0-20,000Leave blank
known_defectsText, as writtenAnyWrite NONE STATED
seller_signature_presentYes / NoYesStop: human review

Three details in that spec do most of the work. Fixed lists ("one of") stop the AI inventing categories. "As written" for free text stops it rewording what the seller said, which matters for anything a buyer might later rely on. And each field has a rule for what happens when it's missing, so a blank never silently becomes a guess.

An estate agency turning seller questionnaires into listing data

An illustrative estate agency takes on about 45 new properties a month. Each seller completes a 12-page questionnaire, about half typed into a PDF form and half printed and filled in by hand. A negotiator retypes the key details into the CRM and the listing, at about 25 minutes a questionnaire: nearly 19 hours a month.

Option one: a saved prompt in a business chat assistant. The negotiator uploads the questionnaire and runs the extraction prompt (below). Output comes back as one row matching the field spec, which she checks against the flagged fields and pastes into the CRM import sheet. About 3 minutes to run and 5 minutes to check: roughly 6 hours a month.

Option two: an extraction tool. Questionnaires are emailed to a parsing address; a tool extracts the fields and an automation writes them to the CRM, with flagged records held for review. At 45 questionnaires of 12 pages, that's 540 pages a month, beyond the smaller plans of a tool like Parseur (Micro is 100 pages for $49 a month) and within its Starter plan (1,000 pages for $129 a month). Checking time falls to about 3 minutes a record.

What they chose. Option one, for now. The saving over option two was about two hours a month, which didn't justify $129 a month and a setup project. They agreed to revisit if volume doubled. They also asked sellers to complete the questionnaire online in future, which is the cheapest document processing of all: no document to process.

The mistake in week two. A handwritten "3" in the bedrooms box, with a small "+1 box room" noted in the margin, came out as 3. The listing went live as a three-bedroom property; the seller rang to complain that the box room wasn't mentioned. The fix was adding "margin_notes: any handwriting outside the boxes, as written" to the spec, so notes like that always surface for a human.

A prompt that returns clean rows, not prose

Extract data from the attached property questionnaire. Return ONLY one row
in CSV format with these columns, in this order:
property_address, bedrooms, heating_type, boiler_install_year, parking,
annual_service_charge, known_defects, seller_signature_present, margin_notes,
low_confidence_fields

Rules:
- heating_type must be one of: gas, electric, oil, heat pump, other
- parking must be one of: none, street, allocated, garage, driveway
- Copy known_defects and margin_notes exactly as written.
- If a field is blank or unreadable, leave it empty. Never guess.
- In low_confidence_fields, list any field you weren't sure about,
  separated by semicolons.

Illustrative output:

"12 Orchard Way",3,gas,2019,driveway,,"Damp patch rear bedroom ceiling, treated 2024","Yes","+1 box room","bedrooms;boiler_install_year"

What you'd check: the two low-confidence fields against the scan, the margin note (which changes the listing), and the blank service charge (right for a house with no shared areas, a question for a flat in a managed block). Everything else can go straight into the import sheet. The low_confidence_fields column is the part people leave out and later wish they hadn't; it tells you where to look instead of rechecking everything.

A storage facility's handwritten move-in forms

An illustrative self-storage site takes about 30 new customers a month on a paper move-in form: name, address, phone, ID type seen, unit number, start date, value of goods for insurance, and an alternative contact. Staff retype them into the unit management software.

AI extraction reads most fields well, but handwritten phone numbers are the weak spot: a 7 read as a 1 means the customer can't be reached when their payment fails. Three validation rules fixed most of it:

  • Phone numbers must have the right number of digits for your numbering plan; anything else is flagged.
  • The unit number must exist in the site's unit list and be marked vacant.
  • The start date must be within 30 days of the form's date.

Flagged fields went to the staff member who took the form, while the customer's details were fresh. Six months later the site moved to a tablet form at the counter, and the extraction step disappeared. For businesses still on paper, turning handwritten forms into spreadsheet data goes deeper into handwriting specifically.

Tables need their own instructions

Tables are where extraction most often goes quietly wrong: rows merge, a table that continues onto the next page loses its header, or a long table is summarised instead of copied. A used-car dealership, for instance, receives an auction condition report for every car it buys, with a damage table listing each panel, the damage type and an estimated repair cost. The workshop wants every line in its preparation sheet.

A prompt that asks for "the damage items" tends to return the five biggest. A prompt that works says: return one row per table row, include the page number and row position for each, copy the cost exactly, and finish with the number of rows you found. Then compare that count with the table itself. In one illustrative report, the AI returned 14 rows; the table had 16, because two rows sat at the top of page 3 under a repeated header the model treated as a new table. Asking "the table continues across pages; treat it as one table" fixed it.

Three habits help with any table:

  • Ask for a row count and check it against the document. It's the fastest completeness test there is.
  • Keep the page number on every row, so checking a doubtful row takes seconds.
  • Extract first, total second. Have the AI copy the rows, then add them up in your spreadsheet rather than asking it for the sum.

For the dealership, the repair costs from the condition report, added up in the prep sheet, became the starting point for pricing each car, which is exactly the kind of number you don't want summarised.

Validation rules that catch bad extractions

Every extraction produces some wrong fields. Rules find them before they cause trouble. The useful kinds, with examples:

  • Format checks. Dates are real dates; postal codes, phone numbers and reference numbers match their pattern; numbers contain only digits.
  • Range checks. Bedrooms between 1 and 10; an invoice total under $50,000; a start date not in the past.
  • Allowed values. Heating type from a fixed list; a unit number that exists.
  • Arithmetic. Line items add up to the subtotal; subtotal plus tax equals total.
  • Cross-checks. The address matches the CRM; the customer already exists; the vehicle identification number matches the stock record.
  • Duplicates. The same document number or file already processed.
  • Confidence. Anything the tool marks as low confidence goes to a person, regardless of the other rules.

In a spreadsheet, most of these are simple formulas in extra columns (for example, a column that shows "CHECK" when bedrooms is blank or above 10). In an automation, they're filter steps. Either way, the goal is that a person only looks at the rows that need them.

Tools by volume and budget

Monthly volumeGood fitCost (list, September 2026)
Up to about 50 documentsChat assistant (ChatGPT, Claude, Gemini) with a saved promptYour existing business plan
50-1,000 pagesNo-code parsing tool feeding a spreadsheet or app via Zapier or MakeParseur from $49 a month for 100 pages; free plan 20 pages a month
Finance documents at any volumeYour accounting software's capture featureUsually included
Thousands of pages, or built into your own systemsCloud extraction services (Google Document AI, Azure Document Intelligence, Amazon Textract)Google lists plain OCR at $1.50 per 1,000 pages and its form parser at $30 per 1,000; Azure has a free tier of 500 pages a month

Chat assistants have practical limits worth knowing. Claude accepts PDFs up to 1,000 pages but only analyses images and layout for PDFs of 100 pages or fewer, so split long scanned files. ChatGPT's help pages give a 512MB file limit and around 2 million tokens of text per document. For what breaks when you upload photos and spreadsheets, see what ChatGPT can and can't read.

Personal data inside documents

Application forms, ID documents and questionnaires are full of personal data, and some, such as ID copies taken for anti-money-laundering checks, are among the most sensitive files a small business holds. Before processing:

  • Extract only the fields you need. If you only need "ID checked: yes, passport, expiry date", don't extract the passport number.
  • Use a business plan or a tool with a data-processing agreement, and check how long it keeps uploaded files.
  • Don't run ID documents through consumer chat apps at all. Use your compliance software or a tool built for identity checks.
  • Redact what the AI doesn't need before uploading; redacting personal data with AI explains how.
  • Delete processed files from the tool once the record is exported and your retention rules allow it.

Measure accuracy on your own documents

Vendor accuracy claims are measured on their test sets, not your forms. A 30-document test takes about an hour and tells you what you need to know. Take 30 real documents, including the worst scans you have, run them through, and score each field right or wrong against the original.

30 questionnaires x 8 fields = 240 fields
Typed PDFs (15):        119/120 correct
Handwritten scans (15): 112/120 correct
  bedrooms              13/15  (both errors: margin notes)
  boiler_install_year   11/15  (digits misread; 3 flagged low confidence)
  known_defects         13/15  (one shortened, one line missed)
Records needing any fix: 8 of 30

Read the result by document type and by field, not as one percentage. Here the typed forms are close to perfect and the handwritten ones need checking on three fields, which tells you exactly where to spend review time. If a field is wrong often and never flagged as low confidence, add a validation rule for it; that's the dangerous combination. Rerun the same 30 documents whenever the form, the prompt or the tool changes, and keep the scores, so you can see whether a change helped or quietly made things worse. Once the checks work, stopping retyping between apps covers moving the rows onward automatically.

Document processing questions

What scan quality do I need for reliable extraction?

Aim for 300 dots per inch, flat pages, even lighting and all four corners visible. Phone scanning apps such as the scanner in the OneDrive or Google Drive app crop and straighten pages automatically, which helps more than a higher resolution. Avoid photos taken at an angle, shadows across the text and pages still folded; those cause more errors than handwriting does.

Can AI read handwriting accurately?

Neat block capitals in boxes, usually well. Joined-up writing, crossings-out and numbers squeezed into small spaces, much less reliably, and digits such as 1 and 7 or 0 and 6 get confused. Treat handwritten fields as needing validation rules and a quick human check, and where you can, replace the paper form with a digital one so there's nothing to read.

Do I need a developer to set up document processing?

Not for most small-business volumes. A chat assistant with a saved prompt needs no setup, and no-code extraction tools connect to spreadsheets and apps through Zapier or Make. You'd only need a developer for cloud extraction services called directly from your own systems, which usually only makes sense at thousands of pages a month.

Further reads

Sources: Google Cloud Document AI pricing; Azure Document Intelligence pricing; Parseur pricing; Claude help article on file uploads; OpenAI file uploads FAQ. Checked September 2026.

Got a pile of documents someone retypes every week?

On a 1:1 call we'll look at the documents, define the fields you actually need, choose between a saved prompt and an extraction tool for your volume, and set up the checks.

Book a 1:1 call with me