Can ChatGPT Read PDFs, Spreadsheets and Photos? What Breaks

Coding Liquids tutorial cover featuring Sagnik Bhattacharya for Can ChatGPT Read PDFs, Spreadsheets and Photos? What Breaks.
Coding Liquids tutorial cover featuring Sagnik Bhattacharya for Can ChatGPT Read PDFs, Spreadsheets and Photos? What Breaks.

Yes. ChatGPT reads text-based PDFs, Excel and CSV files, and photos, on free and paid plans. What breaks is precision: scanned pages, handwriting, charts inside PDFs, merged cells and multi-tab workbooks are where it misreads or quietly skips things. OpenAI itself warns it may not reliably extract exact values from scans or image-based tables.

Size is rarely the problem; structure is. A 40-page contract with a proper text layer is usually read more accurately than one photographed receipt, because ChatGPT has to read the receipt as a picture, and pictures are where a 7 turns into a 1.

Follow me on Instagram@sagnikteaches

The upload limits matter less than you'd think

These are the caps in OpenAI's file uploads FAQ in September 2026. Free plans have tighter limits on how many files you can upload, and the exact allowance changes, so check yours if you upload often.

Connect on LinkedInSagnik Bhattacharya
File typeLimit
Any single file512 MB
Text and document files (PDF, Word and similar)2 million tokens per file
Spreadsheets and CSV filesAbout 50 MB, depending on row size
Images20 MB per image

Two million tokens is roughly 1.5 million words, far more than any quote or contract. The practical limit is attention, not size. With long documents, and with PDFs stored as project files, ChatGPT tends to look up the passages that seem relevant rather than weighing every page for every answer. Ask about clause 14 of a 90-page contract and it will usually find clause 14. Ask "is there anything unusual in here?" and it may never read page 63.

Subscribe on YouTube@codingliquids

The workaround is to make it cover the document in pieces you can see. For a long supplier agreement, ask first: "List every clause and schedule heading with its page number." Check the list reaches the last page. Then ask your question a section at a time: "Read clauses 1 to 12 and list anything about price changes, renewal or termination, quoting the clause number." Work through the remaining ranges the same way. It takes four or five prompts instead of one, and the schedules at the back, where notice periods and price-rise terms often sit, get read instead of skipped.

PDFs: a text layer or a picture of text

The single most useful test takes five seconds. Open the PDF and try to select a sentence with your cursor. If the text highlights, the PDF has a text layer and ChatGPT reads the actual characters. If it doesn't, you have a scan: a photograph of a page, and ChatGPT has to read it the way it reads any image.

There's one catch the five-second test misses. Many scanners run character recognition and hide the result behind the page image, so the text highlights even though it's really a scan, and that hidden text can be poor. Copy one line with figures in it and paste it into a plain text editor. If "Total: $4,800.00" pastes as "Tota1: $4,8OO.0O", with a digit one standing in for a letter and capital letters standing in for zeros, that garbled layer may be what gets read. Treat the file as a scan and ask for the original.

Scans are where exact figures go wrong. A faint 8 becomes a 3, a column of prices shifts one row, a decimal point disappears. OpenAI's own help page on data analysis says ChatGPT may not reliably extract exact values from image-based tables, scanned files or complex visual layouts, and recommends uploading a spreadsheet or text-based file when exact values matter.

Three other PDF features cause trouble even with a text layer:

  • Charts and images. OpenAI documents full reading of images and charts inside PDFs ("visual retrieval") for ChatGPT Enterprise, and says PDFs added as project files are processed as text only. On other plans, don't assume a chart has been read. Test it with a question only the chart can answer.
  • Tables across pages. A price table that breaks over two pages can lose its headers on the second page, so figures get attached to the wrong column.
  • Multi-column layouts, footnotes and sidebars. Text extracted from these can come out in the wrong order, joining the end of one column to the start of another.

The chart test is worth doing properly. Say a marketing agency uploads a client's 20-page annual report and asks, "Which month had the highest website traffic in the bar chart on page 4, and roughly how much?" There are three kinds of answer. A correct one that names the month and a value close to the bar's height means the chart was read. "The document doesn't state monthly traffic figures" means it wasn't, which is honest and useful. The dangerous one is a confident month and number that don't match the bar, because the model has filled the gap from text elsewhere in the report. Only the first lets you rely on anything else it says about the charts.

If you process a lot of scanned documents, such as supplier invoices or delivery notes, a dedicated document-processing tool is usually a better fit than pasting them into a chat. AI document processing for PDFs and scans explains the difference.

Spreadsheets: read with code, so structure matters

ChatGPT analyses spreadsheets by writing and running code, and the code reads the file's cells, not what you see on screen. That's why it can total 5,000 rows correctly, and also why layouts designed for human eyes confuse it. The usual culprits:

  • Headers that aren't in the first row, or headers spread over two rows ("Q1" above "Sales" and "Costs").
  • Merged cells. The value sits in only one of the merged cells; the others read as empty.
  • Subtotal and total rows mixed in with the data, so a sum counts everything twice.
  • Several small tables on one sheet, which the code may treat as one big table with gaps.
  • Hidden rows and columns, which are usually still in the file and still get counted.
  • Meaning carried by colour, such as red for overdue or green for paid. Colours generally aren't read as data.
  • Formulas. Analysis typically works from the values stored in the cells, not the formulas behind them.
  • Multiple tabs. Say which tab you mean, or it may analyse only the first.

The fix is a quick clean copy for AI: one table per sheet, one header row, no merged cells, no subtotal rows, colour-coded meaning moved into a proper column, and only the columns the question needs. For an illustrative small wholesaler's monthly sales sheet, the changes look like this:

In the original sheetIn the clean copy for AI
Title "Sales 2026" in row 1, "Q1" merged across three month columns in row 2, month names in row 3One header row: month, product_code, customer_type, units, value, returned
A "Quarter total" row after every third monthDeleted; the code can total by quarter itself
Returns shown by red textA "returned" column holding yes or no
Customer names and phone numbersRemoved; the question is about products, not people
Five tabs, one per sales repOne tab with a "rep" column, or a separate upload per question

That copy takes 15 to 20 minutes the first time and a couple of minutes each month after, if you keep it as a template. If the file is a mess, cleaning messy data walks through the tidy-up. And once it's clean, whether AI can analyse your sales spreadsheet covers what to ask and what to check.

Photos: receipts, whiteboards and site pictures

ChatGPT reads printed text in a clear, straight-on photo well. It struggles, in roughly this order, with handwriting, glare, sharp angles, low light and small print. Two other limits matter for businesses:

  • Measurements from photos are estimates. It can say a wall looks roughly three metres wide; it can't measure it. Never quote or order materials from an AI's reading of a photo.
  • Counting many similar things is unreliable. Chairs in a hall, boxes on a shelf, people in a crowd: treat the number as a rough guide.

The measurement point is easy to underestimate, so here's a quick sum. A decorator sends ChatGPT two photos of a sitting room and asks for the wall area. The illustrative reply estimates about 3.5 by 4 metres with a 2.4-metre ceiling: 2 × (3.5 + 4) × 2.4 = 36 m². The tape measure says 4.2 by 5.1 metres: 2 × (4.2 + 5.1) × 2.4 = 44.6 m², before taking off doors and windows. The estimate is about 19% short, which on a two-coat job is the difference between finishing on Friday and a second trip to the supplier. Use the photo reading to sense-check a measured figure, never to replace one.

Receipts are a special case. ChatGPT will read a clear receipt, but if you're photographing dozens a month for your accounts, a receipt-capture feature in your bookkeeping software is built for exactly that job, files the image against the transaction and keeps an audit trail, which a chat conversation doesn't.

When you do read a receipt in ChatGPT, ask for the line items as well as the total, because the two check each other. In an illustrative case, a creased receipt from a builders' merchant comes back as "Total: $117.40", but the five line items it lists add up to $177.40. The crease ran through the second digit and turned a 7 into a 1. A total that doesn't match its own lines is the quickest sign of a misread, and it costs one extra line in the prompt: "List each item and price, then the printed total, and say whether they agree."

Crop the photo to the part you care about before uploading, take it in good light, and for anything handwritten, check every figure. For a steady stream of handwritten paperwork, turning handwritten forms into spreadsheet data with AI sets up a process with checks built in.

If you can choose the format, choose this

Much of the trouble above disappears if you ask for, or save, the right format in the first place. Suppliers and colleagues will usually send a different format if you ask.

What you need from the fileBest format to uploadAvoid
Exact figures: prices, quantities, totalsA CSV or a clean single-tab spreadsheetScanned PDFs and screenshots of tables
The wording of a contract or quoteThe original digital PDF or Word fileA phone photo of the printed copy
A table that only exists in a PDFAsk the sender for the spreadsheet behind itRelying on extraction for anything you'll invoice from
Numbers shown in a chartThe data the chart was made fromThe chart alone, especially on plans that read PDFs as text
Handwritten notesA typed copy, even a rough oneLow-light photos taken at an angle

The pattern is simple: give ChatGPT the data, not a picture of the data. Every step from the original source towards an image of it adds a chance for a digit to change.

Worked example: an events company's supplier pack

Take an illustrative events company planning a 180-guest conference. The planner uploads three files and asks ChatGPT for a budget summary and catering numbers:

  1. The venue's quote: a 12-page scanned PDF.
  2. The attendee list: a workbook with two tabs, company names in merged cells, a subtotal row under each company, and dietary needs shown by cell colour.
  3. A phone photo of the handwritten run sheet from the site visit.

The answer looks polished. It's also wrong in three places. Room hire is given as $4,300; the scan says $4,800. The guest count is 196, because the 16 company subtotal rows were counted as guests. And it reports no special dietary requirements at all, because they only existed as colours. From the photo, a 14:30 start was read as 4:30.

The fixes take about 25 minutes: ask the venue for the original digital PDF, or type the six key figures into a short table; make a clean copy of the attendee list with one tab, no merged cells, no subtotal rows and a "dietary" column; and retype the run sheet, which is a dozen lines. Rerun with the clean files, and the numbers match the sources. Twenty-five minutes against a catering order 16 covers out with no vegetarian meals is an easy trade.

A five-minute test before you trust an upload

Run this whenever figures from a file will go into a quote, an order, a budget or a message to a client:

  1. Ask what it received. Page count, sheet names, number of rows and columns. If any of these is wrong, stop there.
  2. Spot-check five values with page or cell references, and look each one up in the original.
  3. Ask a question only the awkward parts can answer: the last page, the second tab, the chart.
  4. Ask what it couldn't read or wasn't sure about.
  5. Compare one total it calculates with a total printed in the file itself.

You can do most of it in one prompt:

Before answering my question, tell me:
1. What you received: number of pages, or sheet names with row
   and column counts for each sheet.
2. Any pages, tables, charts, images or handwriting you couldn't
   read, or read with low confidence.
3. Five figures from the file, each with its page number or cell
   reference, so I can check them against the original.

Then answer this question, citing the page or cell for every
figure you use: [your question]

Here's an illustrative first half of a reply, for a wholesale supplier's price list sent as a workbook:

1. Received: one sheet, "Prices", 212 rows x 6 columns.
2. Low confidence: none. No images or charts found.
3. Five figures:
   - Row 14, "Oak shelf 90cm": $38.50
   - Row 57, "Pine shelf 60cm": $21.00
   - Row 102, "Wall bracket (pair)": $6.20
   - Row 160, "Oak shelf 120cm": $52.00
   - Row 211, "Delivery, per order": $15.00

The five figures all match, and it would be easy to move straight on to the question. But the workbook has three tabs: Prices, Trade discounts and Discontinued. Line 1 shows it only opened the first, so any answer about trade prices would have been built on nothing. That's why the prompt asks what it received before it asks anything else. Name the tab you need, or upload each one as its own CSV, and run the check again.

Once you've run this a few times on a particular supplier's documents, you'll know whether their files are clean enough to trust with a lighter check, or whether they always need the full five minutes. Make a note of which is which.

If the figures pass, sums it calculates with code will be reliable. If it did the arithmetic without running code, check that too; why AI is bad at maths explains what to look for. And before uploading anything confidential, check where it's going: whether chat-with-PDF tools are safe for client files covers the data side.

More questions about uploading files to ChatGPT

Is it safe to upload client files to ChatGPT?

It depends on the plan, not the file type. An uploaded contract or spreadsheet is the same data as pasted text, so the same rules apply: identifiable client information belongs on a business plan such as ChatGPT Business or Enterprise, not a personal account. Before uploading, delete columns or pages the task doesn't need.

Can ChatGPT give me back an edited spreadsheet?

Yes. When it analyses a file with code, it can produce a cleaned or reorganised spreadsheet or CSV for you to download. Check the result before using it: confirm the row count matches the original, spot-check a few values, and look at whether cells contain formulas or fixed values, because generated files often hold values only.

Why does ChatGPT say it can't open my file?

Common causes are password protection, a damaged or unusual file format, a file over the size limit, or reaching your plan's upload allowance. Word or PowerPoint files that are really just images of scanned pages also cause trouble. Remove the password, save a fresh copy in a standard format, or export the part you need as a PDF or CSV.

Further reads

Sources: OpenAI help centre articles on file uploads, data analysis with ChatGPT and visual retrieval with PDFs (checked September 2026).

Want AI to read your documents without the errors?

On a 1:1 call we'll look at the PDFs, spreadsheets and photos your team works from, find where they trip AI up, and decide whether cleaner files or a document-processing tool fits better.

Book a 1:1 call with me