How Accurate Is ChatGPT? What Owners Should Expect by Task

Coding Liquids tutorial cover featuring Sagnik Bhattacharya for How Accurate Is ChatGPT? What Owners Should Expect by Task.
Coding Liquids tutorial cover featuring Sagnik Bhattacharya for How Accurate Is ChatGPT? What Owners Should Expect by Task.

ChatGPT is highly accurate when it works from material you give it, like rewriting your email or summarising a contract, and least accurate when it answers from memory about niche, recent or numerical facts. There's no single percentage: treat drafts as dependable, factual claims as leads to check, and sums as reliable only when it runs code.

OpenAI says much the same in its own help article on whether ChatGPT tells the truth: it can give wrong dates and facts, invent quotes, studies and references, and sound sure of itself when it's wrong, so OpenAI tells users to treat an answer as a first draft, not a final source. Its researchers add a useful detail. Models rarely slip on facts that are written about constantly, but they guess on facts that appear only rarely in their training data, and facts about a small business are about as rare as facts get.

Follow me on Instagram@sagnikteaches

Two questions that predict whether an answer will be right

Before you trust an answer, ask where its facts came from. If you pasted the supplier's spec sheet, the price list or last month's emails, ChatGPT is transforming text that sits in front of it. Its mistakes then look like omissions and small distortions: a clause left out, two figures merged, a promise added during the tidy-up. If you asked it to recall something, it is rebuilding an answer from patterns it learned in training, and where those patterns run thin it fills the gap with something plausible.

Connect on LinkedInSagnik Bhattacharya

The second question is how often that fact has been written down anywhere. OpenAI's September 2025 research paper, Why Language Models Hallucinate, makes the point with birthdays. A famous scientist's birthday appears again and again in training text, and models get it right. A birthday mentioned once, in an obituary, gets guessed. The paper argues that if a fifth of birthday facts appear only once in the training data, a model fresh from its first training stage should be expected to get at least a fifth of birthdays wrong, and that later training still rewards a confident guess over "I don't know". Your shop's opening hours, your supplier's minimum order and a small producer's vintage notes all sit firmly in the "seen once, if at all" group.

Subscribe on YouTube@codingliquids

Two smaller checks round this out. Is the answer a calculation written out in prose? OpenAI's help article credits its code tool, which it calls data analysis, with enabling accurate calculations, and the research paper notes that web search may not help with counting slips. And is the fact recent? Models are trained up to a cut-off date, so anything that changed since then depends on the search tool, which OpenAI says is enabled for all models by default.

Accuracy by task: what owners can expect

No vendor publishes an accuracy figure for "writing my supplier emails", so the table below sorts common business jobs into bands by where the facts come from and how the failures usually look. Use it to decide how much checking each job needs, then measure your own tasks with the test further down.

TaskFacts come fromWhat to expectTypical slipCheck before use
Rewriting or tidying your own textYouDependableAdds an offer or changes a numberRead once against your original
Summarising a document you uploadYouMostly rightLeaves out a clause, merges figuresAsk for clause references; spot-check three
Pulling figures from invoices or spec sheetsYouMostly right on clean text, weaker on scans and photosMisreads a digit, skips a boxed noteCheck every figure a customer will see
Sorting emails, reviews or orders into categoriesYouMostly right, wobbly on borderline casesThe same email lands in different categories on different runsTest on 20 real examples; send "unsure" to a person
Translating customer messagesYouGood for everyday wordingLoses tone, mistranslates trade termsA fluent speaker for anything published
Sums, margins and unit conversionsYou, plus its reasoningReliable when it runs code, shaky in proseWrong formula, carried roundingSee the calculation; redo one line yourself
Spreadsheet formulasIts trainingUsually fine for common formulasWrong range, ignores blanks or numbers stored as textTest on five rows with known answers
Well-known general factsIts trainingMostly rightOversimplifiesCheck if a decision rests on it
Niche facts: small producers, suppliers, peopleIts trainingA lead, not an answerInvents plausible detailVerify at the original source
Facts about your own businessVery littleUnreliableMixes you up with a similar nameGive it your facts instead
Current prices, rules and product changesSearch or stale trainingA lead, not an answerOutdated page, wrong pageOpen the cited page and check its date
Quotes, statistics and referencesIts training or searchOften invented without searchA study that doesn't existFind the original or drop it
Legal, tax, food-safety or HR specificsTraining and searchDon't rely on it aloneConfident, generic, wrong for youA qualified adviser

The pattern holds right down the table: the more of the answer that comes from you, the safer it is. That's why most of the fixes below move facts out of ChatGPT's memory and into material you supply. For the deeper reasons models make things up, see AI hallucinations explained for business owners.

Jobs built on your own material, and how they slip

Rewriting: the offer that crept in

A specialty coffee roaster asks ChatGPT to tidy a quick reply to a café asking about wholesale. The owner's draft reads: "hi, thanks for asking, we do wholesale from 6kg a week, prices on the sheet attached, we deliver tues and fri, can send samples". The illustrative output is smooth:

"Hi [first name], thank you for getting in touch about wholesale. We supply cafés from 6kg a week, and our current price list is attached. We deliver on Tuesdays and Fridays, with free delivery on every order, and we'd be happy to send samples of our full range."

The two phrases in bold were never in the draft. This roaster charges for delivery under 12kg and only samples its two house blends. Neither slip is dramatic; the model polished the message the way an eager salesperson might. The fix is one line in the prompt, "Don't add any offer, price, promise or product that isn't in my draft", plus a habit of reading every rewrite for anything new rather than for style.

Summaries: the clause that went missing

A craft brewery uploads a nine-page malt supply agreement and asks for a summary of the key terms. The summary gets price, delivery windows and quality specs right. It leaves out clause 11: a minimum annual purchase of 40 tonnes, with a charge for any shortfall. Nothing in the summary is false, which is exactly why the gap is dangerous, because the brewer reads "key terms" as "every term that matters".

Ask differently: "List every obligation on us, with its clause number and a quoted line. Include anything about minimums, renewal, price changes, penalties or termination. If a category has nothing, write NONE FOUND." Then search the PDF yourself for five words: minimum, renew, terminate, penalty, exclusive. It takes two minutes and catches what a summary skipped.

Extraction: the allergen note in a footnote box

A delicatessen photographs supplier spec sheets and asks ChatGPT to list allergens for new shelf labels. The illustrative result, set against the sheets:

ProductChatGPT's allergen listWhat the spec sheet says
Chilli and fennel salamiNoneMilk, in a boxed footnote under "ingredients from other suppliers"
Aged sheep's cheeseMilkMilk
Olive tapenadeNoneNone
Seeded crackersSesameSesame; may contain gluten (factory statement at the foot of the page)

Two of four rows are wrong, and both errors sit outside the main ingredients panel. Photos and scans are where extraction gets shaky, a limit covered in more depth in what ChatGPT can and can't read in PDFs, spreadsheets and photos. For allergens the rule is simple: AI can draft the label, a person checks every line against the original sheet, and a "none" is never accepted until someone has read the whole page.

Sums: the margin that was really a mark-up

A butcher asks: "Rib-eye costs me $18.40 a kilo. I want a 30% margin. What should I charge for a 250g steak?" The illustrative answer comes back as prose:

"$18.40 × 1.30 = $23.92 per kilo, so a 250g steak should be priced at $5.98."

The arithmetic is fine and the method is wrong. Adding 30% to cost is a 30% mark-up, which only produces a margin of about 23%. For a true 30% margin, the price is cost divided by 0.70: $26.29 a kilo, or $6.57 for a 250g steak. The model took the most common reading of an ambiguous request. Ask it to "calculate with code, show the formula, and say whether you're using margin or mark-up", then redo one line on a calculator. A 59-cent gap per steak adds up across a week's sales and a 60-line price list.

Jobs that lean on its memory, and how they slip

Product facts: the tasting note with a medal nobody awarded

A wine merchant asks for a 60-word shelf note on a small producer's 2021 red, giving only the producer and the wine's name. The illustrative output, with the problems marked in brackets:

"Deep ruby, with black cherry, violet and a lick of spice. Aged 18 months in new oak [the producer's sheet says 12 months in used barrels], it's full-bodied at 14.5% [the sheet says 13.5%] with fine, ripe tannins. A gold medal winner in 2023 [no medal on record: invented]. Pair with lamb or aged hard cheeses."

Three errors in 60 words, each specific and believable, from a model working on a fact it has probably seen once or never. The corrected note came from pasting the producer's technical sheet and adding "Use only facts from this sheet. If something isn't there, leave it out": "Deep ruby, with black cherry and violet. Twelve months in used barrels keeps the fruit up front; 13.5%, medium to full body, fine tannins. Pair with lamb or aged hard cheeses." Shorter, less exciting, all true.

Your own business: it may describe someone else

A farm shop owner asks ChatGPT for the shop's opening hours and what customers say about it. With search on, the answer may come from an old directory listing, a farm shop with a similar name, or reviews of a café that once shared the site. Without search it may produce generic text in a confident voice. This is the rarest kind of fact there is, so don't ask the model to know it. Keep a one-page facts file (hours, products, delivery wording, returns policy) in a ChatGPT project and attach it whenever you draft anything customer-facing. If the wrong version is already turning up in answers your customers see, what to do when ChatGPT gets facts about your business wrong covers the fixes.

Current prices, rules and product changes

Anything that changed after the training cut-off comes either from search or from stale memory. Search helps, but the cited page can itself be out of date, or be a reseller's page rather than the supplier's. Take a brewery asking for the current price of a canning-line service plan and getting a figure from a two-year-old forum post. Open at least one cited page, check its date and whether it belongs to the supplier, and treat the number as unconfirmed until you've seen it there. Catching outdated information in AI answers has a fuller routine.

Quotes, statistics and references

Ask for "three statistics about coffee subscription cancellations, with sources" and you may get three tidy percentages, each attached to a survey name and a year. Without search, OpenAI lists fabricated studies and references as a known failure; with search, the link may be real but not contain the number. The rule for anything you publish: if you can't open the original and find the figure in it, the figure doesn't go out. Checking sources and citations in AI research shows how to do that quickly.

Measure it yourself: a 20-case accuracy test

The only accuracy figure that matters is the one for your task, your prompt and your source material, and you can get it in an afternoon.

  1. Pick one job you'd actually hand over. Keep it narrow: "shelf notes from producer sheets", not "marketing".
  2. Gather 20 past examples where you know the right answer. Use real ones, including a few awkward cases.
  3. Run all 20 exactly as you would in practice. Same prompt, same files, same plan and settings.
  4. Score each output in one of four boxes: correct as it stands; minor fix (wording, tone, length); factual error (wrong, but harmless if caught); harmful error (would mislead a customer, break a rule or cost money).
  5. Decide from the two error boxes, not from how good the best outputs looked.
Task: ________________________  Prompt version: ____  Date: ________

Case | Correct | Minor fix | Factual error | Harmful error | Note
  1  |         |           |               |               |
  2  |         |           |               |               |
 ... |         |           |               |               |
 20  |         |           |               |               |

Serious error rate = (factual + harmful) / 20 = ____ %
Decision: hand over / hand over with checks / keep doing it by hand
Re-test when: the prompt, the model or the source format changes

Here's the test run by an illustrative wine merchant with 40 new wines a season, each needing a shelf note written from the producer's technical sheet. On the first run the prompt simply said "write a 60-word shelf note from this sheet". Of 20 notes, 11 were correct, 5 needed minor fixes, 3 had factual errors (a vintage taken from the wrong row of a two-vintage sheet, an alcohol figure rounded up, "unfiltered" claimed where the sheet said nothing) and 1 was harmful: "suitable for vegans" on a wine whose sheet didn't mention how it was fined. That's a serious error rate of 4 in 20, or 20%.

The merchant then tightened the prompt: use only facts from the sheet, write NOT STATED where the sheet is silent, and never claim vegan, organic, awards or scores unless the sheet says so. The second run gave 15 correct, 5 minor fixes and no factual or harmful errors. Twenty clean cases don't prove it will never slip, so a quick check stays in place: every name, number and claim against the sheet.

The time sums still work with the check included. Writing a note by hand took about 12 minutes; drafting with ChatGPT and checking it against the sheet takes about 4. For 40 wines that's 480 minutes against 160, a saving of 320 minutes (5 hours 20 minutes) a season. The merchant's rule for future jobs: above 10% serious errors, fix the prompt or keep the job manual; between 1% and 10%, hand it over with a line-by-line fact check; zero in 20 cases, hand it over with a lighter check and re-test whenever anything changes.

Settings and habits that raise accuracy on everyday work

  • Give it the source and fence it in. "Use only the material below. If it isn't there, write NOT STATED." This single instruction removes most invented detail from drafting jobs.
  • Keep standing facts in a project. Put the price list, policies, hours and product sheets in a ChatGPT project so every chat starts from the same facts. On a personal plan, check the model-training switch in privacy settings before uploading business files; business plans don't train on your content by default.
  • Use search for anything current, then open the pages. OpenAI's own advice is to check sources by visiting the links directly when accuracy matters.
  • Push sums into code. Ask it to calculate with code and show the formula, and say which figure it treated as the input.
  • Ask for the line behind each claim. "For each figure, quote the sentence in the source it came from" turns a summary into something you can check in seconds.
  • Ask twice on borderline calls. If two runs sort the same complaint differently, the case goes to a person, and the instructions probably need a clearer rule.
  • Freeze the prompt that passed. Save the version that did well on your 20 cases, don't edit it casually, and re-test after any change.

A second pass that only checks, and never rewrites, catches a surprising share of slips. Paste this into a fresh chat with the source and the draft:

Check this draft against its source. Do not rewrite it.

SOURCE:
[paste the spec sheet, email thread or price list]

DRAFT:
[paste the draft]

1. List every factual claim in the draft: names, numbers, dates,
   prices, promises, allergens, awards, certifications.
2. For each claim, quote the source line that supports it,
   or write NOT IN SOURCE.
3. List anything important in the source that the draft left out.

An illustrative reply for the coffee roaster's email would flag "free delivery on every order: NOT IN SOURCE" and "samples of our full range: NOT IN SOURCE (source says 'can send samples')". The checker is still ChatGPT, so it can miss things too; treat it as a second pair of eyes before your own, not instead of them. For a fuller routine you can run on anything that leaves the business, see a five-minute fact-check routine for AI output.

How much checking each kind of output needs

BandExamplesCheckWho checks
DependableRewrites of your own emails, formatting, first drafts you'll edit anywayRead once for anything new: offers, prices, namesWhoever sends it
Mostly rightSummaries, extraction from clean documents, sorting enquiriesSpot-check against the source; check every number that leaves the businessSomeone who knows the source
A lead onlyNiche facts, current prices and rules, statistics, anything about your business from memoryVerify each claim at the originalThe owner or a named checker
Not on its ownLegal, tax, food safety, allergens, HR, medicalA qualified person signs offAn adviser or trained staff member

A few signals in an answer deserve a second look whatever the band: an exact figure with no source, an award, score or certification, a quotation, a detail you never supplied, anything that reads like a policy, and an answer that changes when you ask the same question again. OpenAI's help article puts the underlying problem in four words worth pinning above the desk: confidence isn't reliability. Once the job has been through the 20-case test, the prompt is frozen and the checks are assigned to someone by name, ChatGPT is accurate enough for a great deal of daily work, and you'll know from your own numbers which jobs those are.

More questions owners ask about ChatGPT's accuracy

Is ChatGPT more accurate on a paid plan?

OpenAI says the tools you get, such as deep research, depend on your plan, and higher plans allow more use of its stronger models, which helps on hard, multi-step questions. The basic pattern doesn't change: work from your own sources stays reliable and recall of niche facts stays risky. Run your 20-case test on the plan you'd actually use before deciding to upgrade.

Is ChatGPT more accurate than Claude or Gemini?

It depends on the task and on which model version each app is running that month, and all three change often. Published benchmark scores rarely match business work like shelf notes or supplier summaries. Run the same 20 cases, with the same prompt and source files, through each assistant you're considering and compare the serious error counts. That comparison is worth more than any ranking.

Does ChatGPT learn from my corrections?

It uses your correction for the rest of that conversation. It won't reliably carry the fix into new chats unless the right fact sits somewhere it reads every time, such as the files or instructions in a ChatGPT project. Memory features change often, so don't rely on them for prices, policies or allergen details. Put business facts in a document you control.

Can I trust the sources ChatGPT links to?

Treat each link as a pointer, not proof. With search on, the page is usually real, but it may be old, be a reseller's copy, or not contain the figure quoted. Open it and find the sentence. Without search, references can be invented outright, which is why OpenAI's own guidance tells users to verify quotes, data and references before relying on them.

Further reads

Sources: OpenAI help article 'Does ChatGPT tell the truth?' (updated mid-2026); OpenAI research paper 'Why Language Models Hallucinate' (Kalai, Nachum, Vempala and Zhang, 4 September 2025); OpenAI help pages on Projects in ChatGPT.

Want to know which of your tasks ChatGPT can take?

On a 1:1 call we'll go through the jobs you'd hand to ChatGPT, sort them by how much checking each needs, and set up the source files and review steps that keep errors out of customer-facing work.

Book a 1:1 call with me