How to Check Sources and Citations in AI Research

Coding Liquids tutorial cover featuring Sagnik Bhattacharya for How to Check Sources and Citations in AI Research.
Coding Liquids tutorial cover featuring Sagnik Bhattacharya for How to Check Sources and Citations in AI Research.

Check AI research citations in three passes. First confirm each source exists: open the link, resolve the DOI, search the exact title in quotes. Then confirm it says what the AI claims by finding the exact figure or quote on the page. Finally, judge whether it's worth citing: primary, current and independent. Anything you can't trace doesn't get published.

ChatGPT fake sources come in more than one kind. The fully invented citation, with a plausible title and a link that goes nowhere, is the famous one. In web-connected tools the more common problem is subtler: a real page that doesn't say what the AI says it does. Checking only that links open catches the first problem and misses the second.

Follow me on Instagram@sagnikteaches

The worked example follows an illustrative software reseller writing a buyer's guide for its small-business customers on backing up Microsoft 365 data. It used an AI research tool to gather evidence, and the guide nearly went out with eight citations, of which only three held up. Examples from an events company, a sign maker, a print shop and a clothing shop show the same failures in different kinds of research.

Connect on LinkedInSagnik Bhattacharya

Why AI tools produce citations that don't hold up

A language model writes by predicting plausible text. A citation is a very predictable pattern (author, year, title, publication, link), so a model can produce one that looks exactly right without any real source behind it. OpenAI's help pages say plainly that ChatGPT can produce incorrect or misleading output while sounding confident, and advise checking quotes, data and references. OpenAI's own research on the subject argues that standard training and evaluation reward guessing over admitting uncertainty.

Subscribe on YouTube@codingliquids

How much that matters depends on how the tool works:

  • A chat assistant answering from memory, with no web search, is the riskiest. Any "source" it gives is reconstructed from patterns, so treat it as a suggestion of what to search for, never as a citation.
  • A chat assistant with web search on links to pages it actually retrieved. Invented links become rarer, but it can still attach a claim to the wrong page or misread a figure.
  • Deep research modes, such as ChatGPT's deep research, run many searches and produce a report with citations or source links. OpenAI lets you point deep research at specific sites, or restrict it to them, which helps a lot for business research. The report is longer, so there are more claims to check.

The general mechanism is covered in the tutorial on AI hallucinations for business owners. For research, the practical upshot is simple: the AI's job is to find candidates, and your job is to confirm them. The comparison of Perplexity and ChatGPT for business research covers which tools make that easier.

Five ways an AI citation fails, and the check for each

FailureWhat it looks likeThe check that catches it
Invented sourcePlausible title, real-sounding publisher, link that 404s or a DOI that doesn't resolveOpen the link; paste the DOI into doi.org; search the exact title in quotes on Google Scholar and the web
Real page, wrong claimThe page exists and is on topic, but the figure or statement isn't on itUse your browser's find function (Ctrl+F or Cmd+F) for the exact number or a key phrase
Right source, wrong detailsThe figure is from a different year, edition, country or sample; a quote is paraphrased as if verbatimCompare the date, edition and exact wording with the original
Weak sourceA vendor's marketing page, an anonymous listicle, a forum post, a page that itself looks AI-writtenAsk who published it, who paid for it and whether it shows its method
Withdrawn sourceA study that has since been retracted or correctedSearch the title in the Retraction Watch database, which Crossref makes freely available

In practice these three passes happen in order: exist, support, worth it. Most time goes on the second.

Pass 1: does the source exist? (about two minutes each)

Start with the cheapest check. For each citation:

  1. Open the link. A 404, a redirect to a homepage or a page about something else is a red flag. Dead links aren't always invented; pages move. Search the title to see whether it lives elsewhere, or check an archived copy.
  2. Resolve any DOI. A DOI (digital object identifier) is the permanent ID of an academic paper. Paste it into doi.org; a real one takes you to the publisher's page for that paper.
  3. Search the exact title in quotation marks. Google Scholar's own help suggests searching a paper's title to check whether it's included. An invented paper usually returns nothing, or only near-misses with different authors.
  4. Check the pieces match. Invented references are often assembled from real parts: real authors who never wrote together, a real journal, a volume number that doesn't exist for that year. If the author list or year differs from what Scholar shows, treat the citation as unverified.

An illustrative invented citation from the reseller's first draft had everything a real one has: two plausible author surnames, a 2023 date, the title "Recovery Time in SaaS Backup Failures", a journal name, and a volume, issue and page range. The journal name returned only unrelated results, the title returned nothing in quotes, and the DOI the AI supplied resolved to an unrelated chemistry paper. Three minutes, and it was out.

Pass 2: does it say what the AI claims?

This is the pass people skip, and the one that catches most problems in web-connected tools. Open the source and find the exact claim. Use the browser's find function for the number ("37%"), then read the surrounding paragraph to check it means what the AI says it means.

Here is an illustrative example of the mismatch, from the same guide. The AI wrote:

AI claim: "Most small businesses that lose their Microsoft 365 data never recover it, according to [link to a backup vendor's page]."

The linked page was real. It was a backup vendor's product page, and it said something much narrower: that deleted items are only recoverable for a limited period without a separate backup, and that customers "should consider" a third-party backup. Nothing about "most small businesses" and nothing about "never". The AI had turned a product page's sales argument into a statistic.

What makes pass 2 manageable is asking the AI to do some of the work up front. Ask for the exact quote supporting each claim along with the link (see the prompt further down). Then your check becomes "is this quote on this page, and does it mean this?", which takes a minute instead of five.

Three things to compare every time:

  • The number itself. 37% and 73% are one keystroke apart. Figures also get rounded up ("nearly half" from 41%).
  • The population. "Businesses with more than 250 staff" quietly becomes "businesses". A survey of 400 IT managers becomes "research shows".
  • The date. A figure from a 2019 survey presented as current is a common error, and it matters in fast-moving areas such as security or AI adoption. The tutorial on catching made-up figures in AI-drafted proposals covers the number-checking side in more depth.

Pass 3: is it a source you'd put your name to?

A source can exist and say exactly what the AI claims and still not be worth citing. Customers judge your guide by its weakest reference. The reseller now scores each source quickly against five questions; three or more "no" answers means find something better.

QuestionStrong answerWeak answer
Is it the original?The survey, study, standard or official documentation itselfA page quoting another page quoting a survey
Who published it?An official body, a vendor's own documentation about its own product, a recognised publisherAn anonymous site, a content farm, a page with no author or date
Who paid for it?Independent, or the sponsor is disclosed and the method is shownA vendor's survey that happens to prove you need the vendor's product
Is it current?Recent enough for the topic (months for AI tools and prices, years for stable facts)Old figures presented as current
Does it show its method?Sample size, dates, how questions were askedA headline number with no method

Vendor sources aren't banned. Microsoft's documentation is the right source for how Microsoft's products work. The problem is a vendor's marketing claim cited as independent evidence about the market.

Eight citations in a reseller's buyer's guide, checked

Here is the whole worked example. The reseller asked a deep research tool for evidence on four questions for its guide: what Microsoft 365 retains by default, how often small businesses lose data, what backup costs, and what customers should ask a backup provider. The draft came back with eight citations. Checking them took about 50 minutes; the illustrative results:

#Claim in the draftResultAction
1Default retention for deleted itemsHeld up: Microsoft's own documentation, currentKept, with the page title and date checked
2Shared responsibility for customer dataHeld up: vendor documentation about its own serviceKept
3What to ask a backup providerHeld up: an independent industry guideKept, reworded in the reseller's own words
4"Most small businesses never recover lost data"Real page, claim not on it (a vendor product page)Claim removed
5A percentage of businesses hit by data lossReal survey, but from 2019 and only large firmsRemoved; no current equivalent found
6Average cost of downtime per hourFigure traced to a vendor-funded survey with no method shownReplaced with a worked example using the customer's own numbers
7A quote attributed to an analystQuote not found anywhere in that wordingRemoved
8"Recovery Time in SaaS Backup Failures" (2023)Invented: title, journal issue and DOI didn't match anythingRemoved

The finished guide had three citations instead of eight and was stronger for it. Where the draft had a dubious downtime statistic, the final version gave customers a quick sum to do themselves: "If eight staff can't work for four hours, at an average cost of $40 an hour each, that's 8 × 4 × $40 = $1,280, before you count lost orders." A reader can check that, and it describes their business rather than someone else's survey.

The reseller also kept a simple sources log in a spreadsheet: claim, source title, link, date accessed, the exact supporting quote, who checked it. When a customer asked where a figure came from six months later, the answer took a minute to find. Two rows from its log, filled in (the quotes are shortened placeholders here):

Claim as publishedSource and dateSupporting quoteChecked by, onRecheck by
Deleted items are recoverable for a limited period without a separate backupVendor documentation page on recovering deleted items, updated 2026"[exact sentence stating the recovery window]"Owner, 3 SeptemberMarch (vendor docs change)
Downtime cost worked exampleOwn calculation, no external source8 staff × 4 hours × $40Owner, 3 SeptemberOnly if the example changes

The "recheck by" column is the one people forget. Vendor documentation, prices and product limits change often, so any claim built on them needs a date when someone looks again. Stable facts, such as a published standard, can go years between checks.

Prompts that make AI research easier to check

You can't stop an AI tool making mistakes, but you can make its mistakes easy to find. This prompt asks for everything pass 2 needs:

Research this question for a guide aimed at small-business customers:
[question]
Search the web. Prefer official documentation, regulators, standards bodies
and independent research. Avoid vendor marketing pages as evidence for
market-wide claims.
For EVERY factual claim, give:
- the claim in one sentence
- the source title, publisher and publication date
- the URL
- a verbatim quote of up to 40 words from the source that supports it
If you cannot find a verbatim quote that supports a claim, label the claim
UNSUPPORTED rather than leaving it in. Do not include any statistic you
cannot quote from its original source.

An illustrative piece of the output, and what still needed fixing:

Sample output: Claim: Deleted mailbox items can be recovered for a limited period by default. Source: vendor documentation page on recovering deleted items, dated 2026. URL: [link]. Quote: "[a sentence stating the default recovery window]". Claim: Many businesses underestimate recovery time. UNSUPPORTED: no verbatim quote found.

The first claim still needed pass 2: the quote was real, but it came from a page about a different mailbox type, so the default it described didn't apply to the customers the guide was for. The UNSUPPORTED label worked exactly as intended; the claim was dropped rather than decorated with a link. If your tool supports it, restricting deep research to a list of sites you trust (official documentation, a regulator, one or two industry bodies) cuts down the weak-source problem before it starts.

Research traps in other kinds of business

Sources fail in similar ways whatever the subject. Some illustrative cases:

  • An events company's trends report. The AI cited "a 2025 industry survey" showing average wedding guest numbers. The survey was real but covered a different market from the one the company works in, and the figure was a median presented as an average. The company rewrote the section using its own bookings from the last two years, which was more relevant anyway.
  • A sign maker's materials guide. The AI claimed a vinyl film lasts "up to 12 years outdoors", citing a manufacturer's datasheet. The datasheet was real, but the figure applied to one premium film in specific conditions; the sign maker uses a different grade. Manufacturer datasheets are good sources, but only for the exact product they describe.
  • A print shop's sustainability page. An AI draft said the shop's paper was "certified carbon neutral", linking to a paper mill's page. The mill's page described a different product line. A claim like that isn't only a sourcing error; environmental claims have to be substantiated, so the shop removed it until the supplier provided a certificate for the stock it actually buys.
  • A clothing shop's fabric care blurb. The AI cited a textile association for "linen gets stronger when wet". Searching the exact phrase found it repeated across dozens of shop pages, none with an original source. The shop kept the practical care advice and dropped the scientific claim.

In every case the fix was the same: trace the claim to the original, check it applies to your situation, or leave it out. If you work in a field where citations have legal weight, the stakes are higher still: courts have fined lawyers for filing documents that cited cases AI had invented. The tutorial on stopping AI inventing case law sets out a checking routine for small firms.

When a claim can't be verified

You will often find a claim that is probably true but can't be traced. Don't publish it with a hopeful link. Choose one of these instead:

  1. Replace the statistic with a worked example using the reader's own numbers, as the reseller did with downtime costs.
  2. Give a range and say where it comes from, if you have several sources that roughly agree.
  3. Describe the evidence honestly: "We couldn't find reliable figures on this; here's what we see with our own customers."
  4. Cut it. A guide with three solid citations is more persuasive than one with eight shaky ones.

Before anything built on AI research goes out, run it through a five-minute fact-check routine as the last step, and keep your sources log up to date. If you publish marketing claims based on this research, the tutorial on backing up every claim AI writes explains how to keep the evidence on file.

Checking AI citations: more questions

Are deep research tools safe from fake sources?

They are better, because they search the web and link to the pages they used, which gives you something to check. They still misread pages, pick weak sources and attach claims to links that don't support them. Treat every citation as a lead to verify, not proof, whichever tool produced it.

How long should checking a citation take?

Usually two to five minutes: under a minute to confirm the source exists, a couple of minutes to find the exact claim on the page, and a moment to judge the source. Statistics that need tracing back to an original survey can take 15 minutes or more, which is a good reason to use fewer of them.

Can I cite a source the AI summarised if I can't open it?

No. If you can't read the source yourself, because it's behind a paywall or the link is dead, you can't know whether it says what the AI claims. Find an accessible copy, use a different source you can read, or leave the claim out.

What should I do if I've already published a fake citation?

Correct it quickly and visibly. Remove or replace the citation, fix any claim that depended on it, and add a short correction note with the date if the piece is public. Then add the check that would have caught it to your routine so it doesn't happen again.

Further reads

Sources: OpenAI help pages (Does ChatGPT tell the truth?; Deep research in ChatGPT); OpenAI research summary on why language models hallucinate; Google Scholar search help; Crossref documentation on the Retraction Watch database. Checked September 2026.

Want AI research you can actually rely on?

On a 1:1 call we'll look at the research your business does with AI, set up prompts that make every claim checkable, and agree a checking routine that fits the time you have.

Book a 1:1 call with me