You can't stop a language model inventing authorities, so stop relying on it for them. Use AI for structure and argument, never as the source of a citation, and put every case it mentions through a fixed check: retrieve the full report, confirm the citation, read the cited paragraph, check it's still good law, then log who verified it.
The reason prompting alone won't fix this is that invention isn't a bug in one tool. A general chat assistant generates the most plausible text, and a plausible case name with a plausible citation is exactly what legal writing looks like. Even tools built on legal databases get it wrong: a Stanford study of Lexis+ AI and two Thomson Reuters research tools, tested in 2024 and later published in a peer-reviewed journal, found they produced hallucinated answers on between 17% and 33% of test queries. The products have changed since, but the lesson holds. The fix is a routine, not a better prompt.
What happens when fake authorities reach a court
The risk isn't theoretical, and it isn't confined to large firms or careless ones.
- In the best-known early case, in 2023, two lawyers and their firm were fined $5,000 after filing a brief that cited six cases ChatGPT had invented. When the court asked for copies, they filed fabricated "decisions" too, and they were ordered to write to each judge falsely named as an author.
- In June 2025, a senior court dealt with two such cases together. In one, the grounds cited five cases that don't exist. The court made wasted costs orders, referred the lawyers involved to their regulators, and stressed that lawyers have a basic duty to check the accuracy of material they put before the court.
- A public database maintained by the researcher Damien Charlotin tracks court and tribunal decisions that address AI-hallucinated content. On 27 September 2026 it listed 2,086 of them. You can browse it at the AI Hallucination Cases database.
The common thread in these decisions is rarely that someone used AI. It's that nobody read the source before relying on it.
Why models produce authorities that look real
A model writing a legal argument has learned what citations look like: party names, a year, a court, a report reference, a paragraph number. When the argument "needs" an authority and the model has no real one to hand, it produces something with the right shape. Three patterns make it worse:
- False-premise questions. "Find me the appellate case holding that X" presumes such a case exists. Models tend to oblige.
- Niche or recent points. The less often a point appears in training data, the more the model fills gaps.
- Asking it to double-check. "Are you sure that case exists?" often gets a confident yes, sometimes with an invented summary. A model can't verify against a source it doesn't have.
Here's an illustrative exchange. The prompt: "Give me two appellate authorities for the proposition that an employer's failure to follow its own disciplinary policy makes a dismissal unfair." The output:
"1. Harlow v Bexmoor Care Ltd (2019), Court of Appeal, at para 41: the court held that departure from a contractual disciplinary procedure 'will ordinarily render the dismissal unfair'. 2. Penrose v Tallis Freight Ltd (2021), Court of Appeal, at para 27..."
Both cases in that example are invented for this tutorial, which is exactly the point: they look like real authorities, with paragraph numbers and a quotable line, and nothing in the output signals doubt. The fix isn't to spot fakes by eye. It's to treat every citation from a general assistant as unverified until it has passed the routine below. There's more on the mechanism in AI hallucinations explained for business owners.
Asking the same assistant to check its own work shows the third pattern in action. A follow-up of "Are you sure Harlow v Bexmoor Care Ltd exists? Please double-check" might return, illustratively: "Yes. Harlow v Bexmoor Care Ltd is a 2019 Court of Appeal decision concerning a care worker dismissed without the hearing her contract required; the court's reasoning at paragraph 41 is frequently cited on procedural fairness." The model has now added a fact pattern and a claim about how often the case is cited, all of it invented, and the extra detail makes the fake more convincing, not less. A second answer from the same tool is not a second source.
Rewriting the request removes the false premise. Compare the two versions:
- Before: "Give me two appellate authorities for the proposition that failing to follow a disciplinary policy makes a dismissal unfair." This assumes the authorities exist and asks the model to produce them.
- After: "Here are the three decisions I've retrieved. Does any of them address whether departing from an employer's own disciplinary policy affects fairness? Quote the paragraph if so; if none does, say so." This gives the model material to read and an explicit way to say no.
In the second version a truthful answer can be "none of the attached decisions addresses that point directly", which is useful research information in itself. The first version leaves the model no honest answer except "I can't find any", and it rarely chooses that one.
The seven-step checking routine
- Draft with placeholders. When AI helps structure an argument, instruct it to write
[AUTHORITY NEEDED: proposition]wherever an authority would go, rather than naming one. The fee earner fills each placeholder from research they have done themselves. - Log every authority. Anything that ends up cited, whoever suggested it, goes into the citation log (template below). One row per authority per proposition.
- Retrieve the full text from an authoritative source. Official law reports, your subscription research database, or the court's own published decisions. Not the AI's summary, not a secondary article, not a search snippet.
- Confirm the citation details. Party names, year, court and report reference all match the retrieved report exactly.
- Read the cited paragraph. Does it actually say what the draft claims? This is where real cases fail: the case exists, but the proposition is from a dissent, an obiter remark, a different issue, or nowhere at all.
- Check it's still good law. Use a citator or your database's treatment history to see whether it has been appealed, overruled, doubted or distinguished on the point you're using.
- Verify quotations character by character and sign off. Any quoted words must appear in the report exactly. The person who checked initials the log row with the date; for court documents, a second person checks a sample.
With a subscription database, steps 3 to 7 take roughly five to ten minutes per authority for someone who knows the area, longer for unfamiliar law. That time is part of the cost of the work, and it should be budgeted like any other research time.
A quick sum makes the budget concrete. Say a small litigation team files four documents a month that cite authority, averaging eight authorities each. That's 32 authorities; at five to ten minutes each, somewhere between two hours forty minutes and five hours twenty minutes of checking a month, or roughly an hour a document. On a fixed-fee matter, that hour belongs in the estimate from the start, not squeezed out of the fee earner's evening before a filing deadline, which is the moment steps get skipped.
Step 6 is the one most often skipped by people who have done steps 3 to 5 properly, because a case that exists and says the right thing feels finished. An illustrative catch: a draft cited a first-instance decision for a point on holiday pay calculation. The case was real, the paragraph said exactly what the draft claimed, and the quote was accurate. The database's treatment history showed it had been reversed on appeal on that very point eight months later. A general assistant trained before the appeal had no way to know, and the retrieved first-instance report didn't say so either; only the citator did.
A citation log you can copy
Keep this in the matter file for every document that cites authority. Here it is filled in for three rows of an illustrative skeleton argument:
| # | Authority as cited in draft | Proposition relied on | Source retrieved from | Details match? | Paragraph supports it? | Still good law? | Checked by, date |
|---|---|---|---|---|---|---|---|
| 1 | Case A, para 23 | Procedural fairness requires a chance to respond | Subscription database, full report | Yes | Yes, para 23 | Yes, followed since | RK, 14 Oct |
| 2 | Case B, para 41 | Departure from own policy renders dismissal unfair | Not found in any source | No | n/a | n/a | RK, 14 Oct: removed, likely AI-generated |
| 3 | Case C, para 12 | Reasonableness judged at time of decision | Court's published decisions | Year wrong in draft | Point is at para 19, not 12 | Yes | RK, 14 Oct: citation corrected |
Row 3 is the more common failure in practice: a real case, cited for roughly the right point, with a wrong year and a wrong pinpoint. It would survive a check that only asks "does this case exist?" and fail the moment a judge turned to paragraph 12.
Which tools help at which step
- General assistants (ChatGPT, Claude, Gemini, Copilot): useful for step 1, structuring an argument, summarising documents you provide, and tightening prose. Never a source of authority.
- Grounded legal research tools such as Lexis+ AI, Thomson Reuters' CoCounsel, or vLex's Vincent AI (vLex is now part of Clio): useful for finding candidate authorities and a first check that a citation matches a real case. The Stanford figures above are a reminder that they still need steps 5 to 7.
- Your research database and citator: the backbone of steps 3, 4 and 6. If a small firm can't justify a full subscription, check what free official sources publish the decisions of the courts you appear in, and what your professional body provides to members.
Whether a general assistant or a legal tool is the better buy for your practice is a separate question; ChatGPT or a legal AI tool for a small firm weighs them up, and checking sources and citations in AI research applies the same logic outside law.
Prompts that lower the risk without removing it
These help, particularly when you give the model the actual decisions to work from:
Structure an argument for [proposition] using only the facts
below. Do not name any case, statute or rule. Wherever an
authority would be needed, write [AUTHORITY NEEDED: the exact
proposition it must support].
Answer using only the attached decisions. For every statement,
give the decision's name and paragraph number, and quote the
words you rely on. If the attached decisions don't support a
statement, say "not found in the attached material".
Run the first prompt on an unfair dismissal argument and an illustrative extract of the output reads: "The dismissal was unfair because the employer decided the outcome before hearing the employee's explanation [AUTHORITY NEEDED: a decision reached before hearing the employee's side is procedurally unfair]. As the courts have consistently held, an employer must also consider alternatives to dismissal." The placeholder worked. The second sentence is the problem: "as the courts have consistently held" is a claim about the case law with no placeholder attached, so it would pass a reviewer who is only looking for case names. Add a line to the prompt ("do not refer to what courts or tribunals have held, even generally") and search every draft for "held", "established", "settled" and "well known" before it goes to the supervisor.
The second prompt is far safer than open-ended research, because the model is working from documents you retrieved. It still needs checking: models sometimes attach a real quote to the wrong paragraph, or quote accurately but from a passage that is summarising a party's argument rather than the court's reasoning.
That last failure is easy to miss because everything checks out mechanically. In an illustrative case, the model answered with "'Any departure from the agreed procedure is necessarily fatal to the fairness of a dismissal' (para 15)". The quote was word for word, and it was at paragraph 15. But paragraph 15 opened with "Counsel for the claimant submits that...", and at paragraph 22 the court rejected the submission. Citing it as the court's view would have put the opposite of the decision before the tribunal. When you read the cited paragraph at step 5, read the paragraph before it as well, so you know whose words they are.
One skeleton argument, checked end to end
Here is how the routine played out on an illustrative skeleton argument in a two-solicitor employment practice. A trainee used a general assistant to structure the argument and to suggest authorities, then ran the routine before the supervising solicitor saw it.
- 11 authorities in the draft. Logging and retrieval took about 40 minutes.
- 2 couldn't be found anywhere. Both had been suggested by the assistant, both had plausible names and paragraph numbers, and both were removed.
- 1 was real but cited for a proposition from a different part of the decision, which on reading was about a separate statutory point. Replaced with the correct authority from the practice's own research.
- 1 had a wrong year and pinpoint, as in row 3 of the log.
- 7 were fine. One had been doubted in a later decision, which the supervising solicitor decided to address head-on in the skeleton.
Total checking time was about an hour and a half. Without the routine, the draft would have reached a tribunal with two invented authorities in it. The practice's response wasn't to ban the assistant but to change its instructions: from then on, the trainee used the placeholder prompt and did the research separately.
In a personal injury practice, the routine applies more narrowly. Most correspondence cites no authority at all, so the routine applies mainly to court documents on disputed points, such as quantum arguments or procedural applications, and the placeholder rule catches most risk at source.
Making the routine stick in a small firm
- Write it into your AI policy. One line does most of the work: "No authority generated by AI is cited unless it has passed the firm's checking routine and been logged." An AI acceptable use policy for a small professional firm shows where it fits.
- Make supervisors ask for the log. A supervisor reviewing a court document should see the log alongside it. If there's no log, the document isn't ready. For the second-person sample in step 7, a workable rule for the 11-authority skeleton above: the supervisor re-checks every authority that originally came from an AI tool, plus the two on which the argument turns, plus one chosen at random. That came to five of the nine left in the final draft, about 25 minutes.
- Train the people most likely to be under time pressure. Trainees and paralegals often use AI most and have the least experience spotting a wrong proposition.
- Check the court's own rules. Some courts and tribunals now publish guidance on generative AI in documents; follow it for each forum.
- Apply it to summaries, not just drafts. An AI summary of a bundle can misstate what a case or a document says just as easily; summarising bundles and transcripts with AI safely covers that side.
AI and case citations: what small firms ask next
Do I have to tell the court I used AI?
It depends on the court. Some have issued guidance or practice directions on generative AI, and a few require a statement about its use in certain documents. Check the rules of the court or tribunal you are appearing in before filing. Whatever the rule, you remain responsible for every authority and proposition in the document, however it was drafted.
What should we do if a filed document contains a citation we can't verify?
Act quickly. Tell the supervising partner, confirm the position against an authoritative source, and if the authority is wrong or doesn't exist, write to the court and the other side to correct it before it is relied on. Tell the client, record what happened, and check your regulator's reporting duties and your insurer's notification terms.
Should we check the other side's citations too?
Yes, as a matter of routine. The other side may be using AI too, and litigants in person increasingly do. Put their authorities through the same retrieval and paragraph check. If one can't be found, raise it politely and promptly with them and, where appropriate, the court, rather than saving it for the hearing.
Can I use AI to check the citations for me?
Only as a first filter. A legal research tool can quickly confirm that a citation matches a real case in its database, which is useful for spotting outright fabrications. It cannot be relied on to confirm the case says what your draft claims, so the paragraph check and the good-law check still need a person reading the source.
Further reads
- A Five-Minute Fact-Check Routine for AI Output Before It Goes Out — A general five-minute check for any AI output before it goes out.
- AI Error Log: Track Mistakes and Stop Them Happening Again — Log every caught citation error so the routine improves.
- How to Set Up Human Review for AI Work Without Slowing Down — How to design review that doesn't get skipped under time pressure.
- How Small Law Firms Use AI to Draft Letters and Routine Documents — Safe AI drafting for routine letters, where citations rarely belong.
- Per-Seat or Pay-As-You-Go? Legal AI Pricing for Small Firms — What grounded legal research tools cost a small firm.
- AI Implementation Plan for a Small Law Firm: The First 90 Days — Where this routine fits in a first-90-days AI plan.
- AI Readiness Checklist for Accountants, Solicitors, Consultants — Twenty checks, grouped and scored, that tell an accountancy, law or consulting firm whether it is ready to pilot AI or has gaps to fix first.
- AI Marketing for Small Law Firms: Content, Reviews and the Rules — How a small law firm can use AI for guides, posts and review replies while staying inside platform, consumer-law and professional conduct rules.
- Can Solicitors Use ChatGPT Without Breaching Confidentiality? — Confidentiality, privilege and data protection are separate tests for a law firm using ChatGPT. What passes each, and the settings to fix first.
- AI Contract Review for Small Firms: What It Catches and Misses — What AI contract review reliably catches, what it misses and why, with a worked review, a test method and a prompt that demands evidence.
- Which Legal Tasks Should a Small Firm Never Hand to AI? — A four-question test for legal AI risks, the seven jobs that stay with a named lawyer, and wording to put the line in your firm's AI policy.
- How to Catch Made-Up Figures in AI-Drafted Proposals — A number audit for AI-drafted proposals: find every figure, trace it to a source, recheck the sums, with a recruitment agency's 23-figure example.
- AI Tools and AI Development: The Complete 2026 Guide — the AI hub, including every tutorial in the AI-for-business series.
Sources: Damien Charlotin, AI Hallucination Cases database (count at 27 September 2026); Stanford RegLab, Hallucination-Free? Assessing the Reliability of Leading AI Legal Research Tools (2024; Journal of Empirical Legal Studies, 2025); published reports of the 2023 sanctions decision and of a June 2025 senior-court decision on fictitious citations.