Use a business-grade tool that doesn't train on your content, load a clean, text-searchable bundle with its page numbers intact, and ask for summaries in which every point carries a page or line reference. Then check those references: every point you'll rely on, plus a spot-check of at least one in five of the rest, against the original before the summary is used.
There are two separate risks, and people usually only think about the first. Confidentiality is about where the bundle goes and who can see it. Accuracy is about what comes back, and the dangerous failure there is less often an invented fact than a missing one: the summary that covers 40 documents fluently and leaves out the one email that undermines your client's account. Models summarising very long inputs can give uneven attention across them, so a summary can read as complete while quietly thinning out in the middle sections. Both risks are manageable, but only with routine.
Clear the tool before the bundle goes anywhere
Before anyone uploads a bundle, the firm should have answered these questions for the specific tool and plan. Here they are filled in for an illustrative firm on Google Workspace considering Gemini Notebook (formerly NotebookLM):
| Question | Why it matters | Illustrative answer for this firm |
|---|---|---|
| Is our content used to train models? | Client documents must not become training data | On work accounts, no, unless someone submits feedback, in which case Google may review the full interaction. Rule: no feedback buttons on client work. |
| Can the vendor's staff read it? | Human review is a disclosure risk | Google says files, chats and outputs on its enterprise-grade tiers aren't reviewed by humans. Confirm your plan qualifies. |
| Who inside the firm can see it? | Shared notebooks and Projects can over-share | One notebook per matter, shared only with the matter team |
| How long is it kept, and can we delete it? | Retention should follow the matter | Notebook deleted at file closure; noted on the closing checklist |
| Are there restrictions on using these documents? | Disclosed documents often carry limits on use | Fee earner confirms nothing in the bundle is subject to an order or undertaking that rules out third-party processing |
| Do our engagement terms cover it? | Clients may expect to be told | Engagement letter mentions secure AI tools for document review |
The same questions apply to ChatGPT Business, Claude Team or Microsoft 365 Copilot, all of which don't train on business content by default; the answers on human review, sharing and retention differ by product and plan, so check each. The professional duty sits with you regardless of tool; whether solicitors can use ChatGPT without breaching confidentiality covers the wider question.
Prepare the bundle so page references survive
Half the value of an AI summary is being able to jump from a point to the page it came from. That only works if the references the AI gives match the page numbers everyone else is using.
- Make it text-searchable. Scanned pages without OCR (optical character recognition, which turns page images into text) are invisible or garbled to the model. Run OCR first and spot-check a few handwritten or faint pages; those are where text extraction fails.
- Match pagination. If the bundle's printed page numbers don't match the PDF's page count (because of a cover sheet or an index), tell the model which to use, or stamp bundle page numbers onto every page before uploading.
- Split long bundles by section. Upload tabs or sections as separate files with clear names ("Tab C - Correspondence, pp. 301-587"). Gemini Notebook, for example, accepts up to 500,000 words or 200MB per source, but a model reading one huge file still does better with sections.
- Keep the index. Upload the bundle index as its own source. It lets you check coverage later.
- For transcripts, keep line numbers and speaker labels, and include timestamps if the transcript came from a recording.
Before running any summary, calibrate. Ask the tool "What is on bundle page C/342? Quote its first line", then open C/342 yourself. If the quote is from C/338, every reference the model gives you will be four pages out, usually because a four-page index at the front of the tab counts as PDF pages 1 to 4. That's a realistic failure, and it shows up late and expensively: a fee earner working through the chronology finds each reference slightly wrong, assumes they're typos, and corrects them one by one. A thirty-second test catches it before anything is generated. Repeat it for each section file, since each may have its own front matter.
OCR errors need the same suspicion on pages that matter. In an illustrative construction dispute bundle, a handwritten delivery note read "rejected 12 pallets, damp" and the OCR text layer said "received 12 pallets, damp". The model summarised it faithfully from the text: "12 pallets received on 6 June [B/214]". The reference was right and the summary was wrong, because the text it read was wrong. For handwritten, faxed or faint pages, the check is against the page image, never the extracted text, and it's worth listing those pages in advance so the checker knows which ones they are.
Ask for summaries that point back to the page
A summary without references can't be checked quickly, so every prompt should demand them. The chronology is the most useful single output for most bundles:
Using only the uploaded bundle sections, produce a chronology of
events relevant to [issue].
For each entry: date | event in one sentence | bundle page(s).
Rules:
- Every entry must have a page reference. If you can't find the
page, don't include the entry.
- Quote exact words in inverted commas where the wording matters.
- Where documents give conflicting dates, list both with pages.
- After the chronology, list any documents in the index that you
did not use, by page.
An illustrative extract of the output, from a contract dispute bundle:
14 Mar Supplier emails revised delivery schedule; states
"week 16 at the earliest" [C/342]
2 Apr Client's operations manager replies accepting
revised dates "subject to board sign-off" [C/351]
9 Apr Board minutes record approval of revised
schedule [D/612]
28 Apr Client emails complaining of late delivery [C/377]
Documents not used: C/355-359, C/360, D/598-604
What the check found. The 2 April entry is accurate, but it cites page 351 when the email is on 352; a small slip, and exactly why references get spot-checked. More important is the "not used" list: C/360 turned out to be a second email from the client on 3 April withdrawing the acceptance, which the chronology skipped because the model treated it as a duplicate of the 2 April thread. That single document changes the case theory. The "list what you didn't use" instruction is what surfaced it.
Other outputs worth prompting for, each with the same reference rule: a document-by-document index with a one-line description of each; an issues table mapping each pleaded issue to the pages that bear on it; and a list of apparent inconsistencies between witness statements, with both page references for each.
An illustrative issues table for the same contract dispute, as it came back and after checking:
| Issue | Pages given by the model | After checking |
|---|---|---|
| 1. Was the revised delivery schedule agreed? | C/342, C/351, D/612 | Add C/352 and C/360; the withdrawal email is the key page |
| 2. Was delivery late against the agreed schedule? | C/377, E/801-806 | Correct |
| 3. Did the client suffer loss from the delay? | E/820-844, C/390 | Remove C/390: it discusses a different contract's schedule and was matched on the word "delay" |
Keyword matching is the usual cause of wrong pages in an issues table. The model finds "schedule" or "delay" and maps the page to the issue without registering that the document is about something else. Reading the first paragraph of each mapped page is usually enough to spot it.
The witness-statement comparison has its own false positives. An illustrative entry: "Inconsistency: the operations manager says the meeting was on 04/05 [F/902, para 11]; the finance director says 5 April [F/931, para 7]." Both mean 5 April; one statement used month-first date format. Before any inconsistency goes into a note for counsel or the client, read both passages and ask whether the conflict is in the facts or only in the formatting.
For general long-document technique outside legal work, summarising long documents without missing details covers chunking and prompts in more depth.
Transcripts fail differently from documents
A hearing transcript produced by an official transcriber and an AI transcript of a recorded client interview are different beasts, and the second carries errors before any summary starts. The failures to look for:
- Speaker misattribution. The transcript assigns a line to the wrong person, and the summary then says a witness conceded something they didn't.
- Negation flips. "I didn't sign it" heard as "I did sign it". Short words are easily lost in poor audio.
- Numbers and names. "Fifteen" and "fifty", similar surnames, company names.
- Tidied hesitation. Summaries often turn "I think... maybe it was Tuesday?" into "the witness said it was Tuesday", removing the uncertainty that matters most.
Here's an illustrative before and after from a recorded interview. The AI summary said: "The client confirmed she signed the variation on 12 May." The transcript line, on checking, read "I don't think I signed it until after the twelfth, maybe the fifteenth?" and the recording made clear the answer was uncertain. The corrected summary line: "Client unsure of signing date; thinks after 12 May, possibly 15 May [transcript 14:22, lines 311-313]." When a transcript came from AI transcription, check any line you'll rely on against the audio, not just the text. AI versus human transcription covers when paying for human transcription is worth it.
A prompt that keeps uncertainty in:
Summarise this transcript by topic. For each point give the
speaker and line numbers. Preserve any hedging or uncertainty
in the speaker's words ("I think", "maybe", "not sure"); do not
turn uncertain answers into firm statements. List any passages
marked inaudible that bear on [issue].
The checking routine before anyone relies on a summary
- Every point you'll use: open the referenced page and read it. Not the neighbouring sentence; the passage itself.
- Random spot-check: at least one in five other entries, chosen before you read them, so you don't only check the ones that look odd.
- Coverage: compare the "documents not used" list against the index. Anything unused but plausibly relevant gets read.
- Adverse material: ask separately, "List every document that is unhelpful to [our client's position], with pages." Summaries lean towards the narrative in your prompt; this question pushes back.
- Record the check: a line in the file saying what was checked, by whom and when. If a summary later turns out to be wrong, you want to know which parts were verified.
Size the spot-check before you start, so it doesn't shrink under time pressure. Say a chronology has 140 entries and the team will rely on 60 of them. All 60 are checked. One in five of the other 80 is 16 more; pick them by taking every fifth entry from a starting number chosen before anyone reads the list. At about two minutes each, 76 checks is around two and a half hours, which is roughly what the checking stage in the example below takes. If a spot-check turns up a wrong reference, check the next five entries around it as well, because errors tend to cluster where a section file was poorly scanned or paginated.
The adverse-material question in step 4 deserves its own read of the answer. Run on the contract dispute bundle, it might return, illustratively: "C/360: client's email of 3 April withdrawing acceptance of revised dates. D/598: internal client email of 20 March noting 'we're short on warehouse space until May anyway'. E/815: client's own delivery log showing goods rejected for reasons unrelated to timing." All three need reading. D/598 had been missed by the chronology entirely because it had no event in it, only an admission, and it bears on whether the delay caused any loss. E/815, on reading, was about a different product line, so it came off the list. The question works best when you treat its answer as a reading list, not a finding.
If the summary includes any reference to case law, even something the model "noticed" in a skeleton argument, put it through the small-firm routine for AI-invented case law before it goes anywhere.
A 1,100-page bundle in a small litigation practice
Consider an illustrative three-fee-earner commercial litigation practice preparing for a mediation on a supply contract dispute, with a 1,100-page bundle across six tabs.
- Preparation (about an hour): OCR on two scanned tabs, bundle page numbers stamped, six section files and the index uploaded to a matter-only notebook.
- Summaries (about 90 minutes): a chronology, a document index, an issues table and a witness-statement comparison, each run section by section and then combined.
- Checking (about three hours): every entry the team planned to rely on (around 60), a one-in-five spot-check of the rest, the coverage comparison and the adverse-documents question. Four wrong page references and one significant omission found and fixed.
- Compared with doing it by hand: a first read and a manual chronology would typically take a fee earner the best part of two days. The AI-assisted version took about five and a half hours, three of them checking.
The proportion is the lesson: more than half the time went on verification, and that's what made the output safe to use. A team that trims the checking to save time has just swapped a slow, reliable process for a fast, unreliable one.
A family practice working on a financial disclosure bundle needs a different kind of summary. Much of the bundle is bank statements, and the summary a fee earner wants is a list of large or unusual transactions. AI is good at finding candidates ("list every transaction over $2,000 and every transfer to an account not listed in the disclosure form, with page"), but it misreads figures in scanned statements more often than in typed documents. Every amount that makes it into correspondence gets checked against the page image, not the extracted text.
Where AI summaries should stop
- Privilege decisions. A model can flag documents that look privileged; deciding is a lawyer's call.
- Anything pasted straight into a court document. Summaries are working papers. Statements of fact in documents for the court should be drafted from the checked sources.
- Documents you're not permitted to upload. If there's any doubt about restrictions on disclosed material or the client's instructions, don't upload; summarise by hand or ask.
- Judgements about credibility. The model can list inconsistencies. Whether they matter, and what they say about a witness, is advocacy.
For firms setting up a shared workspace for matter teams, setting up Claude Projects as a shared team assistant and Gemini Notebook for small business teams cover the practical setup, including sharing controls worth locking down before the first bundle goes in.
Further reads
- Is Gemini Safe for Confidential Business Data in Workspace? — Workspace data controls in more depth before you upload.
- Can Lawyers Use AI Note Takers in Client Meetings? — Creating transcripts of client meetings in the first place.
- Which Legal Tasks Should a Small Firm Never Hand to AI? — Where summaries shade into judgement AI shouldn't make.
- How to Anonymise Client Data Before You Paste It Into AI — When to strip identifiers before a document goes in.
- ChatGPT or a Legal AI Tool: Which Should a Small Firm Use? — Whether a legal-specific tool is worth it for bundle work.
- How HR Consultants Use AI for Policies, Letters and Cases — Three workflows for HR consultants, with a client fact sheet, letter before-and-after, a grievance chronology prompt and the lines AI must never cross.
- AI Tools and AI Development: The Complete 2026 Guide — the AI hub, including every tutorial in the AI-for-business series.
Sources: Gemini Notebook Help (source limits; using Gemini Notebook with a work account); OpenAI and Anthropic business plan data-use pages.