Use AI to find the personal data, then remove it with a proper redaction tool. Run the document through a detection step (an AI assistant on a business plan, or an open-source detector such as Presidio), review every hit, apply true redaction in Acrobat Pro or by deleting the text at source, strip hidden data, and search the output to prove nothing's left.
The dangerous mistakes aren't about AI at all. Black highlighting in Word, or a black box drawn over a PDF, hides the text on screen but leaves it in the file, where anyone can copy it out. And pasting an unredacted client file into a personal chatbot account to "clean it up" is itself a disclosure of the very data you were trying to protect. Get those two right and AI becomes a fast, thorough second pair of eyes.
Decide what has to go before you open any tool
Redaction starts with the recipient, not the document. The same client file needs different redactions for an external compliance reviewer, an insurer, a new adviser taking over the case, or a member of the public who asked for their own data. Write a short redaction brief first. Here's an illustrative one from a financial planning firm sending three client files to an external compliance consultant for a routine file review:
| Data in the file | Action | Why |
|---|---|---|
| Client and partner names | Replace with [Client A], [Partner A] | Reviewer needs to follow who's who, not identities |
| Dates of birth | Replace with age at advice date | Age matters to the advice; the date doesn't |
| Home address, phone, email | Redact | Not needed for a suitability review |
| Employer name | Replace with sector, e.g. "public-sector employer" | Indirect identifier in a small town |
| Policy and account numbers | Redact all but last 3 digits | Lets the reviewer match documents |
| Health disclosures | Redact unless they drove the recommendation | Special category data; share only if essential |
| Income, assets, fund choices, fees | Keep | The substance of the review |
| Adviser and firm names | Keep | The reviewer is assessing the firm's advice |
Twenty minutes on this table saves hours of second-guessing later, and it becomes the instruction you give the AI. Keep it with the file as your record of what was removed and why.
Three ways to find personal data with AI
Detection is where AI helps most. Pick the route that suits your volume and sensitivity.
Route 1: an AI assistant on a business plan
Upload the document to ChatGPT Business, Claude Team, Gemini in Workspace or Microsoft 365 Copilot (plans that don't train on your content by default) and ask for a list of findings, not a rewritten document:
Read the attached document and list every item of personal data, as a table:
Page | Exact text as it appears | Type (name, contact detail, date of birth,
account/policy number, health, employer, other identifier) | Person it relates to
Include indirect identifiers: job titles, employers, family relationships,
unusual events, and anything that could identify someone when combined.
Include text inside headers, footers, signatures, tables and image captions.
Do NOT produce a redacted version of the document. Only the table.
If a page looks like an image you couldn't read, list its page number.
An illustrative extract of what comes back for a 12-page suitability report:
Page | Exact text | Type | Person
1 | Mr Daniel Sample | Name | Client A
1 | 14 Orchard Way | Address | Client A
2 | 03/07/1979 | Date of birth | Client A
2 | head of maths at the | Employer/job title | Client A
| local secondary school | |
4 | Policy no. PX4471902 | Policy number | Client A
7 | Sarah (wife) | Name, relationship | Partner A
9 | recent cancer diagnosis | Health | Partner A
Pages 11-12 appear to be scanned images; text not read.
What you'd fix: it found the job title as an indirect identifier (good), but it listed "Daniel" only once, when the report uses "Dan" in three places. Search for nicknames and first names on their own. And pages 11-12 need OCR before any tool can find anything on them.
For long files, upload 15-20 pages at a time rather than the whole pack. Assistants given 60 pages at once tend to be thorough on the first few and thin on the rest, and they rarely say so. Ask it to state the first and last page it read in each batch, and compare the number of findings per page across batches: a batch that suddenly returns far fewer items is worth re-running.
Route 2: an open-source detector you run yourself
Presidio, an open-source toolkit that began at Microsoft and is now maintained by a community organisation, detects personal data in text and images using named-entity recognition, pattern matching, rules and checksums, and can then redact, mask or replace it. It runs on your own machine, so nothing leaves the building. It needs someone comfortable installing Python tools, and its own documentation is blunt that automated detection gives no guarantee of finding everything. That's true of every route here.
Route 3: the pattern search in your redaction tool
Adobe Acrobat Pro's Find text and redact searches for a single word or phrase, a list of words you type or import from a text file, or built-in patterns such as phone numbers, email addresses, credit card numbers and dates. It isn't generative AI, but combined with the name list from route 1 it catches every repeat of a known identifier across hundreds of pages.
In practice the best result comes from combining routes: the AI assistant or Presidio builds the list, and the redaction tool finds every occurrence of each item on it.
Applying the redaction so it can't be undone
Detection gives you a list. Removal has to happen in the file itself, with tools that delete the underlying content.
PDFs in Acrobat Pro
- Make sure the PDF has a text layer. If pages are scans, run OCR first, or pattern and word searches silently find nothing on those pages.
- Use Find text and redact with your list (import it as a text file if it's long). Review each marked occurrence before accepting.
- Use Redact text and images for anything the search can't find: signatures, photos, stamps, handwritten notes.
- Apply the redactions. Adobe warns that applied and saved redactions are permanent, which is the point, so work on a copy named
… - REDACTED.pdf. - Run the sanitise option to remove hidden information: metadata, attachments, bookmarks, comments, form fields and scripts that can carry the original text.
Word documents
Replace text at source, which is often better than redaction because the document stays readable. Build a replacement map from your AI findings ("Daniel Sample" and "Dan" to [Client A]), use Find and Replace with Match case and Find whole words only, accept all tracked changes, delete comments, then run File, Info, Check for Issues, Inspect Document and remove everything it finds. Export to PDF and share the PDF, not the Word file.
Spreadsheets
Delete the columns rather than hiding them, check for hidden sheets, filter views, comments and named ranges, and remember that pivot tables can cache data from columns you deleted. The safest route is to copy only the columns you need into a new workbook and share that.
Screenshots and images
Cover with a solid, fully opaque shape and export a flattened image. A semi-transparent highlighter from a phone's markup tool can often be reversed by adjusting brightness and contrast.
A financial planner's 64-page file review pack
Walk through an illustrative run. A three-adviser financial planning firm sends an external compliance consultant three client files each quarter. One pack is 64 pages: fact-find, risk questionnaire, suitability report, provider illustrations, and a thread of client emails.
Before: the paraplanner redacted by hand in Acrobat, page by page, about three hours per pack, and still sent one pack last year with a client's mobile number in an email signature on page 51.
The new routine:
- Redaction brief (as above), reused each quarter: 5 minutes.
- OCR on the 9 scanned pages: 3 minutes.
- Findings table from the firm's Claude Team plan, in four uploads of about 16 pages: 15 minutes including reading the output.
- The findings produced a list of 37 distinct identifiers (names, nicknames, numbers, addresses, the employer), imported into Find text and redact: 187 occurrences marked. The paraplanner reviewed them at a glance and rejected 11 false positives, mostly fund names containing surnames: 20 minutes.
- Manual pass for what search can't find: two wet signatures, a handwritten note on the risk questionnaire, a photo of a passport page that shouldn't have been in the pack at all: 10 minutes.
- Apply, sanitise, verify (see the checklist below): 12 minutes.
Illustrative total: about 65 minutes against three hours. The verification step found two misses: a policy number split across a line break ("PX44" at the end of one line, "71902" at the start of the next) that the word search didn't match, and the client's surname in the PDF's document title property. Both fixed before sending. Neither would have been caught by the AI step alone.
What AI detection routinely misses
Every detection route struggles with the same things. Check these by hand on every file:
- Text inside images: signatures, logos with names, screenshots of emails, photographed documents embedded in a PDF.
- Split identifiers: numbers broken across lines or table cells, or written with spaces ("PX 447 1902").
- Nicknames and first names alone: "Dan", "the Samples", "Mrs S".
- Indirect identifiers: "the only female partner at the practice", a rare job, a specific event and date. These need human judgement about the recipient.
- Hidden layers: document properties (author, title), file names, comments, earlier versions in the Word file, attachments inside a PDF, email headers in a forwarded chain.
- Other people's data: a client's email mentioning their neighbour, their child's school, their ex-partner. When you're responding to a request from someone for their own data, it's often these third parties who need redacting.
Redact or replace with labels? It depends who reads it
Black boxes are right when the recipient must not know something exists. Labels such as [Client A] are right when the recipient needs to follow the story. Two illustrations:
A PR consultancy wants to show a prospective client how it handled a past media crisis, using its enquiry log. Journalists' names, phone numbers and outlets come out entirely; the client company is replaced with "[Retail client]". The prospect needs to see the response times and the approach, not who called. Here labels keep the log readable, and nothing about the individuals is needed.
The same consultancy is asked by a member of the public for a copy of the personal data it holds about them, which appears in that same crisis log. Now the requester's own entries stay in, and everyone else's details are blacked out, because replacing another person's name with a label can still let the requester work out who they are. Different recipient, different treatment, same file. With requests like this, check the rules and deadlines that apply with your data-protection adviser before responding.
One legal footnote worth knowing: if you keep the key that maps labels back to real people, data-protection law such as the GDPR usually still treats the labelled document as personal data. Labels lower the risk; they don't make a document public.
Email threads and meeting transcripts need their own pass
Two document types cause more slips than any other, because personal data turns up in places nobody reads closely.
Forwarded email chains carry every earlier message, each with its own signature block, phone number and sometimes a confidentiality footer naming the sender. They also carry quoted text from people who were never meant to be in the file. Before redacting, export the thread to PDF so every message is visible, and ask the AI for findings per message rather than per page. Here's an illustrative before and after from a client email in the financial planner's pack:
BEFORE
From: Daniel Sample <dan.sample@[provider].com>
Sent: 12 March, 08:14
Hi Priya, Sarah and I talked it over. Given her diagnosis we'd rather keep
6 months of spending in cash. Our son Tom
starts university in Sept so we'll need about $9,000 then.
Dan | 07xxx xxx xxx
AFTER (labels, per the redaction brief)
From: [Client A] <[redacted]>
Sent: 12 March, 08:14
Hi [Adviser], [Partner A] and I talked it over. Given [redacted - health]
we'd rather keep 6 months of spending in cash. Our [child] starts
university in September so we'll need about $9,000 then.
[Client A] | [redacted]
The reviewer keeps what matters to the advice (a cash reserve, a known expense and its timing) and loses everything that identifies the family. Note that the adviser's own name was labelled too, because this particular brief asked for it; the firm name stayed.
AI meeting transcripts are the other trap. They name everyone who spoke, capture asides ("sorry, that's my daughter's school calling"), and often mis-transcribe names so that a search for "Sample" misses "Sampel". Ask the AI to list every person mentioned, including misspellings and partial names, then search for each variant. If the transcript isn't essential, a summary written for the recipient is usually safer than a redacted transcript.
Proving nothing's left before you press send
Run this checklist on the redacted copy, not the working file. It takes about ten minutes and it's the step that catches what everything else missed.
- Search for every item on your identifier list in the redacted PDF, including partial strings (last four digits of each number, each surname). Zero results expected.
- Select all, copy, paste into a plain text editor and skim. Hidden text under boxes shows up here.
- Check document properties (title, author, subject, keywords) and the file name.
- Look for attachments, bookmarks and comments in the PDF's side panels.
- Open the file in a different viewer, such as a browser. Some fake redactions look solid in one app and vanish in another.
- Page count and order match the original, so nothing was accidentally dropped (or added, like that passport page).
- A second person spot-checks five random pages against the brief for anything sensitive that isn't in the list, especially indirect identifiers.
Record in a short log what was redacted, by whom, and who checked it. If the recipient ever queries a redaction, you can answer precisely.
Where this fits with pasting data into AI generally
This routine is for documents leaving the business. The related problem, stripping client details before you use a document with an AI tool at all, is covered in how to anonymise client data before you paste it into AI, and the broader controls for a team are in keeping customer data private when your team uses AI. If your scans are poor quality, AI extraction vs traditional OCR explains why some pages come back unreadable, and GDPR and AI tools for a small business sets out the wider duties. For anything contentious, such as documents for a dispute or a formal request from a regulator or court, get your solicitor or data-protection adviser to agree the redaction brief before you start.
Redaction questions that come up before a file goes out
Can I paste a document into ChatGPT and ask it to redact it?
Only on a business plan that doesn't train on your content, and only to get a list of what to redact. Don't send its rewritten text as the redacted document: models can paraphrase, drop sentences or alter figures without saying so. Apply the redactions yourself in the original file with a proper redaction tool, then verify.
Does deleting text in Word count as redaction?
Deleting or replacing the text in Word does remove it, but tracked changes, comments, earlier versions and document properties can still hold the original. Accept all changes, delete comments, run Inspect Document, then save a fresh copy and export it to PDF. Never rely on black highlighting or a black shape laid over the text.
How long should we keep the unredacted original?
Keep it for as long as you'd normally keep that record, stored as usual, with the redacted copy saved alongside and clearly named. Record what you removed and why in a short log. If anyone later questions the redactions, the original and the log show exactly what was withheld.
Is a pseudonymised document safe to share freely?
Not automatically. If you hold the key that maps labels back to real people, data-protection law such as the GDPR generally still treats the document as personal data. Pseudonymising reduces risk and keeps the document readable, but you still share it only with people who need it, under the same care.
Further reads
- Are ChatGPT, Claude, Gemini and Copilot GDPR-Compliant? — Which AI plans are suitable for documents containing personal data.
- What to Check in an AI Tool's Privacy Policy and Terms — What to read in a tool's terms before uploading client files.
- Reusing Past Client Work With AI Without Leaking Client Data — Reuse old client work as templates without leaking details.
- Are Chat-With-PDF Tools Safe for Contracts and Client Files? — The risks of chat-with-PDF tools for confidential files.
- How to Set Up Company AI Accounts Instead of Personal Logins — Keep sensitive documents out of personal AI accounts.
- AI Acceptable Use Policy for a Small Professional Firm — Put your redaction rules into a firm-wide AI policy.
- How to Automate Vaccination Checks for a Grooming Salon — Three layers of vaccine-check automation: software that watches expiry dates, AI-assisted certificate reading, and a full pipeline with a privacy catch.
- Turn Fixed Tickets Into a Searchable AI Knowledge Base — Turn a year of closed IT tickets into tested, searchable fix articles: selection rules, redaction, root-cause clustering, a template and upkeep.
- How Property Managers Use AI to Screen Tenant Applications Fairly — Written criteria, one summary format for every applicant, human decisions with recorded reasons, and a monthly check that your rules aren't quietly unfair.
- Patient Data and AI: A Confidentiality Checklist for Small Practices — A 30-item checklist, with a filled-in example register, for keeping patient information confidential when a small practice uses AI tools.
- How to Create Customer Personas With AI From Real Data — Build personas from booking, order and review data instead of AI guesswork: prompts, a persona template with evidence, and four businesses' results.
- How to Create SOPs From Screen Recordings With AI — Record a task once, talk through the why, and let AI draft the procedure. Covers Gemini video uploads, Loom, Scribe and Tango, redaction and the cold-run test.
- How to Run Staff Surveys and Exit Interviews With AI Analysis — A 12-question survey, an exit interview script, anonymity settings, a theme-coding prompt with checks, and an online clothing shop's first round of results.
- How to Automate Holiday Requests and Absence Tracking — Build a leave request process that checks balances, protects shift cover and records changes without repeated messages.
- How to Turn Handwritten Forms Into Spreadsheet Data With AI — Convert paper forms into checked spreadsheet rows while preserving unclear handwriting, original references and leading zeros.
- How to Prepare a Business Loan Application With AI Help — Build a traceable loan evidence pack, test a weak trading month and use AI to draft explanations without inventing financial claims.
- AI Tools and AI Development: The Complete 2026 Guide — the AI hub, including every tutorial in the AI-for-business series.
Sources: Adobe Acrobat Pro help pages on redaction, Find text and redact, and sanitising PDFs; Presidio repository documentation; Microsoft Office Inspect Document feature. Checked September 2026.