How to Redact Personal Data From Documents With AI Before Sharing

Coding Liquids tutorial cover featuring Sagnik Bhattacharya for How to Redact Personal Data From Documents With AI Before Sharing.
Coding Liquids tutorial cover featuring Sagnik Bhattacharya for How to Redact Personal Data From Documents With AI Before Sharing.

Use AI to find the personal data, then remove it with a proper redaction tool. Run the document through a detection step (an AI assistant on a business plan, or an open-source detector such as Presidio), review every hit, apply true redaction in Acrobat Pro or by deleting the text at source, strip hidden data, and search the output to prove nothing's left.

The dangerous mistakes aren't about AI at all. Black highlighting in Word, or a black box drawn over a PDF, hides the text on screen but leaves it in the file, where anyone can copy it out. And pasting an unredacted client file into a personal chatbot account to "clean it up" is itself a disclosure of the very data you were trying to protect. Get those two right and AI becomes a fast, thorough second pair of eyes.

Follow me on Instagram@sagnikteaches

Decide what has to go before you open any tool

Redaction starts with the recipient, not the document. The same client file needs different redactions for an external compliance reviewer, an insurer, a new adviser taking over the case, or a member of the public who asked for their own data. Write a short redaction brief first. Here's an illustrative one from a financial planning firm sending three client files to an external compliance consultant for a routine file review:

Connect on LinkedInSagnik Bhattacharya
Data in the fileActionWhy
Client and partner namesReplace with [Client A], [Partner A]Reviewer needs to follow who's who, not identities
Dates of birthReplace with age at advice dateAge matters to the advice; the date doesn't
Home address, phone, emailRedactNot needed for a suitability review
Employer nameReplace with sector, e.g. "public-sector employer"Indirect identifier in a small town
Policy and account numbersRedact all but last 3 digitsLets the reviewer match documents
Health disclosuresRedact unless they drove the recommendationSpecial category data; share only if essential
Income, assets, fund choices, feesKeepThe substance of the review
Adviser and firm namesKeepThe reviewer is assessing the firm's advice

Twenty minutes on this table saves hours of second-guessing later, and it becomes the instruction you give the AI. Keep it with the file as your record of what was removed and why.

Subscribe on YouTube@codingliquids

Three ways to find personal data with AI

Detection is where AI helps most. Pick the route that suits your volume and sensitivity.

Route 1: an AI assistant on a business plan

Upload the document to ChatGPT Business, Claude Team, Gemini in Workspace or Microsoft 365 Copilot (plans that don't train on your content by default) and ask for a list of findings, not a rewritten document:

Read the attached document and list every item of personal data, as a table:
Page | Exact text as it appears | Type (name, contact detail, date of birth,
account/policy number, health, employer, other identifier) | Person it relates to

Include indirect identifiers: job titles, employers, family relationships,
unusual events, and anything that could identify someone when combined.
Include text inside headers, footers, signatures, tables and image captions.
Do NOT produce a redacted version of the document. Only the table.
If a page looks like an image you couldn't read, list its page number.

An illustrative extract of what comes back for a 12-page suitability report:

Page | Exact text              | Type               | Person
1    | Mr Daniel Sample        | Name               | Client A
1    | 14 Orchard Way          | Address            | Client A
2    | 03/07/1979              | Date of birth      | Client A
2    | head of maths at the    | Employer/job title | Client A
     | local secondary school  |                    |
4    | Policy no. PX4471902    | Policy number      | Client A
7    | Sarah (wife)            | Name, relationship | Partner A
9    | recent cancer diagnosis | Health             | Partner A
Pages 11-12 appear to be scanned images; text not read.

What you'd fix: it found the job title as an indirect identifier (good), but it listed "Daniel" only once, when the report uses "Dan" in three places. Search for nicknames and first names on their own. And pages 11-12 need OCR before any tool can find anything on them.

For long files, upload 15-20 pages at a time rather than the whole pack. Assistants given 60 pages at once tend to be thorough on the first few and thin on the rest, and they rarely say so. Ask it to state the first and last page it read in each batch, and compare the number of findings per page across batches: a batch that suddenly returns far fewer items is worth re-running.

Route 2: an open-source detector you run yourself

Presidio, an open-source toolkit that began at Microsoft and is now maintained by a community organisation, detects personal data in text and images using named-entity recognition, pattern matching, rules and checksums, and can then redact, mask or replace it. It runs on your own machine, so nothing leaves the building. It needs someone comfortable installing Python tools, and its own documentation is blunt that automated detection gives no guarantee of finding everything. That's true of every route here.

Route 3: the pattern search in your redaction tool

Adobe Acrobat Pro's Find text and redact searches for a single word or phrase, a list of words you type or import from a text file, or built-in patterns such as phone numbers, email addresses, credit card numbers and dates. It isn't generative AI, but combined with the name list from route 1 it catches every repeat of a known identifier across hundreds of pages.

In practice the best result comes from combining routes: the AI assistant or Presidio builds the list, and the redaction tool finds every occurrence of each item on it.

Applying the redaction so it can't be undone

Detection gives you a list. Removal has to happen in the file itself, with tools that delete the underlying content.

PDFs in Acrobat Pro

  1. Make sure the PDF has a text layer. If pages are scans, run OCR first, or pattern and word searches silently find nothing on those pages.
  2. Use Find text and redact with your list (import it as a text file if it's long). Review each marked occurrence before accepting.
  3. Use Redact text and images for anything the search can't find: signatures, photos, stamps, handwritten notes.
  4. Apply the redactions. Adobe warns that applied and saved redactions are permanent, which is the point, so work on a copy named … - REDACTED.pdf.
  5. Run the sanitise option to remove hidden information: metadata, attachments, bookmarks, comments, form fields and scripts that can carry the original text.

Word documents

Replace text at source, which is often better than redaction because the document stays readable. Build a replacement map from your AI findings ("Daniel Sample" and "Dan" to [Client A]), use Find and Replace with Match case and Find whole words only, accept all tracked changes, delete comments, then run File, Info, Check for Issues, Inspect Document and remove everything it finds. Export to PDF and share the PDF, not the Word file.

Spreadsheets

Delete the columns rather than hiding them, check for hidden sheets, filter views, comments and named ranges, and remember that pivot tables can cache data from columns you deleted. The safest route is to copy only the columns you need into a new workbook and share that.

Screenshots and images

Cover with a solid, fully opaque shape and export a flattened image. A semi-transparent highlighter from a phone's markup tool can often be reversed by adjusting brightness and contrast.

A financial planner's 64-page file review pack

Walk through an illustrative run. A three-adviser financial planning firm sends an external compliance consultant three client files each quarter. One pack is 64 pages: fact-find, risk questionnaire, suitability report, provider illustrations, and a thread of client emails.

Before: the paraplanner redacted by hand in Acrobat, page by page, about three hours per pack, and still sent one pack last year with a client's mobile number in an email signature on page 51.

The new routine:

  1. Redaction brief (as above), reused each quarter: 5 minutes.
  2. OCR on the 9 scanned pages: 3 minutes.
  3. Findings table from the firm's Claude Team plan, in four uploads of about 16 pages: 15 minutes including reading the output.
  4. The findings produced a list of 37 distinct identifiers (names, nicknames, numbers, addresses, the employer), imported into Find text and redact: 187 occurrences marked. The paraplanner reviewed them at a glance and rejected 11 false positives, mostly fund names containing surnames: 20 minutes.
  5. Manual pass for what search can't find: two wet signatures, a handwritten note on the risk questionnaire, a photo of a passport page that shouldn't have been in the pack at all: 10 minutes.
  6. Apply, sanitise, verify (see the checklist below): 12 minutes.

Illustrative total: about 65 minutes against three hours. The verification step found two misses: a policy number split across a line break ("PX44" at the end of one line, "71902" at the start of the next) that the word search didn't match, and the client's surname in the PDF's document title property. Both fixed before sending. Neither would have been caught by the AI step alone.

What AI detection routinely misses

Every detection route struggles with the same things. Check these by hand on every file:

  • Text inside images: signatures, logos with names, screenshots of emails, photographed documents embedded in a PDF.
  • Split identifiers: numbers broken across lines or table cells, or written with spaces ("PX 447 1902").
  • Nicknames and first names alone: "Dan", "the Samples", "Mrs S".
  • Indirect identifiers: "the only female partner at the practice", a rare job, a specific event and date. These need human judgement about the recipient.
  • Hidden layers: document properties (author, title), file names, comments, earlier versions in the Word file, attachments inside a PDF, email headers in a forwarded chain.
  • Other people's data: a client's email mentioning their neighbour, their child's school, their ex-partner. When you're responding to a request from someone for their own data, it's often these third parties who need redacting.

Redact or replace with labels? It depends who reads it

Black boxes are right when the recipient must not know something exists. Labels such as [Client A] are right when the recipient needs to follow the story. Two illustrations:

A PR consultancy wants to show a prospective client how it handled a past media crisis, using its enquiry log. Journalists' names, phone numbers and outlets come out entirely; the client company is replaced with "[Retail client]". The prospect needs to see the response times and the approach, not who called. Here labels keep the log readable, and nothing about the individuals is needed.

The same consultancy is asked by a member of the public for a copy of the personal data it holds about them, which appears in that same crisis log. Now the requester's own entries stay in, and everyone else's details are blacked out, because replacing another person's name with a label can still let the requester work out who they are. Different recipient, different treatment, same file. With requests like this, check the rules and deadlines that apply with your data-protection adviser before responding.

One legal footnote worth knowing: if you keep the key that maps labels back to real people, data-protection law such as the GDPR usually still treats the labelled document as personal data. Labels lower the risk; they don't make a document public.

Email threads and meeting transcripts need their own pass

Two document types cause more slips than any other, because personal data turns up in places nobody reads closely.

Forwarded email chains carry every earlier message, each with its own signature block, phone number and sometimes a confidentiality footer naming the sender. They also carry quoted text from people who were never meant to be in the file. Before redacting, export the thread to PDF so every message is visible, and ask the AI for findings per message rather than per page. Here's an illustrative before and after from a client email in the financial planner's pack:

BEFORE
From: Daniel Sample <dan.sample@[provider].com>
Sent: 12 March, 08:14
Hi Priya, Sarah and I talked it over. Given her diagnosis we'd rather keep
6 months of spending in cash. Our son Tom
starts university in Sept so we'll need about $9,000 then.
Dan | 07xxx xxx xxx

AFTER (labels, per the redaction brief)
From: [Client A] <[redacted]>
Sent: 12 March, 08:14
Hi [Adviser], [Partner A] and I talked it over. Given [redacted - health]
we'd rather keep 6 months of spending in cash. Our [child] starts
university in September so we'll need about $9,000 then.
[Client A] | [redacted]

The reviewer keeps what matters to the advice (a cash reserve, a known expense and its timing) and loses everything that identifies the family. Note that the adviser's own name was labelled too, because this particular brief asked for it; the firm name stayed.

AI meeting transcripts are the other trap. They name everyone who spoke, capture asides ("sorry, that's my daughter's school calling"), and often mis-transcribe names so that a search for "Sample" misses "Sampel". Ask the AI to list every person mentioned, including misspellings and partial names, then search for each variant. If the transcript isn't essential, a summary written for the recipient is usually safer than a redacted transcript.

Proving nothing's left before you press send

Run this checklist on the redacted copy, not the working file. It takes about ten minutes and it's the step that catches what everything else missed.

  1. Search for every item on your identifier list in the redacted PDF, including partial strings (last four digits of each number, each surname). Zero results expected.
  2. Select all, copy, paste into a plain text editor and skim. Hidden text under boxes shows up here.
  3. Check document properties (title, author, subject, keywords) and the file name.
  4. Look for attachments, bookmarks and comments in the PDF's side panels.
  5. Open the file in a different viewer, such as a browser. Some fake redactions look solid in one app and vanish in another.
  6. Page count and order match the original, so nothing was accidentally dropped (or added, like that passport page).
  7. A second person spot-checks five random pages against the brief for anything sensitive that isn't in the list, especially indirect identifiers.

Record in a short log what was redacted, by whom, and who checked it. If the recipient ever queries a redaction, you can answer precisely.

Where this fits with pasting data into AI generally

This routine is for documents leaving the business. The related problem, stripping client details before you use a document with an AI tool at all, is covered in how to anonymise client data before you paste it into AI, and the broader controls for a team are in keeping customer data private when your team uses AI. If your scans are poor quality, AI extraction vs traditional OCR explains why some pages come back unreadable, and GDPR and AI tools for a small business sets out the wider duties. For anything contentious, such as documents for a dispute or a formal request from a regulator or court, get your solicitor or data-protection adviser to agree the redaction brief before you start.

Redaction questions that come up before a file goes out

Can I paste a document into ChatGPT and ask it to redact it?

Only on a business plan that doesn't train on your content, and only to get a list of what to redact. Don't send its rewritten text as the redacted document: models can paraphrase, drop sentences or alter figures without saying so. Apply the redactions yourself in the original file with a proper redaction tool, then verify.

Does deleting text in Word count as redaction?

Deleting or replacing the text in Word does remove it, but tracked changes, comments, earlier versions and document properties can still hold the original. Accept all changes, delete comments, run Inspect Document, then save a fresh copy and export it to PDF. Never rely on black highlighting or a black shape laid over the text.

How long should we keep the unredacted original?

Keep it for as long as you'd normally keep that record, stored as usual, with the redacted copy saved alongside and clearly named. Record what you removed and why in a short log. If anyone later questions the redactions, the original and the log show exactly what was withheld.

Is a pseudonymised document safe to share freely?

Not automatically. If you hold the key that maps labels back to real people, data-protection law such as the GDPR generally still treats the document as personal data. Pseudonymising reduces risk and keeps the document readable, but you still share it only with people who need it, under the same care.

Further reads

Sources: Adobe Acrobat Pro help pages on redaction, Find text and redact, and sanitising PDFs; Presidio repository documentation; Microsoft Office Inspect Document feature. Checked September 2026.

Sharing client files and worried what's left in them?

On a 1:1 call we'll look at the documents you send outside the business, set up a detect-and-redact routine with the tools you have, and agree the check that happens before anything leaves.

Book a 1:1 call with me