How to Anonymise Client Data Before You Paste It Into AI

Coding Liquids tutorial cover featuring Sagnik Bhattacharya for How to Anonymise Client Data Before You Paste It Into AI.
Coding Liquids tutorial cover featuring Sagnik Bhattacharya for How to Anonymise Client Data Before You Paste It Into AI.

Copy only the part of the document the task needs, replace every name, address, account number and email with a consistent placeholder such as CLIENT_A, blur the details that could still identify someone (exact dates, unusual job titles, precise amounts), strip hidden metadata, and keep the key that maps placeholders back to real names outside the AI tool.

Be clear about what that gets you. Swapping names for placeholders while you keep the key is pseudonymisation, not anonymisation. Data-protection law such as the GDPR still treats it as personal data, because you can reverse it. That's fine for most tasks, but it means the tool you paste into should still be one your firm has approved for client work.

Follow me on Instagram@sagnikteaches

Pseudonymised or anonymised: know which one you're producing

The GDPR's Recital 26 draws the line. Data is anonymous only when nobody can identify the person using "all the means reasonably likely to be used", taking into account cost, time and available technology. Pseudonymised data, where the link can be restored with extra information such as your key, is still personal data.

Connect on LinkedInSagnik Bhattacharya

For a professional firm this has a practical consequence. Real anonymisation of a single client's file is hard, because client work is specific by nature. There may be exactly one catering company on your books that lost its head chef in the spring, turns over roughly what it does and has a dispute with a venue. Strip every name and a colleague could still tell you who it is. So aim for two things, and be honest about which you've achieved:

Subscribe on YouTube@codingliquids
  • Pseudonymised for the tool: the AI vendor can't tell who the client is, but you can. Good enough for approved business tools and most drafting tasks.
  • Anonymised for sharing: nobody, including someone who knows your client base, could work out who it is. Needed before material goes into an unapproved tool, a training example, a case study or a public forum.

Three layers of identifying detail in client files

Most people deal with the first layer and stop. The leaks come from the second and third.

LayerWhat it includesExamples in professional workWhat to do
Direct identifiersAnything that names or contacts a person or businessNames, trading names, addresses, emails, phone numbers, account and policy numbers, tax or registration references, bank details, signaturesReplace with placeholders from a key
Indirect identifiersDetails that single someone out in combinationExact dates, precise turnover, job titles held by one person, rare events (a fire, a lawsuit, a merger), ages, family details, named suppliers or landlordsGeneralise: ranges, months, roles, "a supplier"
Hidden dataInformation inside the file rather than the visible textAuthor and company fields, tracked changes, comments, hidden rows and sheets, file names, email headers, image metadataDon't upload the original; copy text into a clean document, or inspect and sanitise

The six-step routine, with how long each step takes

For a page or two of text, the whole routine takes five to ten minutes once you've done it a few times. Longer documents are where the tools in the next section earn their place.

  1. Ask whether the AI needs the client at all (one minute). Many questions work without the specifics. "How should a company treat a deposit refunded after the year end?" gets you the same principle as pasting the client's ledger. Only move on if the task truly needs their material.
  2. Work on a copy, in plain text (one minute). Select the passage you need and paste it into a blank document with "keep text only". This leaves behind tracked changes, comments, author fields and embedded objects. Never upload the original file if copying the text will do.
  3. Build a key and replace direct identifiers (two to four minutes). Make a short table: placeholder on the left, real value on the right. Keep it in the client's file on your own system, never in the chat. Use Find and Replace with "Match case" ticked, and replace the longest strings first (the full registered name before the short trading name) so fragments don't survive. Use neutral tokens: CLIENT_A, DIRECTOR_1, SUPPLIER_2, ACCOUNT_1. A filled-in key for one client looks like this (the right-hand column is described here rather than shown, but in your file it holds the real values):
    KEY: stored in client folder only, never pasted into AI    Fee earner: [initials]
    Placeholder     Real value                               Also catch
    CLIENT_A        registered company name                  trading name, web domain
    DIRECTOR_1      finance director's full name             first name alone, email address
    DIRECTOR_2      operations director's full name          nickname used in emails
    SUPPLIER_1      staffing agency's name                   invoice prefix "AGY-"
    ACCOUNT_1       business bank account number             last four digits in notes
    The "also catch" column is the part people skip. Real names turn up in email addresses, web domains and invoice prefixes long after the obvious instances have gone. The opposite slip is just as easy: a director whose first name is three letters that also begin "annual" and "announce" can leave you with "DIRECTOR_1ual accounts" all through the note if you replace without ticking "Find whole words only". Tick it for short names, then search the result for your placeholder followed by a letter to catch any that slipped through.
  4. Generalise the indirect identifiers (two to four minutes). Turn exact dates into months or quarters, unless the date is the point. Round amounts, or scale every figure by the same factor so ratios still work (not when the task involves thresholds or tax bands, which scaling breaks). Replace one-of-a-kind job titles with roles ("the finance lead"). Describe rare events in general terms. A typical before and after:
    Before: On 14 March the company's founder, now 71, handed day-to-day control to
            his daughter after the flood at the second site; turnover fell from
            $2,340,118 to $1,902,455.
    After:  In Q1 the founder handed day-to-day control to a family member after a
            major incident at one site; turnover fell by roughly 19%.
    The age, the relationship, the exact event and the exact figures each narrowed the field. Together they named the client.
  5. Strip hidden data if a file must go in (two minutes). In Word, save a copy, then use the Document Inspector (File, Info, Check for Issues, Inspect Document) and remove what it finds. Microsoft advises running it on a copy because removals can't always be undone. For PDFs, Acrobat Pro's redaction tool removes the text itself and offers to sanitise hidden information as you apply it; a black box drawn over text in an ordinary editor hides nothing. Rename the file, since file names often carry the client's name.
  6. Run the stranger test (one to two minutes). Read the result as someone who knows your sector and your client list. Could they name the client? If the answer is "probably", generalise further or don't use the tool. For anything leaving the firm, ask a colleague who doesn't work on that client to try.

Once the text is already pseudonymised, an approved AI tool can help with the stranger test, because you're no longer exposing anything by pasting it. Something like this does the job:

Below is text I have tried to anonymise. Do not rewrite it. List every word or
detail that could still help someone identify a specific business or person,
including combinations of details. For each, say why and suggest a vaguer version.

[paste the pseudonymised text]

An illustrative reply for a pseudonymised note about a sports equipment shop:

1. "the only climbing-wall installer for 50 miles"  - unique description;
   suggest "a specialist installer".
2. "opened its second branch in May 2024"  - date plus event narrows it;
   suggest "expanded to a second site recently".
3. "DIRECTOR_1, a former professional cyclist"  - distinctive background;
   suggest removing it unless relevant.
4. "SUPPLIER_2 (the frame brand run by a former Olympian)"  - founder plus product
   may identify the supplier; suggest "a frame supplier".

Reading the list critically: the flags are right and the second and third suggestions can go in as they are. For the fourth, if the supplier plays no part in the task, delete the bracket entirely rather than rewording it. The first suggestion loses the detail that made the note useful, so "a specialist installer serving a wide area" is the better edit. What the AI can't know is what your colleagues know. It won't spot that "the client with the late-night loading bay complaint" is famous in your office, so the human read-through still comes last.

Worked through: a bookkeeper's question about a catering client

Here's an illustration. A three-person bookkeeping practice wants help drafting an email to a catering company client about staff costs that jumped in one quarter. The original notes name the company, its two directors, the head chef who left, the agency that supplied temporary cooks and the exact figures.

After the routine, the text that goes into the firm's approved AI tool reads:

Client: CLIENT_A, a catering company (events and corporate lunches), about 25 staff.
Contacts: DIRECTOR_1 (finance), DIRECTOR_2 (operations).

Situation: In Q2, wages fell because a senior kitchen role was vacant for about
ten weeks, but agency costs from SUPPLIER_1 rose sharply. Net staff cost for the
quarter was roughly 18% higher than Q1. Agency invoices are coded to
"subcontractors" rather than "wages", so DIRECTOR_1's own dashboard shows
wages down and hasn't flagged the overall rise.

Task: Draft a short, friendly email to DIRECTOR_1 explaining the rise, why the
dashboard hides it, and suggesting we recode agency staff under staff costs.
Plain English, no jargon, under 200 words.

What changed: the names became tokens; "left in April" became "vacant for about ten weeks"; the exact figures became a percentage, which is all the email needs; the agency's name became SUPPLIER_1. The key (four lines) stays in the client's folder. When the draft comes back, the bookkeeper swaps the tokens back in their own email client, not in the chat.

Notice what wasn't removed: the sector and the staff count. The task needs them, and on their own they identify nobody. Anonymising isn't deleting everything; it's removing what the task doesn't need and blurring what it does.

Speeding it up without handing the job to the chatbot

For regular work, a few tools cut the time down:

  • A saved key per client. Keep a standing placeholder table in each client's folder (CLIENT_A is always the same client for that fee earner) so you don't rebuild it every time.
  • Find and Replace in bulk. Word and Google Docs both handle a list of replacements quickly. For spreadsheets, replace names with a client code column before copying rows out, or use a lookup against your key.
  • Automated detection, for firms with technical help. Presidio, an open-source toolkit that Microsoft started and that now sits with the Data Privacy Stack project under an MIT licence, detects and replaces personal data in text, images and structured data. Its own documentation warns that automated detection won't find everything, so treat it as a first pass that a person checks.
  • Exports that already use codes. Many practice management and accounting systems can export reports with client codes rather than names. Start from those.

Spreadsheets need one extra check, because what you see isn't all that gets copied. Take an illustrative aged-debtors report of 60 rows that a bookkeeper wants help analysing. She replaces the customer-name column with codes from the key, hides the contact-email column rather than deleting it, and uploads the workbook. The hidden column goes with it, and so does a second sheet called "Contacts" that the practice's template always includes. Copying the range into a new workbook doesn't fully save her either: manually hidden rows and columns are copied along with the visible cells unless she uses Go To Special and picks "Visible cells only" first. The safer order is to delete the columns the task doesn't need, copy only the cells that remain into a blank workbook, and check the sheet tabs along the bottom before anything is uploaded. For analysing debt ageing, the AI needs the code, the amount and the days overdue; three columns, not twelve.

What about asking the AI to anonymise the text for you? That means pasting the identified version first, which defeats the point unless the tool is already approved for that client material. Inside an approved business account it's a reasonable way to prepare something for sharing further, as long as you check the result line by line. For whole documents, such as a contract going to a third party, see how to redact personal data from documents before sharing.

Putting the real names back safely

The return journey has its own traps:

  • Swap tokens back outside the AI tool. Do it in the email or document you'll send. Pasting the key into the chat undoes the whole exercise.
  • Check for invented detail. AI sometimes fills gaps with plausible specifics, such as a made-up first name for DIRECTOR_1 or a month you didn't mention. Anything not in your key or your notes shouldn't be in the final version.
  • Check tokens survived intact. "CLIENT_A's" or "Client A" can slip past a Find and Replace. Search for every variant before you send.
  • Re-read with the real names in. A sentence that read fine about SUPPLIER_1 may be tactless once it names a supplier the client likes.

Here's the slip that catches careful people, in an illustrative case: an office manager at a subscription box company's accountants anonymised a payroll query carefully, replaced every name in the body and pasted it in. The email signature block at the bottom, copied along with the text, still carried the client's finance manager's name, direct line and company logo alt-text. Nobody noticed until a colleague reviewing the chat history saw it. The fix is mechanical: delete everything below the last line of the message before you start replacing, and search the finished text for "@", "www" and the client's legal suffix as a final sweep.

When no amount of editing makes it safe to paste

Some material shouldn't go into a general AI tool even after the routine. Stop if any of these apply:

  • Health, criminal records, sexuality, religion or similar sensitive categories where the detail is the point of the task. You can't blur what you need, and the harm from a leak is highest.
  • Very small groups. "The only employee on long-term leave at a four-person firm" identifies someone whatever you call them.
  • Privileged or contractually restricted material. If a confidentiality agreement or a client instruction says the material stays within named systems, anonymising doesn't change that.
  • Audio and video. Voices and faces identify people, and transcripts carry names spoken aloud. Transcribe inside an approved tool and anonymise the transcript instead.
  • Combinations you can't break. If the question only makes sense with the rare event, the exact date and the precise figure, the client is identifiable. Use an approved business tool, or answer it without AI.

If your team does this often, write it into your firm's rules so everyone follows the same routine. There's a sample clause in the AI acceptable use policy for a small professional firm, and the tiering behind it is explained in how to classify business data before using AI tools. For the wider safety question, including what business plans change, read whether it's safe to put customer data into ChatGPT.

Follow-up questions about anonymising client material

Is replacing names with initials enough?

Rarely. Initials combined with a sector, a date and an amount often point straight back to one client, especially in a small firm where staff know the client list. Use neutral placeholders that carry no part of the real name, such as CLIENT_A or SUPPLIER_2, and deal with the indirect details as well. Initials also collide: two clients with the same initials make the AI's answer ambiguous when you map it back.

Do we still need to anonymise if we use ChatGPT Business or Claude Team?

Less often. Those plans don't train on your content by default and come with business data terms, so many firms approve them for ordinary client material with a minimum-necessary rule. Anonymising still matters for restricted material, for anything you'll share outside the firm, for clients who've restricted AI use, and whenever you're unsure which account you're signed into.

What about screenshots, scans and photos?

Images carry identifying detail that text replacement can't reach: letterheads, signatures, faces, number plates and the metadata stored in the file. Retype or copy the text you need into a clean document instead of uploading the image. If an image must go in, crop it to the relevant area, cover identifiers with a solid fill, then export a fresh copy so the original layers and metadata don't travel with it.

Further reads

Sources: GDPR Recital 26 (pseudonymised versus anonymous data); Microsoft Support, Document Inspector; Adobe Acrobat Help, redaction and sanitisation; Presidio project repository and documentation.

Want a safe way for your team to use client data?

On a 1:1 call we'll look at the client material your team most wants to use with AI, decide which tasks need anonymising and which need a business account instead, and set up a routine people will stick to.

Book a 1:1 call with me