Reusing Past Client Work With AI Without Leaking Client Data

Coding Liquids tutorial cover featuring Sagnik Bhattacharya for Reusing Past Client Work With AI Without Leaking Client Data.
Coding Liquids tutorial cover featuring Sagnik Bhattacharya for Reusing Past Client Work With AI Without Leaking Client Data.

To reuse past client work without leaking client data, build a separate, sanitised library rather than pointing AI at your old client folders. Check each contract allows reuse, strip names, figures and identifying details, keep the reusable parts (structures, methods, anonymised excerpts), and store the library on a business plan that doesn't train on your content, open only to staff who need it.

Pointing AI at the shared drive is how the leak usually happens, and the vendor is rarely the cause. Microsoft 365 Copilot, for instance, can surface any file a user already has permission to open, so a folder shared with everyone years ago becomes searchable in seconds. And a model asked to "draft a proposal like the one we did for the housing trust" will blend that client's specifics into a new document for someone else. The risk is recombination, not theft.

Follow me on Instagram@sagnikteaches

Three ways past work leaks through AI

  • Recombination in drafts. The model borrows a figure, a quote or a distinctive phrase from one client's report and places it in another's. Nobody notices until the second client does.
  • Permission sprawl. AI tools connected to your whole drive, mailbox or document store inherit every sharing mistake you've ever made. Old "anyone with the link" shares and all-staff folders are the usual culprits. Cleaning up SharePoint permissions before Copilot covers the Microsoft side.
  • Plan and settings. A consultant pastes a client report into a personal chat account with model training switched on. Business plans (ChatGPT Business, Claude Team, Microsoft 365 Copilot, Gemini in Workspace) don't train on your content by default; personal accounts need the training switch turned off, and still aren't the place for client material.

Check you're allowed to reuse it

Before sanitising anything, look at the contract. Consultancy agreements usually separate the deliverables, which often belong to the client once paid for, from your pre-existing know-how, methods and templates, which stay yours. Some go further: an NDA may forbid using anything learnt on the project, and public-sector or funder contracts sometimes restrict reuse of commissioned material.

Connect on LinkedInSagnik Bhattacharya
Usually yours to reuseUsually the client'sCheck the contract
Your methods, frameworks, interview guides, report structuresThe final report and its findingsAnonymised excerpts from deliverables
Templates you brought to the projectClient data, survey responses, financialsCase-study use, even anonymised
General lessons ("small charities underestimate volunteer training time")Anything marked confidentialWork under an NDA or restricted funding terms

Two illustrative clauses show how different the answer can be. The first is common in consultancy terms: "The Consultant retains ownership of all methodologies, templates and know-how developed before or independently of this engagement, and may use general knowledge and experience gained during it." That leaves your frameworks in tier 1 and lets you keep anonymised lessons. The second turns up in NDAs: "All information disclosed or generated in connection with the Project is Confidential Information and shall not be used for any purpose other than the Project." Read literally, that covers even a sanitised finding, so the whole engagement goes to tier 3 unless the client agrees otherwise in writing. When wording sits between those two, ask your solicitor rather than deciding it in a spreadsheet.

Subscribe on YouTube@codingliquids

If your contracts don't say, fix that for future work. Mentioning AI use in client contracts and proposals includes wording for reuse and AI tools.

Sort every document into one of three tiers

Don't sanitise everything. Most of the value sits in a few well-chosen documents, and most of the risk sits in the rest. An illustrative sort for a community interest company that writes evaluations and funding bids for other charities:

DocumentTierWhat goes in the library
Evaluation framework used on 12 projects1: reuse freelyThe whole framework
Final evaluation report for a youth charity2: sanitise firstStructure, method section, two anonymised findings
Funding bid that won a large grant2: sanitise firstSection headings, the theory-of-change layout, budget structure without figures
Survey responses from service users3: neverNothing
Board paper on a client's financial difficulties3: neverNothing
Interview guide for volunteers1: reuse freelyThe whole guide

Tier 1 is your own intellectual property. Tier 2 is where the work is. Tier 3 stays in the client folder and never goes near the library or a prompt.

Sanitising a document, with a prompt that helps

Sanitising means more than deleting names. Remove or generalise anything that identifies the client, the people, the place or the moment. Here's a sentence from an evaluation report, first as written and then sanitised:

Before: "[Charity name]'s mentoring programme, led by its founder [name] after she left the council's youth service in 2021, supported 214 young people across the two estates in its first year, with a budget of $186,000 from the city's recovery fund."

After: "A youth charity's mentoring programme, in its first year, supported just over 200 young people across two neighbourhoods on an annual budget of under $200,000 from a single public funder."

The client name, the founder's history, the year, the exact count, the exact budget and the funder have all gone or been rounded. What remains is still useful: scale, structure, funding model.

AI can do a first pass. The prompt:

Sanitise the text below for reuse in an internal library.
Replace: organisation names with a generic description; people's names
and job titles with roles; places with generic terms; dates with
relative terms; exact figures with rounded ranges.
Remove: quotes from named or identifiable people; anything describing
a unique feature that would identify the organisation.
Then list every change you made, and list anything you were unsure
about.

TEXT:
[paste section]

For a funding bid, the change list might come back like this (illustrative):

Changes: "[farm name]" -> "a community farm";
"[CEO name], CEO" -> "the chief executive"; "March 2024" ->
"in its third year"; "$48,250" -> "just under $50,000".
Unsure: "the only community farm offering equine therapy for
children with autism in the region" - kept as it describes the
service.

The change list is useful, and the "unsure" item is exactly the one to fix. A description that says "the only" anything identifies the client as surely as its name. Generalise it to "a community farm offering animal-assisted therapy". The model flagged it but kept it; a person has to make the call. For the mechanics of the first pass, anonymising client data before you paste it into AI goes into more detail.

The text is only part of the file. A realistic slip: a sanitised report went into the library as a Word document whose body was clean, but the file still carried three margin comments from the client's finance director, tracked changes showing the original figures, and the client's name in the document properties as the author's company. The AI in the library read all of it, and a later draft quoted one of the comments. Before uploading, run Word's Document Inspector (File, Info, Check for Issues, Inspect Document) to remove comments, revisions and personal information, then check any charts: an embedded chart often keeps its original spreadsheet data, labels included, even after you've retyped the caption. Pasting the clean text into a fresh document is the quicker route for short excerpts.

The recognition test

Before a sanitised document goes into the library, ask two questions. Could the client recognise themselves if they read it? Could a competitor or funder who knows the sector work out who it's about? If either answer is yes, generalise further or drop the excerpt. Small sectors make this harder: in a field with a dozen organisations, "a regional arts charity with a touring puppet theatre" might as well be a name.

Here's an excerpt that passes the name check and still fails the test, from an HR consultancy's library: "A 140-person engineering business moved its whole workforce to a four-day week within six weeks of losing a case brought by staff over working hours." No name appears, but the headcount, the sector, the unusual policy and the legal trigger together point to one firm that anyone in that trade press would know. The version that went in: "A mid-sized manufacturer changed its working-hours policy quickly after a legal dispute; the lesson was that rushed consultation with staff created more grievances than it settled." The reusable point is the lesson, and it survives without any of the identifying detail.

Have someone who didn't work on the project do this test. The person who wrote the report can't un-know the client.

The riskiest material is email threads and meeting transcripts. They carry the most identifying detail (names in signatures, off-the-record remarks, forwarded attachments) and the least reusable value. In an illustrative case, a CIC that tried to sanitise a year of client emails for "tone examples" spent a day on it and ended up keeping four paragraphs. Leave emails and transcripts in tier 3 by default, and if a message contains a genuinely good explanation, rewrite that explanation from scratch as a tier 1 note in your own words.

Where the library lives and who can open it

  • A dedicated folder or project, not the client drive. ChatGPT and Claude both offer Projects that can be shared on business plans; Gemini Notebook (formerly NotebookLM) does the same job for Google Workspace teams. Upload only tier 1 and sanitised tier 2 documents. The ChatGPT side is covered in using ChatGPT Projects to keep client work separate.
  • Access for people who write proposals and reports, not the whole organisation by default.
  • No live connection to client folders. If you use Copilot or another assistant across your drive, it will still see client folders according to their permissions, so tighten those separately.
  • An owner. One person approves what goes in and removes what should come out.

The owner's main tool is a register kept next to the library, one row per document. An illustrative extract:

Library fileDerived fromTierContract checkSanitised / tested byAdded
evaluation-framework-v3Own IP1Not neededn/aWeek 1
youth-mentoring-methodClient file 2023-0142Standard terms, know-how clauseStaff A / Staff BWeek 2
bid-structure-capital-grantClient file 2024-0062Funder terms allow anonymised reuseStaff B / Staff AWeek 2

The "derived from" column is the one that earns its keep. When a former client later asks you to delete everything connected to their project, you can find and remove every excerpt that came from their file in minutes, and tell them honestly that you have. Without it, you'd be reading the whole library, and still guessing.

Drafting from the library with a paper trail

When you draft from the library, make the AI tell you what it used. That way you can check nothing slipped in from memory or from the wrong file:

Draft the "Our approach" section of a proposal for [new client, in
general terms]. Use ONLY documents in this project. After the draft,
list each library document you drew on and the sentence it informed.
Do not include any figure, name or quote that is not in the new
client's brief below.

A realistic mistake this catches: in one illustrative case, a draft for a new client included "increasing volunteer retention from 58% to 81%", a figure from a sanitised report where the rounding had been missed. The source list pointed straight to it. Without the list, it would have gone into the proposal as if it were a promise. For the drafting itself, writing proposals in under an hour picks up from here.

A six-person CIC's first month

The illustrative community interest company here has six staff and about 80 past reports and bids on its drive. Sanitising all 80 would take weeks and add little. Instead, in the first month:

  • Week 1: the director lists the 20 documents staff most often go back to and sorts them into tiers. Six are tier 1, eleven tier 2, three tier 3. About 3 hours.
  • Week 2: two staff sanitise the eleven tier 2 documents using the prompt, then swap to run the recognition test on each other's work. About 1.5 hours per document, 16 hours in total.
  • Week 3: the library goes into a shared Project on a two-seat business plan (the minimum on ChatGPT Business and Claude Team, about $50 a month on monthly billing). The director checks sharing links on the old client folders and removes eight "anyone with the link" shares.
  • Week 4: the team drafts two proposals from the library with the source-list prompt.

The drafting time for a proposal's method and approach sections fell from about half a day to under two hours, mostly because staff stopped hunting through old folders for "that framework we used". The eight removed sharing links were the unplanned win. Adding documents after that is part of closing each project: whoever writes the final report sanitises one excerpt for the library before the file is archived. The general approach to building a knowledge base like this is in building a company knowledge base AI can answer from.

A quarterly leak check

  • Search the library for every current and past client name. There should be no hits.
  • Search for exact currency figures and percentages with decimals; sanitised documents should mostly have ranges.
  • Review who has access to the library and to client folders; remove leavers and anyone who no longer needs it.
  • List external sharing links on client folders and remove any that aren't needed.
  • Ask the AI in the library project: "Which organisations are mentioned in these documents?" Anything it names is a document to fix.
  • Check each AI tool's plan and training settings haven't changed, especially after renewals or plan switches.

The organisations question is worth showing, because its answer is rarely empty on the first run. An illustrative reply from a library of 17 documents:

Organisations mentioned:
- "a community farm" and "a youth charity" (generic descriptions)
- "the [funding programme name] pilot" in youth-mentoring-method,
  section 3
- "[Client initials] Trust" in bid-structure-capital-grant,
  footnote 4
- Microsoft, Google (tools mentioned in the method notes)

The generic descriptions and the software names are fine. The other two need fixing: a named funding programme narrows the client down almost as well as a name, and initials in a footnote are a name that the search for full client names missed. Both came from sections the sanitising pass didn't reach, footnotes and a pasted table, which tells you where to look first next quarter. Log each fix in the register so the next check starts from a known state.

Twenty minutes a quarter is cheap insurance against the one email that starts "we noticed our figures in someone else's report".

Reuse, anonymising and confidentiality: follow-up questions

Is anonymised client data still personal data?

Only truly anonymised data falls outside data-protection law such as the GDPR, meaning nobody could reasonably re-identify the people in it. Replacing names with codes while keeping a key is pseudonymisation, and that data is still personal data. Most sanitised consultancy documents sit somewhere in between, so treat the library as confidential anyway, and take advice if it holds anything about individuals rather than organisations.

What if client work was already pasted into a personal ChatGPT account?

Find out what went in and when, then check that account's data settings: switch off the model-training option, delete the relevant chats, and move future work to a business plan. If the material was confidential under a client contract or included personal data, record what happened and consider whether you need to tell the client or take data-protection advice. Don't guess; it depends on what was shared.

Should we tell clients we reuse our past work?

Clients generally expect a firm to bring experience from earlier projects; that's what they're paying for. What they don't expect is their own material turning up elsewhere. Saying in your proposals and terms that you draw on anonymised methods and templates from past work, and that client-specific material is never reused, sets the expectation clearly and gives clients the chance to ask for stricter handling.

Further reads

Sources: Microsoft 365 Copilot documentation on permissions and data access; ChatGPT Business and Claude Team plan pages (business-data defaults, shared Projects); Google help on Gemini Notebook.

Want past work reusable without the confidentiality risk?

On a 1:1 call we'll look at where your past client work sits, agree what can be reused, and plan a sanitised library and access setup your team can maintain.

Book a 1:1 call with me