Questions to Ask Before Buying AI That Touches Client Data

Coding Liquids tutorial cover featuring Sagnik Bhattacharya for Questions to Ask Before Buying AI That Touches Client Data.
Coding Liquids tutorial cover featuring Sagnik Bhattacharya for Questions to Ask Before Buying AI That Touches Client Data.

Before any client data goes in, get written answers to five things: whether the tool trains on your data, how long it keeps prompts, files and outputs, which model provider and sub-processors see it, who can access it at the vendor and in your firm, and how breaches are reported. Then check the answers appear in the contract.

The contract is the part people skip. A sales call will tell you "we never use your data", but what binds the vendor is the terms of service and the data processing agreement (the contract that sets out how the vendor handles personal data on your behalf). If a reassuring answer isn't written into one of those, you don't have it. The questions below are written for professional firms whose client files carry confidentiality duties, not for general software buying; for that, start with questions to ask an AI vendor before you sign anything.

Follow me on Instagram@sagnikteaches

Training and reuse of your clients' material

  1. Is our content, including uploads and outputs, used to train or improve any model? Good answer: "No, by default, and that is stated in section X of our terms." Red flag: "Only anonymised data" or "only to improve our services", which can mean exactly what you feared.
  2. Can your staff read our content, and when? Good answer: only for support you request, or for abuse investigations, with access logged. Red flag: routine human review of conversations for quality.
  3. Do you use aggregated or "de-identified" data from customers? Good answer: a clear description, such as usage counts only. Red flag: vagueness about what is aggregated. Client matters can be identifiable from details even without names.

The difference between business and consumer plans matters here. Whether ChatGPT trains on client data on business and free plans explains it for the most common tool.

Connect on LinkedInSagnik Bhattacharya

Retention, deletion and leaving the vendor

  1. How long do you keep prompts, uploaded files and outputs? Ask for each separately. Files are often kept longer than chat text.
  2. Can we set a shorter retention period ourselves? Good answer: an admin setting, with the options named.
  3. When we delete something, is it deleted from backups and logs, and how long does that take? Good answer: a stated period, for example "removed from backups within 30 days".
  4. When we cancel, how do we export our data and how do you confirm deletion? Good answer: a documented export and a written deletion confirmation on request. What happens if your AI vendor shuts down covers the worst case.

The model underneath and who else touches the data

Most AI tools sold to small firms are built on a model from a large provider such as OpenAI, Anthropic or Google. Your data travels to that provider too, on the vendor's contract terms rather than yours. So:

Subscribe on YouTube@codingliquids
  1. Which AI model providers do you send our data to? Good answer: named providers, and whether this changes by feature.
  2. What retention applies at the model provider? This is where the detail lives. OpenAI's API, for example, keeps abuse-monitoring logs for up to 30 days by default, and zero data retention is available only to eligible customers with OpenAI's prior approval. Anthropic's commercial API deletes inputs and outputs within 30 days by default, with zero retention by agreement.
  3. Does retention differ by model? It can. In June 2026 Anthropic began keeping prompts and outputs for 30 days on its most capable "covered" models even for customers with zero-retention arrangements, and later added a separate route back to zero retention by application. A vendor that says "we have zero retention with our provider" should be able to tell you whether that holds for every model it uses.
  4. Where is our data stored and processed, and who are your other sub-processors? Good answer: a published sub-processor list, the countries involved, and a commitment to notify you before adding one. Your data-protection adviser can tell you whether those countries raise issues under data-protection law such as the GDPR.

Access controls inside your own firm

Professional firms often forget this group. A tool can be secure against outsiders and still show one client's file to the wrong colleague.

  1. Do you support single sign-on and multi-factor sign-in? Good answer: yes, with your Microsoft or Google accounts.
  2. Can we restrict what each user sees, down to client or matter level? Good answer: permissions that mirror your practice system, or separate workspaces per client. Red flag: "all users in your organisation share one knowledge base".
  3. Is there an audit log of who uploaded, viewed or exported what? Good answer: a log admins can export, kept for at least a year.
  4. If the tool searches our files, does it respect the permissions we already set? Good answer: yes, and they can explain how. This matters for any firm running information barriers between clients.

Questions 13 and 15 tend to fail in a trial rather than on paper. An illustrative case: a consultancy advising two competing retailers loads a sample of past work into a tool's shared knowledge base. A consultant on the first client's team asks it for "examples of our pricing recommendations" and gets back a slide from the second client's pricing review. Nothing had been misconfigured; the tool simply had one library for the whole firm. That is a red on question 13 however good the other answers are, unless the vendor offers separate workspaces per client.

Question 14 can be tested before you sign. On a trial account with dummy files only, upload a document from one user, open it from a second user, export it from the second, then download the audit log. All three events should appear with names and times. If the upload is logged but the view isn't, the log won't tell you who read a client's file, which is usually the question you'll need answered.

Security evidence and when things go wrong

  1. Can we see your SOC 2 Type II report or ISO 27001 certificate? Type II means an auditor tested controls over a period, not just on one day. Check the certificate's scope covers the product you are buying. SOC 2 and ISO 27001 explained covers what to look for. Some vendors also certify against ISO/IEC 42001, the AI management system standard published in December 2023; useful, but not a substitute.
  2. How quickly will you tell us about a breach affecting our data? Good answer: a fixed number of hours in the contract. Red flag: "without undue delay" and nothing else.
  3. What is your liability cap for a data breach? Many vendors cap liability at the fees paid in the last twelve months. For a tool holding client files, ask whether data breaches have a higher cap. The sum shows why: twelve seats at $25 a month is $3,600 a year, so a fees-based cap means $3,600 is the most you could recover, which may not cover the partner time spent writing to every affected client, let alone any claim. A separate, higher cap for data-protection breaches, often a multiple of annual fees, is a common ask. Small firms won't always get it, but asking costs nothing.
  4. What happens to our data if you are acquired? Good answer: the terms continue to bind the new owner, and you can terminate and have data deleted.

The data processing agreement is where most of these answers should end up. What to check in an AI vendor's data processing agreement goes through it clause by clause.

Answers that sound fine and aren't

Vendors rarely say anything false. They answer a slightly different question. Three illustrative replies, and what's wrong with each:

Vendor saysWhat it doesn't sayAsk instead
"We take privacy seriously and are GDPR-compliant."Anything about training, retention or model providers"Please answer questions 1, 4 and 8 specifically, with contract references."
"We don't train on your data."Whether the model provider underneath does, or whether staff review content"Does that also apply at your model provider, and can your staff read our content?"
"All data is encrypted at rest and in transit."Who holds the keys, and who can decrypt it to read it"Which of your staff roles can access our content in readable form?"

Scoring the replies

Nineteen answers are hard to compare across two or three vendors unless you score them the same way. Give each answer one of three marks:

MarkMeaningExample
GreenClear answer, and it is in the contract, DPA or security pack"No training on customer data: Terms, section 6.1"
AmberClear answer, but only in an email, or a gap you can manage yourself"We delete files after 30 days" with no contract reference
RedNo answer, a vague answer, or an answer that fails your obligations"Aggregated data may be used to improve our models"

Then apply two rules. First, any red on questions 1, 4, 8, 13 or 17 stops client data going in until it is fixed, because those cover training, retention, the model provider, access within your firm and breach notice. Second, ambers on those same five should become contract wording before you sign: ask the vendor to add their email answer to an order form or side letter. Everything else can be amber if you record how you are managing it.

Turning an amber into contract wording is usually one sentence. If the vendor's email said "we never train on your data", the order form might carry something like this (illustrative wording for your own adviser to check, not legal advice):

"The Supplier confirms that Customer Content, including prompts, uploaded files and outputs, will not be used to train or improve any AI model, whether by the Supplier or by any sub-processor, and that this applies to every feature and model used to provide the Service."

The last clause matters. Without it, a vendor can honour the promise on its main model and not on a newer one it adds next year.

Here are the big five scored for two illustrative vendors a small law firm was comparing:

QuestionVendor AVendor B
1. TrainingGreen: terms, section 6.1Green: DPA, clause 4
4. RetentionAmber: "30 days" in an email onlyGreen: admin setting, 7 to 90 days
8. Model providersGreen: two named providersRed: "leading AI providers", none named
13. Access within the firmGreen: workspaces per matterGreen: mirrors practice system permissions
17. Breach noticeAmber: "without undue delay"Green: 48 hours, in the DPA

Vendor B looks stronger on four of the five, but its red on question 8 stops client data going in until it names its providers. Vendor A needs two ambers turned into contract wording. The firm sent both requests the same day; Vendor A agreed to add retention and a 72-hour breach notice to the order form within a week, and Vendor B's reply still named no provider, so the firm chose Vendor A.

A vendor with two or three reds on minor questions and greens on the big five is usually a better choice than one with no reds but ambers everywhere, because the second has told you nothing binding. If you are comparing vendors on price, support and fit too, fold these marks into a wider scorecard rather than keeping two separate lists.

The email to send, and how to check the answers with AI

Send the questions in one email so the answers come back in one place. A filled-in version:

Subject: Data questions before we purchase [product], [firm name]

Hello [name],

Thanks for the demo on Tuesday. Before we can put client files into
[product], we need written answers to the attached 19 questions.
Where an answer is covered by your terms of service, data processing
agreement or security documentation, please give the section number.

We'd also like copies of: your current DPA, your sub-processor list,
and your SOC 2 Type II report or ISO 27001 certificate (we're happy
to sign an NDA for the report).

We're aiming to decide by [date]. Thanks,
[name], [role]

When the answers and documents come back, a business AI plan can cross-check them. Paste the vendor's answers and the relevant DPA sections into a new chat:

Below are (A) a vendor's written answers to our data questions and
(B) sections of its data processing agreement. For each answer in A,
say whether B supports it, contradicts it, or doesn't address it.
Quote the relevant words from B. List any answer that relies only
on the sales email and not the contract.

Illustrative output: "Q4 (retention): A says 'files deleted after 30 days'. B, clause 7.2, says 'Customer Data will be deleted within 90 days of termination'. These differ: A describes files during the subscription, B only covers termination. Q9 (model provider retention): not addressed in B."

That is useful, and it still needs checking. Models miss things: a later clause allowing retention "as required for legal or compliance purposes" is easy for them to pass over. Read the flagged clauses yourself, and treat "not addressed" as a question to send back.

A surveying firm puts one tool through the questions

An illustrative eight-person surveying firm wanted a tool that turns site notes and photos into draft condition reports. Its reports include client names, property addresses and sometimes occupants' details.

The vendor's answers were good on training (no, stated in the terms), single sign-on and audit logs. Three answers were weak. Photos were kept for the life of the account, with no admin retention setting. The model provider was named, but the vendor couldn't say what retention applied there. And all users in a firm shared one library of past reports, so a junior surveyor could search every client's reports.

The firm didn't walk away. It asked for, and got, a contract addition committing to delete photos 90 days after a report was finalised, and written confirmation of the provider's retention period. The shared library it handled itself: past reports were not uploaded, and only anonymised template reports went into the library. It recorded the decision and the conditions in its AI tool register, and set a reminder to re-ask questions 8 to 11 at renewal, because model providers change.

The reminder earned its keep a year later. The vendor had added a second model provider for a new photo-description feature, listed on its sub-processor page but never emailed to customers. The firm asked question 9 about the new provider before approving the renewal, got a written retention answer, and added the provider to its register. Without the reminder, client photos would have gone to a company the firm had never heard of.

The mistake it avoided is common. Two surveyors had already tried the tool on a free trial, which ran on the vendor's standard terms, not the negotiated ones. Trial accounts count as purchases for data purposes. Close them, or keep client data out of them.

Follow-up questions on vetting AI vendors

Is a vendor's privacy policy enough?

No. A privacy policy describes what the vendor does with personal data in general and can change. What binds the vendor is the contract: the terms of service, the data processing agreement and any security schedule. Ask for those documents, and make sure the answers you were given in sales emails appear in them or are attached to them.

What if the vendor won't answer in writing?

Treat it as a no for client data. A vendor that sells to professional firms should be able to answer these questions in writing, usually from a prepared security pack. You can still use a tool like that for work with no client information in it, but record the decision in your tool register.

Do we need client consent before using the tool?

It depends on your engagement terms, your clients' own contracts and your professional rules. Many firms cover AI in their engagement letter and ask specific consent only where a client's contract requires it. If you are unsure whether your terms are enough, ask your solicitor or data-protection adviser before client files go in.

Further reads

Sources: OpenAI API data controls documentation; Anthropic privacy article on organisational data retention and Claude help article on covered models; vendor documentation cited in the body. Checked September 2026. The surveying firm example is illustrative.

Vetting an AI tool that will hold client files?

On a 1:1 call we'll go through the vendor's answers and documents with you, flag the gaps that matter for your client obligations, and decide whether to proceed, negotiate or walk away.

Book a 1:1 call with me