Sort your data into four tiers (public, internal, confidential, restricted) by asking what harm a leak would cause, then attach one AI rule to each: any tool for public, company AI accounts for internal, approved business tools with identifying details removed for confidential, and no AI at all for restricted. A small firm can do this in an afternoon.
The classification itself is the easy part. What makes it work is attaching a plain rule to each tier that staff can apply in the three seconds before they paste something into a chat window. Below are the tiers, a scoring test for anything borderline, a fitness studio classified line by line, and a one-page rule to hand to your team.
Four tiers, each with its AI rule
Large companies use five or six levels. A small business needs four at most, because every extra tier is another judgement call someone gets wrong on a busy Friday.
| Tier | The test | Typical examples | AI rule |
|---|---|---|---|
| Public | Already published, or you'd be happy to publish it | Website copy, price list, opening hours, published reviews | Any AI tool, including free ones |
| Internal | Not secret, but you wouldn't post it; a leak would be awkward, not harmful | Rotas, internal how-to notes, marketing plans, supplier lists | Company AI accounts only (business plans, or personal accounts with training switched off if that's all you have) |
| Confidential | A leak would harm a person or break a promise to a client | Client names with contact details, client notes, staff pay, contracts, quotes to named clients | Approved business AI tools only, and remove names and identifiers unless the task needs them |
| Restricted | A leak could cause serious harm, or the law sets special rules | Health information, card numbers, bank details, passwords, ID documents, disciplinary records | No AI tool, unless the owner has approved a specific tool and task in writing |
The rule column assumes you've already moved client work onto business plans. If you haven't, do that alongside this exercise, because the internal and confidential rules depend on it. Stopping AI tools training on your business data covers the settings and plans.
Step 1: List data types, not files (30 to 60 minutes)
The instinct is to open the shared drive and start labelling files. Don't. A small business has thousands of files and a few dozen kinds of information. Classify the kinds, and the files follow.
The fastest way to list them is to walk one customer through your business from first contact to final invoice, writing down every piece of information you collect or create at each point. Then do the same for one employee, from job advert to leaving. Finally, list the business's own records: finances, contracts, plans, passwords.
An AI assistant is useful for making sure you haven't missed anything, as long as you describe your business rather than paste real records into it:
I run a [type of business] with [number] staff. We take bookings by
[channels], keep client records in [tools], and handle payments through
[tools].
List every type of information a business like mine typically collects
or creates about (a) customers, (b) staff, (c) the business itself.
Group them by stage: enquiry, booking, delivery, payment, aftercare,
and for staff: hiring, employment, leaving.
For each item, give one line on why it might be sensitive. Don't
assume anything about where we operate.
Expect 25 to 45 data types. Cross out anything you don't actually collect, and add anything the list missed. The AI's list is a prompt for your memory, not an authority.
Here's an excerpt of the kind of list you get back for a small personal training studio (illustrative output, trimmed to the booking and delivery stages):
BOOKING
- Name, email, phone: identifies the client; misuse enables spam or fraud
- Preferred session times: reveals routine and when someone is away from home
- Payment card details: direct financial risk if exposed
- Wearable fitness data (heart rate, sleep): health-related
DELIVERY
- Session attendance records: shows patterns of behaviour
- Exercise programmes and weights lifted: tied to a named person
- Body composition scans: health data
- Trainer observations about mood or energy: can reveal health issues
Three fixes before you use it. The studio doesn't collect wearable data or run body-composition scans, so both lines go. It takes card payments through its booking system and never sees the card number, so that line moves to "held by our payment provider" rather than disappearing, because staff still need to know never to type a card number into a chat. And the list missed two things only the owner would know: the signed waiver form every client completes at their first session, and the CCTV at the entrance. Both end up restricted.
Step 2: Three questions that settle the tier
Most items classify themselves. For the ones that don't, ask three questions and take the highest tier any answer points to:
- Would a leak hurt a person? If it could embarrass, endanger or cost someone money, it's at least confidential. If it reveals health, finances, identity documents or anything someone would expect to stay between them and you, it's restricted.
- Would a leak hurt the business or break a promise? Pricing strategy and supplier terms are internal. Anything covered by a client contract, a confidentiality clause or an NDA is confidential at least.
- Does a law or an industry rule treat it specially? Data-protection law such as the GDPR gives health information and some other categories stricter protection. The card-payment industry's security rules strictly limit where full card numbers may be stored. Anything in this group is restricted by default.
If you're unsure whether a category of information carries special legal duties in your line of work, that's a question for your data-protection adviser or a solicitor, not for a chatbot.
To see the test at work outside a gym, take three borderline items from an illustrative five-person bookkeeping practice. Its client list with company names and the monthly fee each pays: question 1 is "barely" (the contacts are business ones), question 2 is "yes" because fees are covered by engagement letters, so it's confidential. A spreadsheet of staff holiday dates: questions 1 and 2 are both "slightly", nothing in question 3, so internal. A folder of scanned passports collected for anti-money-laundering checks: question 3 settles it instantly as restricted, however rarely anyone opens it. Taking the highest answer is what stops the fee list sliding down to internal because it "only has company names".
Timing raises two more questions. Information can move down a tier, but only by your own decision: next quarter's price rise is internal until the day you announce it, then public. Someone else publishing something doesn't move your copy. A studio client who posts her own progress photo on social media hasn't made the studio's copy public; it stays restricted, and only the testimonial she agreed the studio could publish counts as public. The tier follows how you came to hold the information, not what the person has chosen to share elsewhere.
A personal training studio, classified line by line
Here's how an illustrative studio with an owner, three trainers and a part-time administrator might classify its data. Fitness businesses are a useful example because they hold more sensitive information than owners tend to assume.
| Data type | Tier | Why |
|---|---|---|
| Class timetable, prices, trainer bios | Public | Already on the website |
| Client testimonials posted with permission | Public | Published by consent |
| Marketing calendar, campaign ideas | Internal | Harmless if leaked, but not for competitors |
| Staff rota, equipment supplier list | Internal | Operational, no personal harm |
| Monthly revenue by class type | Internal | Aggregated; no individual can be identified |
| Client names, emails, phone numbers | Confidential | Personal data; harm if misused |
| Session notes (exercises, weights, attendance) | Confidential | About an identifiable person |
| Individual payment history and package balances | Confidential | Financial information about a person |
| Trainer pay rates and contracts | Confidential | Staff personal and financial data |
| Health questionnaires and injury history | Restricted | Health information has special legal protection |
| Progress photos and body measurements | Restricted | Intimate, identifiable, health-related |
| Incident and accident reports | Restricted | Health details plus possible legal claims |
| Card details, bank details, system passwords | Restricted | Direct financial harm if leaked |
The line that surprises this studio's trainers is session notes. A note saying "Tuesday: squats 3x8 at 60kg, knee felt better" looks harmless. It becomes confidential because it's about a named client, and it tips into restricted when it mentions the knee injury. In practice the studio told trainers to write programme ideas for AI in general terms ("a client in their fifties returning from a knee injury, three sessions a week") and never to paste the actual notes. Using ChatGPT safely for client programmes shows how trainers can get useful output that way.
The exercise took the owner and administrator about three hours: an hour listing, an hour classifying, and an hour arguing about the borderline cases, which is time well spent.
Five cases where small-business schemes go wrong
Mixed documents
A document takes the tier of the most sensitive thing in it. A class-planning spreadsheet is internal until someone adds a column called "injuries to watch", at which point the whole file is restricted. Keep restricted details in their own place so everyday files stay in a lower tier.
AI output inherits the tier of its input
A summary of confidential session notes is confidential. People forget this and paste an AI-written summary into a tool they'd never have put the originals in.
"Anonymised" that isn't
Removing a name isn't enough if the remaining details identify the person. "The only 70-year-old client who trains at 6am on Tuesdays" is still identifiable to anyone who knows the studio. Anonymising client data before you paste it into AI covers the techniques properly.
A before/after makes the difference plain. The prompt a trainer first wrote:
Write a 6-week plan for [client's full name], 68, who trains with me
at 6am Tuesdays and Thursdays, had a hip replacement at the start of
the year and is scared of falling since her husband died.
The version that gets the same quality of plan and sits comfortably in the confidential tier:
Write a 6-week strength plan for a client in their late sixties,
returning to training after hip surgery earlier this year. Two
sessions a week. Low confidence about balance, so build in
progressions that feel safe. No floor work in the first fortnight.
The name went, the exact age became a range, the time slots went (they identify her to any other early-morning member), and the bereavement disappeared entirely, because the plan doesn't need it. The detail the AI actually uses, "low confidence about balance", survived.
Screenshots, photos and voice notes
Staff classify documents and forget the screenshot of a booking screen or the voice memo recorded after a session. These carry the same tier as the information in them, and they're the files most often dropped into a chat app on a phone.
Here's how it typically shows up. A trainer wants a friendly reminder text for tomorrow's early clients, so she screenshots the day view of the booking app and drops it into a free chat app with "write a reminder for these people". The reminder is fine. The screenshot, though, showed five full names, five mobile numbers and a red "medical note" flag beside two of them, all now sitting in a personal account with training switched on. Nobody noticed until the owner asked where the draft had come from. The fix was a line added to the staff rule and a saved template reminder, so the screenshot was never needed again.
Connecting a tool changes the exposure
Classification is about what you paste in. Once an AI tool is connected to your inbox or drive, it can read whatever that user can open, whatever the tier. Before you connect anything, check the folders holding restricted data are shared only with the people who need them.
Step 3: Make the tiers visible where people work
A scheme that lives only in a policy document gets forgotten. Put the tier where people see it:
- Folder names. A top-level folder per tier for anything restricted ("RESTRICTED - health forms") is crude but effective. People hesitate before dragging files out of a folder with that name.
- File-name prefixes. Starting confidential files with "CONF_" costs nothing and shows up in every search result and attachment.
- Built-in labels. Microsoft 365 Business Premium includes sensitivity labels you can create and apply by hand to emails and files; automatic labelling needs a higher licence or an add-on. Google Workspace supports classification labels for Drive files and Gmail on Business Standard and Business Plus, but not Business Starter.
- Email subject tags. "[CONFIDENTIAL]" at the start of a subject line reminds the recipient too.
Don't label every file on day one. Label the restricted material first, then confidential client folders, and let internal be the unlabelled default. That order puts the effort where a mistake costs most.
For the studio, the top of the shared drive ended up looking like this, with public material left where it already lived on the website:
Studio shared drive
Admin and planning (internal: rotas, marketing, suppliers)
CONF - Clients (one folder per client: programmes, notes)
CONF - Staff (contracts, pay)
RESTRICTED - Health forms (questionnaires, waivers, injury history)
RESTRICTED - Incidents (accident reports)
shared with: owner, administrator only
The trainers can still see their own clients' programmes. What changed is that the health questionnaires no longer sit inside each client folder, which is where they had been, so a trainer opening a client folder to draft a programme no longer sees the injury history next to it.
Step 4: The one-page rule for staff
This is the document people will actually read. Adjust the tool names and keep it to one page:
WHAT CAN GO INTO AI TOOLS
PUBLIC (website, prices, published reviews)
Any AI tool.
INTERNAL (rotas, plans, how-to notes, supplier lists)
Only our company AI accounts: [list].
CONFIDENTIAL (client names and contacts, client notes, quotes,
staff pay, contracts)
Only [approved tools]. Remove names and contact details
unless the task needs them. Treat AI summaries of this
material as confidential too.
RESTRICTED (health info, photos of clients, card or bank details,
passwords, ID documents, incident reports)
Never into any AI tool without written approval from [name]
for that specific tool and task.
NOT SURE? Treat it as the higher tier and ask [name].
Screenshots, photos and voice notes follow the same rules.
Pin it next to the computers and include it in onboarding. It also slots straight into a fuller document if you're writing an AI usage policy.
How to check the scheme is working
Run three checks a month after you introduce it:
- Spot-check ten files in the shared drive at random. Is each one in a folder that matches its tier? If more than two are misplaced, the folder structure needs work, not the staff.
- Ask each person three questions, informally: which tier is a client's email address, which is an injury note, which is the rota? If people hesitate on the same item, rewrite that line of the rule.
- Review the AI tool list. Has anyone started using a tool that isn't on the approved list, and what tier of data has gone into it?
At the studio, the first spot check found eight of ten files where they belonged. The two that didn't were telling. A signed waiver had been scanned straight into a client folder, because the office scanner still saved to "CONF - Clients" by default; changing that one setting fixed it for good. A pay spreadsheet sat in Admin and planning because the payroll template lived there. Neither was carelessness. Both were the structure pulling files to the wrong place, which is why the check blames the folders first. In the question round, two trainers hesitated over whether a client's first name alone in a session note was confidential. It is, because the note sits with the rest of that client's record, so the staff rule gained one line: "A first name plus anything about a session is confidential."
After that, revisit the classification whenever you start collecting a new kind of information or connect a new tool. Once the tiers are settled, they also tell you which data is safe to use when you prepare your business data for an AI project, and which needs to be stripped out first.
Further reads
- How to Keep Customer Data Private When Your Team Uses AI — Day-to-day habits that keep confidential data out of the wrong tools.
- Is It Safe to Put Customer Data Into ChatGPT? — The specific question most owners start with, answered plan by plan.
- Do You Need a DPIA Before Using AI Tools? — When restricted data needs a formal risk assessment first.
- What Is Data Loss Prevention, and Does a Small Business Need It? — Software that enforces the tiers once labels are in place.
- How to Clean Up SharePoint Permissions Before Turning On Copilot — Folder permissions matter as much as labels once Copilot is on.
- GDPR and AI Tools: What a Small Business Must Do — The data-protection duties behind the confidential and restricted tiers.
- What to Check in an AI Tool's Privacy Policy and Terms — Seven checks for any AI tool's privacy policy and terms, the words to search for, and how ChatGPT, Claude, Gemini and Copilot compare.
- How to Brief a Developer on a Custom AI Workflow You Need — The eight parts of a developer brief for an AI workflow, a copyable requirements template, and a worked example from a dog-grooming salon.
- How to Get Started With AI in Your Small Business: First 7 Steps — Seven steps that take about a month, from a five-day time log to a keep-or-drop decision, with a yoga studio's numbers at each stage.
- How to Clean Up Customer Records Before You Add AI — Four clean-up passes that stop AI emailing people twice, or at all when they said no, with matching rules and a merge log.
- How to Write a One-Page AI Strategy for Your Business — The seven boxes a one-page AI strategy needs, a filled-in garden centre example, and five tests that show whether your page will guide real decisions.
- How to Train Staff to Use AI in a Small Business — A three-layer training plan with learning outcomes, a 75-minute session agenda, a copyable error-spotting exercise and a language-school example.
- How to Build a Shared Prompt Library for Your Team — An entry template, three sample prompts from a physiotherapy clinic, five storage options compared and a monthly routine that keeps the library useful.
- AI Use Case Template: Score Every Idea on One Page — A one-page template with scoring anchors, knock-out questions and a worked veterinary example for ranking AI ideas before you spend anything.
- What Business Data Should You Start Collecting Now for AI? — Seven datasets worth capturing from today (enquiries, quotes, job actuals, questions, complaints, prices, feedback), with the fields that make them usable.
- What to Do Before You Buy Any AI Tool: A 10-Point Checklist — Ten checks to run before paying for any AI tool, from pricing the job it replaces to reading the exit terms, with a kitchen-fitter worked example.
- AI Tools and AI Development: The Complete 2026 Guide — the AI hub, including every tutorial in the AI-for-business series.
Sources: Microsoft Learn and Microsoft 365 admin documentation on sensitivity labels in Business Premium; Google Workspace Admin Help on classification labels and supported editions. Checked September 2026.