Is Your Business Data Ready for AI? A Clean-Up Checklist

Coding Liquids tutorial cover featuring Sagnik Bhattacharya for Is Your Business Data Ready for AI? A Clean-Up Checklist.
Coding Liquids tutorial cover featuring Sagnik Bhattacharya for Is Your Business Data Ready for AI? A Clean-Up Checklist.

Your data is ready enough for AI when, for the specific job you want AI to do, it sits in one findable place, a sample of 50 records has few errors (under about 5% duplicates or blanks in the fields that matter), it's current, and you're allowed to use it that way. You can check that in an hour.

The mistake most businesses make is trying to clean everything before starting. Readiness depends on the job. A website chatbot needs current prices and policies, and doesn't care whether your customer list has duplicates. An automated quote follow-up needs clean customer emails and accurate job statuses, and doesn't care about last year's price list. Pick the job first, then check only the data it touches.

Follow me on Instagram@sagnikteaches

Start with the job, then list the fields it relies on

Write down the one AI job you want working first, and the handful of fields that must be right for it to work. Some common pairings:

Connect on LinkedInSagnik Bhattacharya
AI jobData it readsFields that must be right
Website or phone chatbot answering questionsPrice list, services, policies, opening hours, service areaCurrent prices, what's included, cancellation terms
Automatic quote follow-upsCRM or quotes spreadsheetCustomer email, quote date, quote status, amount
Cash-flow forecast from invoicesAccounts software or invoice exportInvoice date, due date, amount, paid date, customer
Copilot or Gemini searching your filesShared drives, email, documentsFile permissions, current versions, sensible file names
Rebooking remindersJob historyCustomer, service type, last visit date, contact consent

That list of fields is what you test. Everything else can wait.

Subscribe on YouTube@codingliquids

The 50-record sample test

Take 50 records at random from the data behind your chosen job. Not the first 50, which are often the oldest, and not ones you picked because you remember them. In a spreadsheet, add a column with the formula =RAND(), sort by it, and take the top 50. Then check each record against the fields on your list and tally the problems.

A filled-in result for an illustrative cleaning company testing its client list before automating rebooking reminders:

CheckProblems in 50RateVerdict
Same client appears twice (duplicate)48%Fix before automating
Email missing or clearly wrong612%Fix before automating
Last visit date missing24%Acceptable; fix as you go
Service type written inconsistently ("deep", "Deep clean", "DC")1938%Standardise; quick fix
Client has asked not to be contacted, but no flag12%Serious: one is too many

Reading it: an 8% duplicate rate would mean some clients get two reminders. A 12% bad-email rate means one in eight reminders goes nowhere. The inconsistent service types look alarming at 38% but take ten minutes to fix with find-and-replace. And the one unflagged "don't contact" client matters more than all the others, because it's the kind of mistake that loses a customer and can breach data-protection rules. As a rule of thumb, under 5% on a field is fine to start with and fix as you go; above 10% needs fixing first; any consent or permission problem gets fixed regardless of rate.

If the job relies on documents rather than records, as a chatbot or a file-searching assistant does, adapt the test: pick ten documents the AI would read and ask of each one whether it's the current version, whether its text can be selected, whether it has a date and owner, and whether anything in it contradicts another document. Two contradictions in ten documents is a red flag, because the AI will pick one of the two answers at random and say it confidently.

Checklist 1: can the AI find it?

  • There is one master copy. Why: AI reading "Prices 2025 FINAL v3" and "Prices NEW" will quote from both. Check: search your drive and email for the document name; archive every copy but one.
  • It's in a format the tool can read. Why: a scanned PDF is a picture; a spreadsheet with merged cells and colour-coding as the only status is unreadable to most tools. Check: can you select the text in the PDF? Does each spreadsheet column hold one kind of thing, with one header row?
  • It's where the tool can reach it. Why: a chatbot reads your website or uploaded documents, not the notebook in the van. Check: list where each source lives and whether the chosen tool connects to it.
  • File and folder names say what's inside. Why: tools that search your files lean on names and headings. Check: could a new starter find the current price list in under a minute?

A realistic example: a removals firm wanted an AI tool to answer staff questions from its procedures. Most procedures were photos of printed sheets taken on phones years ago. The AI read some of them, garbled others and missed handwritten amendments entirely. Retyping twelve documents, a day's work, did more for the project than any setting.

Checklist 2: can the AI trust it?

  • Duplicates are merged. Why: duplicates double-count in forecasts and double-send in automations. Check: in Excel, Data > Remove Duplicates on a copy, matching on email or phone; in Google Sheets, Data > Data cleanup > Remove duplicates. Look at what it would remove before accepting.
  • One field holds one thing. Why: "[mobile number] / wife's mobile" in the phone field breaks any automation that dials or texts. Check: scan the key columns for extra text.
  • Formats are consistent. Why: dates as 03/04/25, 3 April and "Apr" confuse both formulas and AI. Check: sort each date column; text dates sort in the wrong place.
  • Categories come from a fixed list. Why: AI can group "Deep clean", "deep" and "DC", but it may also merge things that differ. Check: list the distinct values in each category column and cut them to a short set.
  • Stray spaces are gone. Why: "Smith " and "Smith" look identical and don't match. Check: Excel's TRIM function or Sheets' Trim whitespace option.
  • Numbers are numbers. Why: currency symbols typed into the cell, or "approx 450", stop totals working. Check: a SUM of the column should equal what you expect.

Checklist 3: is it current?

  • Prices and policies have a "last checked" date. Why: AI repeats whatever it's given, with confidence. Check: each document states its date and owner at the top.
  • Closed and lost records are marked. Why: a follow-up automation will chase a quote you lost in March unless its status says so. Check: filter for records with no status change in six months.
  • Old versions are archived, not deleted. Why: you want history available for questions, but out of reach of the tool. Check: an "Archive" folder the AI tool isn't connected to.

An example of this going wrong: an illustrative plumbing firm connected a chatbot to its shared drive. The drive held both the current call-out charge and a price sheet from two years earlier, and the bot quoted the older, lower figure to several customers in its first week. Nothing was wrong with the bot. The fix was moving the old sheet to an archive folder and adding a "prices valid from" line to the current one.

A price and policy sheet an AI can rely on

For chatbots and assistants, the most important "data" is often a single document. A filled-in header and first rows for an illustrative plumbing firm's sheet, written so both staff and an AI tool read it the same way:

PRICE AND POLICY SHEET
Owner: office manager    Last checked: 1 September 2026
Valid from: 1 September 2026 until replaced
Replaces: all earlier price sheets (archived, not for quoting)

Call-out charge (first 30 minutes, weekdays 8am-6pm): [amount]
Each further 30 minutes: [amount]
Evenings, weekends: call-out [amount]; emergency only
Boiler service: fixed price [amount], includes safety check
We do not: quote fixed prices for leaks or blockages by phone;
  give prices for work outside our service area.
Cancellation: free up to 24 hours before; after that [amount].

Three details make this AI-friendly: a date and owner at the top, a "we do not" section that stops a bot improvising, and one fact per line. The same sheet is what you'd load into a chatbot's knowledge, as covered in training a chatbot on your FAQs, policies and prices.

Checklist 4: are you allowed to use it this way?

  • You know what personal data is involved. Why: names, contact details, addresses and anything about health or finances carry obligations under data-protection law such as the GDPR. Check: list the personal fields the job needs, and drop the rest.
  • Customers were told their data might be used this way. Why: marketing reminders need a lawful basis, and "don't contact" requests must be honoured. Check: your privacy notice and your consent or opt-out field.
  • The AI tool's terms fit the data. Why: business plans such as ChatGPT Business, Claude Team, Microsoft 365 Copilot and Gemini in Workspace don't train on business content by default; consumer plans rely on the model-training switch in privacy settings. Check: the plan and settings for the account you'll use.
  • File permissions are right before AI searches them. Why: Copilot and similar tools show people anything they already have access to, so a payroll sheet shared with "everyone" becomes easy to find. Check: review sharing on folders holding pay, HR and client-confidential files.

The permissions item surprises people most. Picture a roofing contractor that switched on an AI assistant across its shared drive. An apprentice asked it for "the rates for the church roof job" and got a summary of the pay rates sheet, which had been shared with the whole company years earlier by mistake. The AI didn't break any rule; it surfaced one. Tighten sharing first, as covered in cleaning up SharePoint permissions before Copilot.

Checklist 5: will it stay clean?

  • Entry is controlled. Why: cleaning a list that fills back up with mess wastes the effort. Check: drop-down lists (Excel's Data Validation) for categories, required fields in forms, a date picker for dates.
  • Each dataset has a named owner. Why: "everyone" maintains nothing. Check: a name at the top of each master document or list.
  • There's a monthly spot check. Why: problems creep back. Check: rerun the 50-record test on 20 records once a month; it takes 15 minutes.
  • A backup exists before any bulk change. Why: a clean-up that goes wrong can't be undone without one. Check: a dated copy saved somewhere separate; see the backup checklist before connecting AI tools.

One customer record, before and after

An illustrative record from a landscaping firm's customer list, as found:

Name:      mrs a. SAMPLE / Tom (husband)
Phone:     [mobile] or ring house after 6
Email:     asample@@example.com
Address:   12 Oak Rd (side gate code in notes)
Service:   maint - fortnightly? (check)
Last visit: sometime in june
Notes:     Gate code 4471. Dog. Prefers Tues. Owes for May.

And after cleaning, with each problem fixed or moved to the right place:

First name: A            Last name: Sample
Phone:      [mobile]
Email:      asample@example.com
Address:    12 Oak Road
Service:    Garden maintenance     Frequency: Fortnightly
Last visit: 17/06/2026 (from job history)
Access:     Side gate, code held in secure job notes
Preferences: Tuesday visits; dog on site
Account:    May invoice overdue (in accounts system, not here)
Contact consent: Service messages yes, marketing not asked

Three changes matter for AI. The gate code moved out of a free-text field an AI tool might repeat in a text message. The overdue balance moved to the accounts system, where it belongs, so a friendly reminder automation won't mention it by accident. And the consent field now exists, so a marketing automation knows to leave this customer out.

Using AI to do the tidying

An assistant is good at the tedious middle of a clean-up: standardising categories, splitting fields and flagging likely duplicates. Work on a copy, remove columns the task doesn't need, and use a plan whose data settings you've checked. A request that stops the AI guessing:

Attached is a CSV of 400 customer records (no payment data).
1. Standardise the Service column to exactly one of: Garden
   maintenance, Hedge cutting, Lawn treatment, One-off clearance.
   If a value doesn't clearly fit, write REVIEW.
2. Split Name into First name and Last name.
3. List likely duplicates (same phone or email, or same name and
   address) as pairs, without merging them.
4. List emails that look invalid.
Return a new CSV plus a short summary of what you changed.
Don't guess missing values.

Its summary might read like this (illustrative):

Service: 371 standardised, 29 marked REVIEW (e.g. "maint + hedges",
"spring tidy").
Likely duplicates: 17 pairs (11 same phone, 6 same name + address).
Invalid emails: 23 (double @, missing domain, spaces).
Name split: 12 records had two people's names; kept the first and
marked REVIEW.

Check its work on a sample before trusting it. In a test like this, expect a few wrong calls: "spring tidy" may deserve its own category rather than REVIEW, and some duplicate pairs will be two different households at the same address, such as flats. The AI does the sorting; you make the decisions. The same approach for a CRM is in cleaning up a messy CRM with AI, and spreadsheet techniques for the fiddly parts are in the tutorial on cleaning messy data.

A one-day clean-up plan for an amber result

Most amber results can be fixed in a working day if you plan it. For the cleaning company's client list above, a realistic schedule:

  • 9.00: save a dated backup copy of the list; agree the final list of service types.
  • 9.30: find-and-replace the service types; add a drop-down so new entries use the list.
  • 10.00: run duplicate detection on phone and email; review each pair and merge by hand, keeping the most recent address.
  • 11.30: filter invalid or missing emails; the office calls or texts those clients over the next week, and the automation skips them until fixed.
  • 13.30: add a "contact consent" column; mark every known "don't contact" request from the inbox and notes.
  • 15.00: rerun the 50-record sample test on a fresh random sample.
  • 15.30: write the owner's name and the monthly spot-check date at the top of the list.

The result for a list of about 600 clients: duplicates down from 8% to under 1%, service types fully standardised, bad emails isolated so the automation ignores them, and a consent flag in place. That is green for the rebooking reminder, which can start the following week.

Reading your results: red, amber or green

  • Green: every key field under 5% problems, no consent or permission issues, one master copy. Start the AI project now, and fix small issues as they appear.
  • Amber: one or two fields between 5% and 15%, or duplicate copies of documents. Spend a day on the fixes above, rerun the sample test, then start.
  • Red: key fields above 15%, data scattered across personal inboxes and notebooks, or any permission or consent problem. Fix the process that creates the data before any AI project; otherwise the AI will automate the mess.

An electrician's firm that scores red on its job records because engineers write notes on paper sheets doesn't need a better AI tool. It needs a simple form on the engineers' phones with four required fields. Once that's been running for a month, the same checklist will come back amber or green, and the AI job that prompted it becomes a small project instead of a rescue. For the full preparation process, step by step, see preparing your business data for AI.

Questions about getting data ready

Do I need to clean all my data before using AI?

No. Clean the data behind the one job you want AI to do first, such as the price list for a chatbot or customer emails for a follow-up automation. Cleaning everything at once takes months, stalls projects and often tidies data no AI will ever read. Once the first job works, repeat the checklist for the next one.

Can AI read my old scanned paperwork?

Usually, but check first. If you can't select or search the text in a PDF, it's a picture of a page, and the AI has to read it with optical character recognition, which makes mistakes on handwriting, stamps and faint copies. Test ten typical documents and check the extracted figures by hand before relying on it for anything with numbers.

How long does a data clean-up usually take a small business?

For one use case, a few hours to a few days. Testing a 50-record sample takes about an hour. Fixing a customer list of a few thousand rows with spreadsheet tools and an AI assistant is often a day's work. The slow part is usually decisions, such as which of two addresses is right, rather than the editing.

Is it safe to paste my customer list into ChatGPT to clean it?

Only on a plan and settings you've checked. Business plans such as ChatGPT Business and Claude Team don't train on your content by default; consumer plans let you switch model training off in privacy settings. Even then, remove columns the task doesn't need, such as bank details or notes about health, and keep the original file.

Further reads

Sources: facts sheet for AI plan privacy defaults; Microsoft Excel and Google Sheets help for duplicate removal, trimming and data validation. Checked September 2026.

Want a second pair of eyes on your data?

On a 1:1 call we'll pick the AI job you care about most, run the sample test on your real records together, and list what to fix first.

Book a 1:1 call with me