No. Most small-business AI uses, such as drafting, summarising, sorting enquiries and answering from your own documents, run on models the vendor has already trained, so you need a handful of good examples and the right documents, not a dataset. Volume only matters for prediction, such as forecasting or scoring, where you want two or more years of consistent records.
What you do need is data that's findable and current for the one job you're starting with. The table below gives rough amounts for common tasks, followed by the cases where volume genuinely matters and a 30-minute inventory of what you already hold.
Three different things people mean by "data"
Most of the confusion comes from one word covering three separate things.
- Training data. The vast amount of text the vendor used to build the model. It isn't yours, you can't add to it in any practical sense, and you don't need to. This is why a new assistant can write a decent letter on day one.
- Context. What you hand the model for a particular task: your instructions, a few examples of good work, the document it's working from. The amounts are small and quality matters far more than quantity.
- Records. The history in your systems: sales, jobs, enquiries, invoices. Records matter when you want AI to predict something from patterns, and when you want to measure whether AI is helping.
When someone says "we don't have the data for AI", they're usually thinking of records. For the tasks most firms start with, the thing that's needed is context.
How much each common task needs
These amounts are rules of thumb for getting reliable results, not vendor requirements. The right-hand column is where most projects actually succeed or fail.
| Task | What the AI needs from you | Rough amount | What matters more than volume |
|---|---|---|---|
| Drafting emails or letters in your style | Examples of good ones, plus the facts for this one | 3-5 examples | Examples are recent and typical, not your one best-ever letter |
| Summarising calls or meetings | The transcript or notes | One at a time | Recording quality and knowing who said what |
| Answering questions from your policies or product information | The documents that hold the answers | Often 10-40 documents | Only current versions; one source of truth per topic |
| Sorting enquiries or emails into categories | Clear category definitions, plus past examples to test against | 20-50 labelled examples for testing | Rules for the awkward cases that fit two categories |
| Pulling fields out of documents (invoices, forms) | Sample documents | 10-20 of each layout for testing | Scan quality and consistent layouts |
| Forecasting demand, sales or cash | Consistent history | At least two full seasonal cycles, so 24+ months if monthly | The same definitions used throughout the period |
| Scoring leads or predicting which clients will leave | Outcomes, including the failures | Hundreds of closed outcomes, wins and losses | Whether lost deals were ever recorded at all |
Notice that the first five rows involve tens of items at most, and in several of those the "data" is only needed to test the setup, not to make it work.
The "typical, not best-ever" note in the first row catches people out. Say the owner of a small removals firm gives the assistant one example: a long, warm letter written to a care home after a difficult three-day move, the letter everyone was proudest of. Every draft that follows comes out at 350 words, thanks customers for "trusting us with your precious memories" and offers "our heartfelt support" to someone who asked for a quote on a two-bedroom flat. Swap in four ordinary quote emails from last month, each about 120 words, and the drafts settle into the firm's real voice. The volume barely changed; the choice of examples did.
The sorting row works the same way: the labelled examples are there to test your categories, and the awkward ones matter most. An illustrative joinery workshop testing enquiry sorting wrote its first few rows like this:
Email (shortened) Label
"Can you quote for fitted wardrobes in two bedrooms?" New quote
"The oak door you fitted in May is sticking" Aftercare
"Price for 40 metres of skirting, supply only?" Trade order
"Loved the staircase! Could you also do the landing?" ??? (aftercare or new quote)
The fourth row is the one that teaches you something. It's praise, a past job and a new piece of work in one email, and a person would treat it as a new quote with a friendly opening. So the workshop wrote the rule down ("any request for new work is a new quote, whatever else the email says") and made sure its 30 test emails included three more like it. Collecting more ordinary emails wouldn't have found that gap; one awkward one did.
For document extraction, scan quality beats volume. An illustrative café sent phone photos of supplier invoices, taken at an angle on the counter, to an AI extraction step. One total of $1,850.00 came through as $185.00 because a fold hid a digit, and nobody noticed until the month's food costs looked impossibly low. Ten more sample invoices wouldn't have helped. Flat, well-lit scans would: Microsoft Lens has been retired, so use the scan feature in the OneDrive mobile app or Google Drive's scanner, and add a check that flags any invoice total more than 50% below that supplier's usual amount.
When volume really does matter
Prediction is different in kind. A forecasting or scoring model learns from patterns in your own history, so the history has to contain the pattern. Two conditions decide whether you have enough.
Enough cycles. If your business is seasonal, a forecast needs to have seen each season at least twice to tell a regular December peak from a one-off. With 14 months of data, it has seen one December; it can't tell the difference.
Both sides of the outcome. A model that predicts which leads convert needs to see leads that didn't. Many small firms only record the wins, because the losses just stop getting replies. A CRM with 400 won deals and no lost ones can't be scored meaningfully.
If you fall short on either, you still have options. Simple methods work surprisingly well on small histories: a spreadsheet trend line, or Excel's Forecast Sheet on the Data tab, which fits a standard exponential-smoothing forecast to whatever history you give it and shows the uncertainty range. Pair that with your own judgement, and start recording what's missing now. What business data to start collecting for AI lists the fields worth capturing from this month onwards.
Chat assistants will happily forecast from too little. Give one 14 months of a garden centre's monthly sales and ask "Forecast next December", and an illustrative answer reads:
Based on your data, I forecast December sales of about $48,200, around 12% up on last December, reflecting the upward trend across the year and your strong seasonal peak.
It sounds analytical, but it rests on a single December, and the "12% up" is the year's general trend stretched over a season the data has seen only once. A more useful question is: "What can and can't you conclude about December from one year of data? Give me a low, middle and high figure and say what each assumes." That turns a false point estimate into a range the owner can plan stock around, and it's honest about which part is a guess.
Whatever method you use, test it on months you already know. Take an illustrative bike shop with 30 months of sales. Hide the last three, build the forecast from the other 27, and compare it with what actually happened:
Month Forecast Actual Difference
Jul $22,400 $24,100 -7%
Aug $21,900 $20,300 +8%
Sep $17,600 $16,900 +4%
Three months: forecast $61,900, actual $61,300, about 1% over
Individual months were out by up to 8%, while the three-month total was close. So this forecast is good enough for ordering stock by the quarter and not for rota planning week by week. That's a more useful conclusion than any accuracy claim, and it took ten minutes with data the shop already had.
The real problem is usually scattered data, not scarce data
In most small firms, the information exists; it's just in four places with three different dates on it. That breaks AI tools faster than a shortage ever does, because an assistant answering from an outdated document gives an out-of-date answer with the same confidence as a correct one.
Typical patterns:
- A marketing agency has a client's brand guidelines in four versions across shared drives, and the assistant picks up the 2023 one with the old colour palette.
- A web design studio's pricing lives partly in a spreadsheet, partly in old proposals, and partly in a founder's head.
- A mortgage adviser's case notes are split between the CRM, email threads and handwritten notes from client meetings.
The fix is modest: one current folder for the job, old versions archived out of it, and a named person who updates it. Organising shared files so AI tools can use them covers folder structure and naming, and preparing your business data step by step covers the wider tidy-up if you need it.
Three AI ideas at a two-adviser mortgage brokerage
This is illustrative. Say a two-adviser mortgage brokerage with one administrator has three ideas for AI and assumes it lacks the data for all of them.
Idea 1: draft the progress-update emails clients get during an application. What it needs: five past update emails the advisers were happy with, with client details removed, plus the current case note each time. The brokerage has hundreds of past emails. Ready to start the same week, with an adviser reviewing every draft before it goes.
Idea 2: an internal assistant that answers "which of our usual lenders accept this situation?" from lender criteria documents. What it needs: the current criteria for each lender the firm regularly uses. The brokerage has 14 lender documents saved, of mixed ages; four are more than a year old. Criteria change often, so the real work is a monthly routine to replace old versions, not collecting more. Ready after about three hours of clearing out and re-downloading, plus 30 minutes a month to keep it current. The adviser still checks the lender's own current criteria before recommending anything.
Idea 3: predict which past clients are due to remortgage. This one sounds like it needs a prediction model. It doesn't: it needs one field, the date each client's current deal ends. The firm has around 600 completed cases over six years, but that date is recorded for only about 60% of them. It's a completeness problem solved with a filter, plus a one-off effort to fill the gaps from old files. Using AI to spot remortgage opportunities in your client bank goes further into that workflow.
The lesson: none of the three ideas was blocked by a lack of data. One was blocked by out-of-date documents and one by an empty field.
The same question in two other kinds of firm
A marketing agency writing social posts for 20 clients. Each client needs its own voice, so the "data" is per client: the brand guide and five approved posts. That's about 100 examples in total, which sounds like a lot until you notice they already exist in the agency's approval history. The work is filing them into one folder per client, not creating anything. The agency should also check each client contract for terms about AI use before loading that client's material.
A web design studio that wants help estimating project hours. It has tracked hours on about 40 past projects. That's far too few to train a prediction model on, and it doesn't need to. An assistant can read a new brief alongside a summary table of the 40 projects, pick out the three most similar, and draft an estimate range with the reasoning shown. A founder then adjusts it. The history works as reference material rather than training data, which is a much lower bar.
When a vendor says its tool needs your data
Some software sells prediction features: lead scores, churn warnings, demand forecasts. Those claims about data are real, and worth testing before you pay. Ask:
- What is the minimum history the feature needs, in months and in records, and what does it do if you have less?
- Can it show its accuracy on your own past data before you commit, rather than a demo account's?
- Does it need outcomes you don't currently record, such as lost deals or reasons for cancellation?
- Is your data used to improve the vendor's models for other customers, and can you switch that off?
A vendor that can't answer the first two clearly is selling a feature that may simply switch itself off for a firm your size.
A 30-minute inventory of what you already have
Copy this into a document and fill it in for the one job you want AI to help with first. If you can complete every line except the last two, you have enough to start.
Job I want AI to help with:
What a good result looks like (paste one example):
Past examples of good work I can share (how many, where stored):
Reference documents it must answer from (list each, with date last updated):
Anything confidential in those that must be removed first:
Who keeps these examples and documents up to date:
Records needed? (only if predicting: how many months, same definitions throughout?)
Outcomes recorded both ways? (only if scoring: wins AND losses?)
Filled in by the removals firm from earlier, for its first job, it might read (illustrative):
Job I want AI to help with: follow-up emails for quotes not yet accepted
What a good result looks like: [last Tuesday's 110-word follow-up]
Past examples of good work: 4 ordinary follow-ups from last month,
in the "Quotes 2026" folder
Reference documents: price list (updated March), storage terms
(updated January), cancellation policy (updated last year: CHECK)
Confidential, remove first: customer names, addresses, key-safe codes
Who keeps these current: office manager
Records needed? No, nothing is being predicted
Outcomes recorded both ways? Not needed for this job
Every line except the last two is filled in, so the job can start this week. The one flag, an out-of-date cancellation policy, is exactly the kind of thing that would otherwise surface as a wrong promise in a customer's inbox.
Once it's filled in, giving AI your business context shows how to turn those examples and documents into instructions the assistant will actually follow. If the documents list is long, building a knowledge base AI can answer from is the next step.
Start now, tidy first, or record and wait
- Drafting, summarising or sorting: start now. Three to five good examples and a clear brief are enough.
- Answering from your documents: start once each topic has a single current version and a person who keeps it current.
- Predicting or scoring: check for two full seasonal cycles or several hundred outcomes including failures. If you don't have them, start recording now and use simple spreadsheet methods in the meantime.
Follow-up questions about data and AI
Can I train ChatGPT on my own business information?
Not in the sense most people mean. You don't retrain the model; you give it context. ChatGPT and Claude both offer Projects, where you attach reference files and standing instructions that every chat in the project can use. OpenAI is retiring custom GPTs, which stop running on 11 December 2026, so a Project is the place to set this up. Fine-tuning a model on your own data is a developer feature that is rarely worth the cost or effort for a small firm's everyday work.
Do I have to clean all my business data before using AI?
No. Clean the data for the one job you are starting with. If you want drafts of client updates, that means tidy case notes and a few good example emails. A company-wide data clean-up before you begin is a common way to spend months without starting. Widen the tidy-up only as you add workflows.
Is it safe to upload client documents so the AI has context?
Use a business plan that doesn't train on your content by default, upload only the documents the task needs, and remove details the task doesn't require. Check your client contracts and professional rules first, since some forbid third-party processing without consent. If personal data is involved and you are unsure of your duties, ask your data-protection adviser.
Further reads
- How to Classify Business Data Before Using AI Tools — Sort what can and can't go into AI tools before you upload.
- How to Clean Up Customer Records Before You Add AI — The tidy-up to do when records really are the problem.
- How to Forecast Next Quarter's Sales With AI Using Your History — What forecasting with your own history looks like in practice.
- Is It Safe to Put Customer Data Into ChatGPT? — The privacy side of handing AI your documents.
- Generative AI vs Traditional AI: Which Does Each Task Need? — Why prediction tools need records and drafting tools don't.
- AI Myths That Stop Small Businesses Getting Started — The other beliefs that stop firms starting.
- AI for Small Business Owners: A Plain-English Beginner's Guide — What an owner should know before starting with AI: how chat tools really work, the four kinds of product, data rules, costs and a first fortnight.
- Predictive Maintenance With AI: Realistic for a Small Workshop? — When AI predictive maintenance pays in a small workshop, the five cheaper rungs below it, and a CNC shop costed end to end.
- Is Your Business Data Ready for AI? A Clean-Up Checklist — A 50-record sample test and a five-part clean-up checklist, each item with why it matters and how to check it, plus a record cleaned before and after.
- AI Tools and AI Development: The Complete 2026 Guide — the AI hub, including every tutorial in the AI-for-business series.
Sources: OpenAI and Anthropic help pages on Projects; Microsoft support page on Excel Forecast Sheet (checked September 2026).