Start with records of what happens at each step of a job: every enquiry and where it came from, every quote with its outcome and the reason, actual time and cost against the estimate, customer questions and complaints, and what you charged. Capture the same fields every time, dated, in one system.
Twelve months of consistent records beats five years of scattered notes, and you need less volume than you'd think: most small-business AI uses work from hundreds of records, not millions. The field most firms forget is the reason: why a quote was lost, why a job overran. Without it, AI can tell you what happened but never why, and the why is usually where the money is.
What AI does with a small firm's data
Almost every practical AI use in a business your size falls into one of three kinds, and each needs a different sort of data:
- Answering from your knowledge. A chatbot or internal assistant replying with your prices, policies and processes. It needs text: questions customers actually ask and your best answers.
- Spotting patterns and forecasting. Which enquiries turn into jobs, which jobs overrun, what next quarter looks like. It needs rows: one line per event, with the same columns every time.
- Drafting in your style. Quotes, follow-ups and replies that sound like you. It needs examples: a folder of your best real ones.
You don't need a data warehouse or a data scientist for any of these. If you're unsure whether you have enough to start at all, do you need a lot of data to use AI? answers that directly. This tutorial is about what to start recording now so that in six to twelve months you have something worth analysing.
How many records you'll produce decides which of the three kinds to aim at. A furniture maker who builds 25 commissioned pieces a year will never have enough rows for pattern-spotting to beat the owner's own memory; after two years, 50 rows is still anecdote. For a business like that, the valuable collection is text and examples: every enquiry email and the reply that won the commission, the questions clients ask about timber and finishes, and the ten best quotes with their drawings. That's what an assistant needs to draft replies and quotes in the maker's voice. A takeaway handling 400 orders a week is the opposite case, with rows piling up fast enough that a month of consistent fields is already worth analysing.
The seven datasets worth starting now
1. Enquiry log
Fields: date, channel (phone, web form, email, social), how they heard about you, service wanted, job size, response time, outcome (quoted, not quoted, lost contact).
What it lets AI do later: tell you which marketing sources produce jobs rather than just enquiries, and which enquiries you're too slow to answer. Start this week: add a "how did you hear about us?" drop-down to your web form and a line in the phone script.
2. Quote register
Fields: job ID, date sent, job type, size measure (rooms, square metres, volume), price quoted, days from enquiry to quote, outcome (won, lost, no reply), reason lost, competitor price if the customer volunteered it.
What it lets AI do later: show your real win rate by job type and price band, and check new quotes against similar past ones. "Reason lost" is the field almost nobody records and the one that tells you the most. Start this week: a drop-down with six reasons (price, timing, went elsewhere, no reply, changed plans, other) filled in whenever a quote closes.
The difference between a note and a record is easiest to see side by side. A painting and decorating firm's quote, as it usually gets written down:
3-bed house, lounge/hall/landing, about 2.2k inc ceilings, she's
thinking about it, maybe went with the other lot?
The same quote as a register row (illustrative):
job_id date_sent job_type size_m2 price_quoted days_to_quote outcome reason_lost competitor_price
J1047 2026-09-03 interior 62 2200 4 lost went elsewhere 1850
The note holds nearly the same information, but nothing can count it. Forty rows like the second one will show whether interior jobs over 50 square metres are losing on price, and by roughly how much. Forty notes like the first will show nothing without someone retyping them.
3. Job actuals
Fields: job ID, estimated hours, actual hours, crew size, vehicles, materials used, extras added on the day, and a complication code (access, weather, customer not ready, scope change).
What it lets AI do later: find which job types you consistently under-quote, which is often the single most valuable pattern in a service business. Start this week: two extra fields on the job sheet, "actual finish time" and "what slowed us down".
Watch the complication code's "other" option. An electrical contractor that added the field might find, after two months, that 38 of its 60 overrunning jobs were coded "other", which makes the column almost useless. Reading the free-text notes on those 38 usually shows why: here, 22 said some version of "waiting for the plasterer" or "kitchen fitter not finished", and 9 said "part not on the van". Neither was on the list. Adding "waiting for another trade" and "part not carried" as codes cut "other" to a handful, and the second code turned out to point at a van-stock problem worth more than any AI project. Check the "other" share after the first month; if it's above about one in five, the list is missing something.
4. Customer questions and your best answers
Fields: the question in the customer's own words, the channel, your answer, and whether it needed a person.
What it lets AI do later: power a website chatbot or an assistant that drafts replies, and feed a company knowledge base AI can answer from. Start this week: a shared document where whoever answers the phone adds any question asked twice.
In a ten-room guest house, the first entries might read:
| Question (customer's words) | Channel | Our best answer | Needed a person? |
|---|---|---|---|
| "Can we bring our dog?" | Web chat | Yes, in rooms 1 and 2 (ground floor, garden door), $15 a night, one dog per room | No |
| "What time can we check in if our train's early?" | Rooms from 3pm; bags can be left from 10am | No | |
| "Is the breakfast room step-free? My mum uses a walker." | Phone | One small step at the door; we put a ramp out on request | Yes: manager checked the ramp is available |
Keep the customer's wording exactly, "my mum uses a walker" and all. A future chatbot needs to recognise questions the way people actually ask them, and the third column is what it will answer from, so write it as you'd want it sent.
5. Complaints, claims and fixes
Fields: job ID, date, what went wrong, cause category, cost to put right, days to resolve.
What it lets AI do later: spot repeat causes (the same crew, the same job type, the same supplier) before they become a pattern customers notice. Start this week: a single tab in your spreadsheet; the discipline is logging small complaints, not only the big ones.
6. Price history and discounts
Fields: date of each price-list change, old and new price, every discount given, how much, and why.
What it lets AI do later: connect price changes to win rates, and show how much discounting costs you a year. Start this week: save each price list with its effective date instead of overwriting the last one.
Discounts are where a few weeks of logging tends to surprise owners. Suppose a four-chair hair salon records every discount for a quarter, with a reason chosen from a short list. The log shows 94 discounts averaging $6: $564 in the quarter, or roughly $2,250 a year if the pattern holds. The reason column holds the part nobody knew. Most were "regular client", and nearly all came from one stylist, who had been rounding bills down for years out of kindness. That isn't necessarily wrong, but it's now a decision the owner can make on purpose, perhaps by turning it into a loyalty reward that every client gets. None of it was visible in the till totals.
7. Feedback and what happened next
Fields: job ID, review or feedback text, score if any, and whether the customer came back or referred someone.
What it lets AI do later: summarise what customers praise and complain about, and link it to repeat business. Start this week: copy each review into the log with the job ID, including the lukewarm ones.
The one field that makes the rest usable: a job ID
Seven separate lists are only mildly useful. Seven lists that share a job ID are a history of every job from first call to final review, and that's what lets AI answer questions like "which enquiry sources produce jobs that overrun?" Give every enquiry an ID the moment it arrives (J1001, J1002 and so on) and carry it onto the quote, the job sheet, the invoice and any complaint.
A few formatting rules, agreed now, will save weeks of cleaning later:
| Rule | Why it matters |
|---|---|
| Dates as YYYY-MM-DD in their own column | Sorts correctly and never gets confused between date formats |
| Drop-downs for categories, ten options at most | "Price", "too expensive" and "$$$" become one value instead of three |
| Money as plain numbers, no text in the cell | "About 1,200 inc. extras" can't be added up by anything |
| One row per event, never merged cells | AI tools and spreadsheets read rows; merged cells break both |
| A notes column for everything else | Stops people squeezing comments into the structured fields |
Where to keep it
Use the system you already have if it can hold these fields: job-management software, a CRM (HubSpot's free CRM handles enquiries and quotes), or your accounting software for prices and invoices. If none of them fits, a spreadsheet with one tab per dataset is perfectly good for the first year. Airtable vs Google Sheets for data AI can use compares the two most common choices.
The rule that matters more than the tool: one home per dataset. Quotes in three places (email, a notebook and a spreadsheet) are worse than quotes in one imperfect place.
What not to collect
Collecting more feels like future-proofing. It mostly creates risk. Data-protection rules such as the GDPR expect personal data to be "adequate, relevant and limited to what is necessary" for the purpose, and kept no longer than necessary. In practice:
- Record the need, not the reason. If a customer mentions a health condition, note "ground-floor access needed" or "extra time for packing", not the diagnosis. So a removals surveyor's note that reads "recovering from hip surgery, can't lift anything" becomes "customer can't lift: crew packs and carries all boxes". The crew gets what it needs to plan the job, and the spreadsheet holds no health data.
- Don't copy ID documents or bank details into spreadsheets or AI tools. Keep them in the system designed for payments, if you need them at all.
- Don't record calls without telling callers, and don't keep recordings indefinitely "in case AI can use them one day".
- Set a keep-for period per dataset, such as enquiries that never became jobs deleted after two years. Ask your data-protection adviser what's reasonable for your records.
- Keep customer names out of the analysis copy. For pattern-spotting, the job ID is enough; names add risk and no insight.
Worked example: a removals firm's first 90 days
An illustration: an 11-person removals firm with three vans. Before: enquiries arrived by phone and web form, surveys were done on paper or by video call, quotes lived in the sent-items folder, and nobody recorded why quotes were lost. The owner suspected three-bedroom house moves were under-quoted but couldn't prove it.
Setup took about six hours: a spreadsheet with four tabs (enquiries, quotes, jobs, complaints) keyed by job ID, a phone form for the surveyor that captured rooms, estimated volume and access notes, and drop-downs for source, reason lost and complication code. The owner already used a video survey process for quotes, so the volume estimate simply became a field instead of a note.
After 90 days the log held 240 enquiries, 150 quotes and 70 completed jobs. Pasted into an AI assistant with names removed, three findings came out, each checked by hand in the spreadsheet:
- Three-bedroom house moves ran on average 1.4 hours over the estimate, mostly coded "customer not packed". The firm added a packing-readiness question to the survey and a surcharge line to the quote.
- Web-form enquiries answered within two hours were won at roughly twice the rate of those answered the next day.
- "Price" was the stated reason for only a quarter of lost quotes; "no reply" was the largest group, which pointed to follow-up, not pricing.
None of this needed advanced AI. It needed three months of the same fields filled in every time.
A 30-day capture plan
- Week 1: pick the three datasets closest to the AI use you care about most. For most service firms that's enquiries, quotes and job actuals. Write down the fields.
- Week 2: build the forms and drop-downs, start issuing job IDs, and brief the team in ten minutes on why "reason lost" matters.
- Week 3: capture live. Spend five minutes each afternoon checking the day's rows for gaps.
- Week 4: check quality (below), fix any field people keep skipping, then add the next dataset.
How to tell the data will be usable
After a month, three checks. First, completeness: at least nine in ten rows should have the outcome fields filled in. Second, consistency: each category column should contain only the drop-down values. Third, ask an AI tool to describe it, using a prompt like this on an export with names removed:
Attached is a spreadsheet export of our [enquiries / quotes / jobs] for
the last month. Do not draw conclusions yet. Instead:
1. List each column and what it seems to contain.
2. Flag columns with blanks, inconsistent values or text in number fields,
with the row numbers.
3. List three questions this data could answer once we have six months
of it, and any extra field that would make those answers more reliable.
For the removals firm's first month, a reply might include lines like these (illustrative):
2. Issues found
- source: 14 different values, including "google", "Google",
"Google ads", "internet" and "web". Rows 4, 9, 17, 22...
- quoted_price: text in rows 12, 19 and 33 ("tbc", "approx 900",
"see email")
- reason_lost: blank in 11 of 19 lost quotes
3. Questions for six months' time
- Which sources produce the highest-value moves?
- Does quote speed affect win rate by job size?
- Adding customer age would help predict which quotes are won.
Each issue points at a fix. Fourteen spellings of "source" means that column was typed, not picked from a list, so it needs a drop-down. Text in a price field means the quote went out before the price was final; add a "draft" status instead. Blank reasons in more than half the lost quotes means nobody is asking, which is a habit to fix in a team briefing. And ignore the last suggestion. Customer age adds personal data for a guess at a pattern, which is the opposite of collecting only what you need.
If the reply is mostly about blanks and inconsistencies, fix the capture process before collecting more. If you're also sitting on years of older, messier records, cleaning up customer records before you add AI covers that separate job. Once you have six months of clean rows, forecasting next quarter's sales from your history becomes a realistic next step.
Further reads
- How to Prepare Your Business Data for AI, Step by Step — What to do with the records you already have.
- How to Classify Business Data Before Using AI Tools — Label what's safe to put into which AI tool.
- How to Keep Customer Data Private When Your Team Uses AI — Rules for staff once the data exists.
- Is Your Business Data Ready for AI? A Clean-Up Checklist — A clean-up checklist to run after three months.
- How to Build a KPI Dashboard With AI When You Have No Data Team — Turn the new records into a simple dashboard.
- How to Build the FAQ Your AI Chatbot Needs Before Launch — Put your customer-questions log to work.
- Reduce Owner Dependency: Use AI to Capture What Only You Know — Capture the know-how that never makes it into records.
- AI Tools and AI Development: The Complete 2026 Guide — the AI hub, including every tutorial in the AI-for-business series.
Sources: GDPR Article 5 (data minimisation and storage limitation principles). Business figures are illustrative.