RAG, short for retrieval-augmented generation, means the AI first searches your own documents for the passages relevant to a question, then writes its answer from those passages, ideally showing where each point came from. A small business can use it today by loading policies, price lists or manuals into Gemini Notebook or a ChatGPT or Claude project.
The idea matters because an AI model on its own only knows what it learned in training, which doesn't include your returns policy, your prices or last month's change to the staff handbook. RAG hands the model the right pages at the moment of asking, like giving a new employee the handbook open at the right section. It doesn't make the model smarter; it makes the answer depend on your documents instead of the internet's general average, and that's why the quality of those documents decides everything.
Retrieve, then generate: the two steps in plain English
Step one, retrieve. Before anyone asks anything, the system splits your documents into short passages, usually a few paragraphs each, called chunks. It turns each chunk into a numerical fingerprint of its meaning, called an embedding. When a question comes in, the question gets a fingerprint too, and the system pulls out the handful of chunks whose fingerprints are closest. That's why "can a customer send their hearing aids back?" can find a passage headed "Trial and returns" even though the words don't match.
Step two, generate. The retrieved chunks are pasted, behind the scenes, into the prompt alongside your question and an instruction along the lines of "answer using these sources". The model then writes an ordinary-sounding answer, and good tools add citations pointing back to the chunks used.
Almost every "AI trained on your documents" product works this way, from website chatbots that learn your FAQ to the knowledge tools built into office software. Nothing is retrained. Your documents sit in a searchable store, and the model reads a few pages of them each time it answers.
One policy question, traced through a RAG system
Follow one question through the notebook of a two-branch hearing-aid shop (an illustrative set-up) that has loaded 14 documents, about 38,000 words in total, into a notebook: the returns and trial policy, the staff handbook, price lists, warranty terms from two manufacturers and a few old leaflets nobody got round to deleting. A new member of staff asks:
A customer was fitted 40 days ago and wants to return her aids. Can she, and does she get everything back?
What retrieval pulls out. The system finds the three closest chunks:
- Returns and trial policy (updated January 2026), section 2: "Customers may return hearing aids within 45 days of the fitting date for a refund of the device cost."
- The same policy, section 4: "Custom earmoulds and the fitting fee are non-refundable."
- A 2024 in-store leaflet: "Try your new hearing aids for 30 days, risk free."
What generation produces. With a good instruction, the answer reads (illustrative):
Yes. The returns policy allows returns within 45 days of the fitting date, so at 40 days she is within the trial (Returns and trial policy, s.2). She is refunded the device cost, but custom earmoulds and the fitting fee aren't refundable (s.4). Note: a 2024 leaflet mentions a 30-day trial; the January 2026 policy is more recent and should take precedence.
That's RAG working well: the right answer, the sources named, and the conflict flagged rather than hidden. Now the version that goes wrong. Without an instruction to prefer the newest document, the model can just as easily lean on the leaflet and answer "No, the trial is 30 days". Nothing in the system errored; it retrieved a real passage and wrote a fluent answer from it. The shop's fix took ten minutes: delete the old leaflet from the notebook, add "Updated: month year" to the top of every policy, and add one line to the instructions telling the model to prefer the most recent source and mention any conflict.
The example shows the central truth about RAG: most wrong answers are retrieval problems (the wrong passage was found) or document problems (the wrong passage existed), not the model being "stupid".
The vocabulary you'll meet, translated
- Chunk: a slice of a document, often a few paragraphs. If a table is split across two chunks, its headings and figures can end up separated.
- Embedding: the numerical fingerprint of a chunk's meaning, used to find passages that mean the same thing in different words.
- Vector database: the store that holds those fingerprints. Everyday tools hide it from you completely.
- Context window: how much text the model can read at once. Retrieval exists because your documents may not all fit, and because shorter, relevant context tends to produce sharper answers.
- Grounding or citations: the answer is tied to specific sources you can click and check. If a tool can't show you where an answer came from, you can't audit it.
- Token: the unit models count text in. A useful rule of thumb is that 1 million tokens is roughly 750,000 English words.
Everyday RAG you can use this week
You don't need to build anything. Three mainstream tools already do retrieval over files you upload. The limits below are as of September 2026 and change often, so check the vendor's page for your plan.
| Tool | How it uses your files | Limits worth knowing |
|---|---|---|
| Gemini Notebook (formerly NotebookLM) | Answers only from the sources in a notebook, with citations to each passage | Free: 100 notebooks, 50 sources per notebook, 50 chats a day; Google AI Pro raises that to 300 sources and 500 chats. Each source can be up to 500,000 words |
| ChatGPT projects | Files and instructions shared across every chat in the project | Files per project: 5 on Free, 25 on Plus, 40 on Pro, Business and Enterprise; each file up to 512MB |
| Claude projects | Project knowledge read in full while it fits, then searched automatically as it grows | Free allows up to 5 projects; on paid plans Claude switches to retrieval when project knowledge nears the context limit, expanding capacity up to 10 times |
Gemini Notebook is the purest example, because it's designed to answer from your sources and to show you the passage behind each claim. Google says notebook data isn't used to train Gemini Notebook unless you give feedback, and for Workspace accounts it isn't reviewed by people or used for training even when you do. Our tutorial on Gemini Notebook for small business teams covers setting one up for a team.
ChatGPT and Claude projects suit ongoing work: a project for policies, another for a big client, each with its own instructions. OpenAI is retiring custom GPTs (they stop running on 11 December 2026), so projects are the place to put shared business context in ChatGPT now. If the documents include client or patient details, use a business plan (a Workspace account, ChatGPT Business or Claude Team), where your content isn't used for model training by default.
When you don't need retrieval at all
If everything the AI needs fits comfortably in one conversation, you can skip RAG and simply attach the documents. The shop's 38,000 words is roughly 50,000 tokens, well within what current models can read in one go. Attaching the whole returns policy to a question means nothing can be missed by a search step, which is why Claude only switches its projects to retrieval once the knowledge gets large.
Retrieval earns its place when the collection outgrows the context window, when many people ask many different questions of the same library, or when you want every answer tied to a specific passage. A small practice with a 20-page handbook can attach it; a firm with years of job files, manuals and supplier contracts needs retrieval.
Five small-business jobs where RAG earns its keep
- Staff questions about procedures. A pharmacy team asking "what do we do when a delivery arrives damaged?" gets the answer from the current procedure document, with the section cited, instead of interrupting the manager mid-dispensing. The measure of success is new staff no longer asking the same five questions every week.
- Replies to customer emails. An electrician's office loads its price list, booking terms and standard answers. The AI drafts a reply to "do you fit EV chargers and roughly what does it cost?" using the real prices, and the office checks it before sending.
- Quoting from past work. A plumbing firm keeps two years of quotes as PDFs. Asking "find our last three quotes for a full bathroom refit and list what was included and excluded" turns a 20-minute search into a two-minute read, and new quotes stop forgetting the same exclusions.
- Warranty and supplier terms. The hearing-aid shop can ask "which manufacturer covers loss in the first year, and what's the excess?" across both sets of warranty terms at once, instead of opening two long PDFs.
- A website chatbot. Most chatbots "trained on your website" are RAG underneath: they retrieve from your pages and FAQ, then answer. The same document rules apply, which is why an osteopathy clinic's bot is only ever as accurate as the FAQ it reads.
The pattern is the same each time: a question people ask repeatedly, answers that live in documents you control, and a person checking anything that goes out to a customer.
Getting documents ready: a before-and-after
RAG rewards tidy documents. Here's part of a podiatry clinic's staff FAQ as it was, then after a 30-minute tidy.
Before:
Cancellations - we normally ask for a day's notice but Dr P is flexible
with regulars. Nail surgery is different (see Sarah's email from last year,
might have changed). Prices went up in spring I think, check the new sheet.
After:
CANCELLATION POLICY (Updated: September 2026)
Routine appointments: 24 hours' notice. Later cancellations are charged
in full, except for illness with a note.
Nail surgery: 72 hours' notice, because the slot and anaesthetic are
prepared in advance.
Prices: see "Price list September 2026" only. Older price lists are archived.
The "before" version will produce vague or contradictory answers however good the AI is, because it points to an email nobody loaded and a price sheet nobody named. The rules that make the "after" version work: one topic per heading, a date on every policy, numbers stated rather than referred to, and old versions removed rather than left alongside new ones. Building a company knowledge base AI can answer from goes further on structuring a whole library.
Four ways RAG answers go wrong, and how each shows up
- Old versions survive. The trial-period mix-up above. It shows up as answers that were true last year. Archive old documents outside the notebook or project, and see how to catch outdated information in AI answers.
- Scans with no text. A photographed supplier agreement may contain no searchable text at all, so retrieval never finds it and the AI answers as if the clause doesn't exist. It shows up as "the documents don't mention that" for something you know is there. Run scans through text recognition first, or re-save them from the original file.
- Tables broken apart. A price table split across chunks can lose its column headings, so a price gets attached to the wrong product. It shows up as figures that exist in the document but in the wrong row. Keep price tables short, repeat the headings, or put each product on its own line with its price.
- Gaps filled from general knowledge. Ask a question your documents don't answer and many tools quietly answer from the model's training instead. It shows up as confident answers with no citation. Add an instruction such as: "Answer only from the sources. If they don't cover it, say 'not in our documents' and suggest who to ask."
A 20-question test to prove it's working
Before staff rely on it, write 20 questions whose answers you already know, including a few awkward ones: a question the documents don't cover, one where an old and a new version disagree, and one that needs figures from a table. Score each answer. Here are five rows from the podiatry clinic's sheet (results illustrative):
| Question | Expected answer and source | Result |
|---|---|---|
| How much notice for cancelling nail surgery? | 72 hours, cancellation policy | Correct, right source |
| What does a biomechanical assessment cost? | From the September 2026 price list | Correct figure, but cited the archived list: fixed by removing it |
| Can we treat a patient who is on blood thinners? | Not in the documents; ask a clinician | Answered from general knowledge: instruction added |
| Do we open on public holidays? | Opening hours sheet: closed | Correct |
| What's the fee for a home visit in the evening? | Price list, "home visits" row | Wrong row: price table rebuilt with one item per line |
Aim for at least 18 of 20 correct with the right source before anyone relies on it, and rerun the same 20 whenever you add or replace documents. The answers that fail tell you exactly which document or instruction to fix.
RAG, fine-tuning and custom GPTs: which is which
People use these terms interchangeably, but they're different tools. RAG looks things up at the moment of asking, so updating it is as simple as replacing a document. Fine-tuning changes the model itself by training it on examples, which suits a style or a format, not facts that change monthly. Custom GPTs were ChatGPT's packaged assistants with their own instructions and files, and OpenAI is retiring them, turning them into plugins whose knowledge files become reference files. For facts about your own business (policies, prices, procedures), RAG is almost always the right answer. RAG vs custom GPT vs fine-tuning compares the three in detail, and what a chat-with-your-documents system costs prices the step up from everyday tools to a built system.
For most small businesses, the path is short: tidy five to ten key documents, load them into a notebook or project, write the "answer only from the sources" instruction, and run the 20-question test. If it passes, you have working RAG for the price of a subscription, or for nothing on a free plan.
Further reads
- AI Knowledge Base Options for Small Businesses Compared — Compares the knowledge-base tools that use RAG behind the scenes.
- How to Train ChatGPT on Your Business Information Safely — How to give ChatGPT your business information safely.
- How to Train an AI Chatbot on Your FAQs, Policies, and Prices — Using the same ideas to power a customer chatbot.
- How to Use ChatGPT Projects to Keep Client Work Separate — Keeping each client's documents in its own project.
- Are Chat-With-PDF Tools Safe for Contracts and Client Files? — Whether it's safe to upload contracts and client files at all.
- How to Build the FAQ Your AI Chatbot Needs Before Launch — Writing the FAQ document a RAG system answers from best.
- What Is an AI Agent? A Plain-English Guide for Business Owners — What an AI agent actually is, one task followed step by step, the agents small businesses can switch on today with costs, and how to give one its first job.
- Best Free AI Tools for Small Business Owners and Their Catches — Twelve genuinely useful free AI tools for small businesses, each with a worked example, the catch that matters and the setting to change on day one.
- AI Tools and AI Development: The Complete 2026 Guide — the AI hub, including every tutorial in the AI-for-business series.
Sources: Google's Gemini Notebook help pages on plan limits and privacy; OpenAI help centre pages on Projects and file uploads; Anthropic's help pages on RAG for projects and the Claude pricing page; OpenAI's custom GPT retirement FAQ (checked September 2026).