RAG vs Custom GPT vs Fine-Tuning: What Does Your Business Need?

Coding Liquids tutorial cover featuring Sagnik Bhattacharya for RAG vs Custom GPT vs Fine-Tuning: What Does Your Business Need?
Coding Liquids tutorial cover featuring Sagnik Bhattacharya for RAG vs Custom GPT vs Fine-Tuning: What Does Your Business Need?

Almost certainly RAG: an assistant that looks up your own documents before answering, which a ChatGPT Project, Claude Project or Gemini Gem (becoming a skill from November 2026) with uploaded files already provides. Don't start a custom GPT, because OpenAI retires them on 11 December 2026, and skip fine-tuning unless you need a fixed output format at very high volume.

The confusion comes from treating these as three rival products. They're three different levers. RAG changes what the AI knows at the moment it answers. A custom GPT was only ever a package of instructions plus files, which is itself a small RAG set-up with a name on it. Fine-tuning changes how a model behaves by training it on examples, and it has become harder to get: OpenAI stopped new organisations creating fine-tuning jobs on 7 May 2026 and ends new jobs for everyone on 6 January 2027.

Follow me on Instagram@sagnikteaches

Three levers: knowledge, behaviour and packaging

Before comparing costs, be clear about what each one changes. Most bad decisions here come from using a behaviour tool to fix a knowledge problem.

Connect on LinkedInSagnik Bhattacharya
ApproachWhat it changesWhat it doesn't changeEffort for a small firmStatus in September 2026
RAG (retrieval-augmented generation)The facts in front of the model when it answersTone, format or reasoning styleHours, using a project; weeks, if custom-builtBuilt into ChatGPT Projects, Claude Projects, Gemini Notebook and most chatbot platforms
Custom GPTNothing new: bundles instructions, files and optional actionsThe underlying modelAn hour or twoRetiring: stops running 11 December 2026; creation has closed
Instructions plus examples ("few-shot")Format, tone and decision rules, per requestWhat the model knowsAn afternoon of writing and testingWorks everywhere
Fine-tuningThe model's habits, learned from hundreds or thousands of examplesFacts that change; it doesn't reliably store themWeeks, plus a developer and a training setOpenAI self-serve closed to new organisations; limited elsewhere

Notice the fourth row that the title leaves out. Standing instructions with two or three worked examples solve most of the problems people bring to fine-tuning, at no extra cost, and you can change them in a minute.

Subscribe on YouTube@codingliquids

What happened to custom GPTs, and what replaces them

OpenAI is retiring custom GPTs across its ChatGPT plans. They stop running on 11 December 2026; Enterprise customers with an approved deferral get until 11 February 2027. Building new ones has already ended. OpenAI's migration path turns a GPT into a plugin: the GPT's instructions become a skill and its knowledge files become reference files.

For a small business the practical replacement is simpler. If your GPT held a policy manual and some instructions, recreate it as a shared ChatGPT Project on a Business plan, or a Claude Project on Claude Team, or a Gemini Gem if you're on Google Workspace. All three do the same job the GPT did: fixed instructions, a set of files, shared with named colleagues. The detailed trade-offs are in custom GPT vs Claude Project vs Gemini Gem. What you lose is the GPT Store listing and, for some GPTs, custom actions that called outside systems; those now belong in apps and plugins, which are covered in connecting ChatGPT or Claude to your business apps.

So the real comparison in the title is now RAG versus fine-tuning, with the custom GPT reduced to "a RAG set-up you need to move before December".

RAG in plain terms: one patient question traced through

Retrieval-augmented generation means the AI searches a store of your documents for the passages relevant to a question, then writes its answer using those passages. The model itself learns nothing; it reads, like a new receptionist glancing at the leaflet before replying. A plain-English explainer lives in what is RAG; here's what it looks like in one practice.

A seven-person dental practice loads its patient leaflets, fee list and aftercare sheets into a Claude project. A receptionist types a question a patient has just emailed: "I had a filling this morning. When can I eat?"

  1. Search. The tool breaks the practice's files into passages and finds the ones that best match the question: "Aftercare after fillings", section 2, and a line in the general FAQ.
  2. Assemble. Those passages are placed in front of the model along with the question and the project's standing instructions ("answer in plain English, name the leaflet you used").
  3. Answer. The model writes a reply based on those passages and cites them.
  4. Update. When the practice changes its aftercare advice, someone replaces the leaflet. The next answer uses the new text. No retraining, no developer.

That last step is the whole argument for RAG in a small business. Prices, opening hours, policies and team members change every few months, and a system where you fix a fact by editing one file is the only kind a busy owner will keep accurate.

Projects also handle growth for you. On paid Claude plans, a project whose files outgrow Claude's working memory switches automatically to retrieval, searching the files rather than reading them all each time; Anthropic says this lets a project hold up to ten times more content. You don't configure anything; you just keep adding documents.

RAG does have its own failure: the search step can pick the wrong passage. In the practice's first week, "How much is a hygiene appointment?" kept returning the price from a 2025 leaflet that was still in the project. The model did exactly what it was told; it was handed stale text. Retrieval is only as current as the files, which is why one owner per document matters more than which tool you choose.

Fine-tuning in 2026: what it's for and who still offers it

Fine-tuning takes an existing model and trains it further on your examples: hundreds or thousands of "input, ideal output" pairs. The model picks up habits, such as always replying in a particular structure, consistently sorting messages into your own categories, or matching a house style closely. It's the right tool when the problem is consistency of behaviour at scale and prompting has genuinely run out of road.

It's the wrong tool for teaching facts. A model fine-tuned on this year's fee list will still produce this year's fees next year, mixed unpredictably with what it learned elsewhere, and the only fix is another training run. That's the single most common reason small firms regret it.

Availability has also narrowed:

  • OpenAI has wound down its self-serve fine-tuning platform in stages. From 7 May 2026, organisations that had never fine-tuned couldn't start. From 2 July 2026, organisations that hadn't used a fine-tuned model in the previous 60 days lost the ability to create new jobs. On 6 January 2027, active customers lose it too. Existing fine-tuned models keep working until their base model is deprecated.
  • Google offers supervised tuning of Gemini models on its Gemini Enterprise Agent Platform (the product formerly called Vertex AI). At the time of writing, its documentation lists selected Gemini 2.5 models rather than the newest flagship, so check which models are supported before planning around it.
  • Anthropic doesn't offer fine-tuning through its own API. The managed route has been through Amazon Bedrock for an older, smaller Claude model; check whether it's still offered for the model you'd want.
  • Open-weight models (ones you can download and run yourself or through a cloud host) can be fine-tuned freely, but that puts you in developer territory: training data, hosting, monitoring and updates are all yours.

For a business under 50 people, if fine-tuning is being proposed to you in 2026, ask the supplier which platform they'll use, what happens when that base model is retired, and why instructions with examples weren't enough. A good supplier will have a clear answer to all three.

Six business needs, matched to the right lever

Here are the requests that come up most often, with the lever that fits each and what it looks like in practice.

Staff questions from the policy manual

RAG, through a shared project. A nursery's staff asking "what's the ratio for the toddler room on a trip?" or "who do I call if a parent is late for collection?" need answers from the current policy set, with the policy named so they can check it. A Claude or ChatGPT project holding the policies does this on day one.

Patient or customer FAQs on the website

RAG again, but inside a chatbot platform that adds handover to a human and logging, rather than a staff project. The knowledge is the same leaflets; the wrapper is different because the audience is the public.

Letters in the house style

Instructions plus examples, not fine-tuning. The dental practice wants treatment-plan letters that follow its structure: greeting, what we found, options with costs, next step, sign-off. Three anonymised example letters in the project, plus a list of rules, gets there. You'll see the difference in the sample further down.

Sorting thousands of emails into your own categories

Start with a small, cheap model and a prompt containing your categories and a dozen labelled examples. On OpenAI's API, gpt-5.6-luna costs $0.20 per million input tokens and $1.20 per million output tokens. Sorting 1,200 emails a month at about 800 tokens in and 50 out comes to roughly 0.96 million input tokens (about $0.19) and 0.06 million output tokens (about $0.07): well under a dollar a month. Fine-tuning becomes worth discussing only if accuracy stalls below what you need and volumes are far higher. The API maths is worked through in what the ChatGPT API costs for an automation.

Output in an exact shape for another system

Structured outputs, a feature where you give the model a strict schema (a template of named fields) and it must fill it, solve this without training. A language school that wants every enquiry turned into "name, level, preferred days, course type" for its booking sheet should use that, not a fine-tuned model.

Today's diary, stock or order status

Neither RAG over documents nor fine-tuning. Live facts need a live connection to the system that holds them, through an app, connector or API. A project full of exported appointment lists is out of date by lunchtime.

A dental practice decision, costed

Take the illustrative seven-person practice from earlier. It has three jobs in mind: answer reception's questions from its leaflets and policies, draft treatment-plan letters in its format, and sort the shared inbox into "appointment request", "billing", "clinical question", "complaint" and "other". A supplier has quoted for a fine-tuned model covering all three.

JobLever chosenToolMonthly cost at listSet-up time
Reception questionsRAGClaude Team project, 2 seats$40 on annual billing ($50 monthly)3 hours, mostly removing old leaflets
Treatment-plan lettersInstructions plus 3 examplesSame projectIncluded2 hours writing rules and anonymising examples
Inbox sortingSmall model, prompt with 12 examplesAn automation calling gpt-5.6-lunaUnder $1 in API usage, plus the automation tool's planHalf a day, including testing on 100 past emails
All three, fine-tunedFine-tuningSupplier's quoteHosting and retraining on top of the buildSeveral weeks, plus a training set of past letters and emails

The practice goes with the first three rows. The deciding factor wasn't the money, although that helped. It was that the fee list changes twice a year and the practice wanted to update it by replacing a file, not by paying for a retraining run. The inbox sorter reached 94 correct out of 100 on past emails in the illustration; the six misses were all borderline "clinical question or complaint" messages, which now go to the practice manager whatever the label says.

The same letter request, with and without examples

This is the test that usually ends the fine-tuning conversation. First, a bare request:

Write a treatment-plan letter for a patient who needs two fillings
and has a choice between white and silver fillings.

An illustrative result: a friendly but generic letter that opens with "I hope this letter finds you well", invents a price for each filling type, lists the risks of both in a long paragraph, and signs off "The Dental Team". Wrong structure, invented numbers, not the practice's voice.

Now the same request inside the practice's project, whose instructions include this:

Letters follow this structure, always:
1. "Dear [first name]," then one sentence on the visit date.
2. "What we found": plain words, no jargon, one short paragraph.
3. "Your options": a numbered list; each option has what it involves,
   the fee from the current fee list (quote it exactly, never estimate)
   and how many visits it takes.
4. "Next step": one sentence on how to book.
5. Sign off with the dentist's name and "for the practice".
Never use "I hope this letter finds you well".
Three approved example letters are attached as Example-1 to Example-3.
If a fee isn't in the fee list, write [FEE NEEDED] instead.

The illustrative result now follows the five-part structure, takes fees from the fee list, and marks the one treatment missing from the list as [FEE NEEDED]. What still needs fixing: it described one option as "the best choice", which the dentist doesn't want in writing. One extra rule ("present options neutrally; the dentist recommends in person") fixed it for every letter after. Fine-tuning would have needed a new training run to make that change.

Symptoms that you picked the wrong lever

If you already have something in place, these signs tell you it's the wrong kind of fix.

  • Answers go stale after every price rise or policy change. Facts are being carried by instructions, training or a copy-pasted prompt instead of a document that gets replaced. Move them into files the tool retrieves from.
  • The assistant answers confidently about things that aren't in your documents. Retrieval is in place, but nothing tells the model to stop when the files are silent. Add "if the documents don't cover it, say so" and test with questions you know aren't covered.
  • The facts are right but the format wanders. This is a behaviour problem. Add a fixed structure and two or three approved examples before considering anything heavier.
  • Every request is slow and costs more than expected. Someone is pasting the whole manual into each prompt. Use a project or retrieval so only relevant passages are read.
  • Nobody can say what happens when the model is retired. Typical of fine-tuned or custom-built set-ups with no owner. Write down the base model, the provider's retirement policy and who will rebuild it.

A three-question decision test

Run any AI idea through these in order and you'll land on the right lever.

  1. Does the right answer depend on facts that change? If yes, the facts must come from documents (RAG) or from a live system (an app connection). Never from training.
  2. Is the problem the shape, tone or consistency of the output? If yes, write instructions and add approved examples. Test on 20 real cases. Only if you're running many thousands of requests a month and still can't reach the accuracy you need is fine-tuning worth pricing, and check first that the provider still offers it.
  3. Is it currently a custom GPT? Move it to a Project, Gem or plugin before 11 December 2026, and test it with the same questions you use today.

For most small businesses, the answer to the title is RAG through a project you already pay for, a well-written set of instructions, and nothing trained at all.

Follow-up questions on RAG, GPTs and fine-tuning

Can I fine-tune ChatGPT on my documents so it just knows them?

Not in the way most people imagine. Fine-tuning teaches a model patterns from example conversations; it doesn't reliably store facts you can update. OpenAI has also closed self-serve fine-tuning to new organisations. Uploading documents to a ChatGPT or Claude project, which uses retrieval, gets you the 'it knows our stuff' result and lets you change a price by editing one file.

What happens to the custom GPTs we already use?

They stop running on 11 December 2026, with approved Enterprise deferrals to 11 February 2027. OpenAI offers a migration that turns a GPT into a plugin, with its instructions becoming a skill and its knowledge files becoming reference files. Alternatively, copy the instructions and files into a shared ChatGPT Project. Do either before December, and test the result with your usual questions.

Is RAG safe for confidential documents?

It's as safe as the tool holding the documents. On business plans such as ChatGPT Business, Claude Team or Gemini in Workspace, files aren't used for training by default and access follows the project's sharing. The real risks are sharing a project too widely and uploading material staff shouldn't see, so decide who can open each project before adding sensitive files.

Further reads

Sources: OpenAI API deprecations page (fine-tuning availability changes, 2026); OpenAI help page on custom GPT retirement and migration FAQ; Claude help pages on retrieval for projects; Google Cloud documentation on supervised tuning for Gemini models and the Gemini Enterprise Agent Platform name change; OpenAI API pricing page. Checked September 2026.

Unsure which lever your AI problem needs?

On a 1:1 call we'll take the job you want AI to do, work out whether it's a knowledge, format or live-data problem, and set up the simplest version with the tools you already have.

Book a 1:1 call with me