Open-Source vs Paid AI Models: What Small Businesses Should Know

Coding Liquids tutorial cover featuring Sagnik Bhattacharya for Open-Source vs Paid AI Models: What Small Businesses Should Know.
Coding Liquids tutorial cover featuring Sagnik Bhattacharya for Open-Source vs Paid AI Models: What Small Businesses Should Know.

Most small businesses should start with paid models, because they arrive as finished products: a chat app, admin controls, support and business terms for $20-$25 a seat a month. Open models make sense in three cases: documents that must stay on your own hardware, automations where you want one model version fixed for years, and software you're building yourself.

The word "open" hides more variety than most owners expect. Nearly all of these models are open-weight: you can download the trained model and run it anywhere, but you don't get the training data or training code, which the Open Source Initiative's definition of open-source AI requires. Each also carries its own licence, from the permissive Apache 2.0, through Meta's Llama 4 Community License with its attribution rules, to licences that forbid commercial use outright. And "open" says nothing about privacy. An open model used through someone else's app or website is as much a cloud service as any paid one.

Follow me on Instagram@sagnikteaches

Four ways to reach a model, and who holds your text in each

Owners often frame this as "ChatGPT or an open-source model", but the real decision is how you reach the model and who ends up holding what your staff type. There are four routes.

Connect on LinkedInSagnik Bhattacharya
RouteHow you payExamples, September 2026Who holds what you type
Paid model in a finished appPer seat, monthly or annualChatGPT Business, Claude Team, Gemini in Google WorkspaceThe vendor, under business terms that exclude training by default
Paid model through an APIPer token, billed separately from any subscriptionOpenAI API, Anthropic APIThe vendor, under API terms; OpenAI doesn't train on API data by default
Open-weight model on a cloud hostPer token, at the host's pricesGemma 4, gpt-oss, Mistral or DeepSeek models on Amazon BedrockThe host, under its own terms
Open-weight model on your own machineHardware and staff time; the software can be freeA downloaded model running in Ollama or LM StudioYou, and nobody else

Only the last row keeps text entirely in your hands, and it's also the row that turns you into the IT department. Running AI on your own computers covers the hardware side. Whichever route you use, the model family you pick brings its own licence, price and trade-offs.

Subscribe on YouTube@codingliquids

Licences that decide what a free model lets you do

A downloaded model comes with a licence, and the licence, not the word "open", decides what you're allowed to do with it. These are the terms the publishers state on their own pages for families a small business is likely to meet, as of September 2026:

Model familyLicenceWhat it means for a small business
OpenAI gpt-oss (20b and 120b)Apache 2.0Use it, change it and build products on it commercially; it comes "as is", with no warranty
Google Gemma 4 (released March 2026)Apache 2.0The same freedoms. Earlier Gemma versions (1 to 3n) sit under Google's Gemma Terms of Use instead, which add a prohibited-use policy
Meta Llama 4Llama 4 Community License AgreementCommercial use allowed with conditions: follow Meta's acceptable use policy, show "Built with Llama" if you make a product containing it available, start any derived model's name with "Llama", and get a separate licence above 700 million monthly users
Mistral Small 4, Mistral Large 3, Ministral 3Apache 2.0Permissive, like gpt-oss
Mistral Voxtral TTS (a text-to-speech model)CC BY-NC 4.0Non-commercial only, so not for anything customer-facing in a business
DeepSeek V4.1 FlashMITPermissive, commercial use allowed; at 763 billion parameters it's far too big for office hardware, so businesses reach it through a host

Two details in that table catch people out. First, licences change between versions of the same family: a developer who built on Gemma 3 is bound by different terms from one who built on Gemma 4, so always record the exact version. Second, the 700-million-user clause in Llama's licence is irrelevant to a small firm, but it's the reason the licence isn't open source in the strict sense. The parts that do affect you are attribution and the acceptable use policy.

Here's how that plays out. A painter and decorator's web developer builds a colour-advice chatbot for the website on a Llama 4 model. Customers use it, so the business is making a service that contains Llama available to others. Under the licence, the site should prominently display "Built with Llama", and the bot has to stay within Meta's acceptable use policy. Nobody mentioned either at handover. Neither costs much to fix, but both belong on the list of questions you ask a developer before signing off, and a solicitor is worth a short call if the product is central to your business.

Where paid models win, and where open ones do

Once the licence is acceptable, the choice comes down to a handful of criteria. This is the table I'd put in front of a small team before it commits either way.

What matters to youPaid model (app or API)Open-weight model
Up and running this weekWins: sign up, invite staff, doneNeeds a host or hardware, plus setup time
Text never leaves your machinesNot possible; business terms limit what the vendor does with itWins, when you run it yourself
User list, single sign-on, data exportBuilt into business plansYou assemble them, for example with a self-hosted front end
Hard reasoning and very long documentsFrontier paid models are built for this; test on your own workLarge open models compete, but the ones that fit office hardware are small
The same model for yearsVendors retire products: OpenAI closed Sora's apps in April 2026 and its API in September 2026Wins: a downloaded file doesn't change or vanish
Cost at light use$20-$25 a seat covers most staffThe software is free; the hardware and your time aren't
Cost at high volumeCheap tiers cost cents per million tokensHosted open models cost about the same; local is a fixed cost
Help when it breaksVendor support and status pagesForums, or whoever set it up
Rights in the outputOpenAI and Anthropic assign you their rights in outputs "if any"Governed by the licence; Apache 2.0 and MIT place few limits

The version-stability row deserves more weight than it usually gets. Paid vendors change products constantly: OpenAI's custom GPTs stop running on 11 December 2026, and when the calendar tool Clockwise shut down in March 2026 it deleted user data rather than transferring it. If an automation depends on a model behaving exactly as it did when you tested it, an open-weight file you control is the one thing nobody can withdraw. Keeping your data and prompts portable is the cheaper insurance for everything else.

The price gap is smaller than "free" suggests

Open means cheaper only when you run the model on hardware you already own. When a cloud host runs an open model for you, it charges per token (roughly three-quarters of an English word), and the cheapest paid models now sit in the same range. Prices per million tokens, input then output:

ModelTypeInput / output per million tokens
gpt-5-nano (OpenAI API)Paid$0.05 / $0.40
gpt-6-luna (OpenAI API)Paid$0.10 / $0.50
Gemma 4 31B on Amazon BedrockOpen, hosted$0.14 / $0.40
Mistral Large 3 on Amazon BedrockOpen, hosted$0.50 / $1.50
Claude Haiku 4.5 (Anthropic API)Paid$1 / $5
Claude Sonnet 5 (Anthropic API)Paid$2 / $10

The Bedrock figures are its listed on-demand prices in September 2026 and vary by region; Anthropic's Batch API halves its prices for work that can wait. At small-business volumes, the bill is rarely what separates the options, as the next example shows. For the staff-time costs hiding behind "free", the true cost of open-source AI goes line by line.

A worked choice: an architect practice reading tender documents

Consider a nine-person architect practice that wants every incoming tender document and client brief read, with the requirements pulled into a checklist: deliverables, deadlines, submission format and any standards the client names. It receives about 60 documents a month, averaging 8,000 words each.

The volume sum. 8,000 words is roughly 10,700 tokens, so 60 documents is about 640,000 input tokens a month. A checklist of around 800 words is about 1,070 tokens, so output adds roughly 64,000 tokens. At the listed prices:

  • Claude Sonnet 5: 0.64 million × $2 plus 0.064 million × $10, about $1.92 a month.
  • Gemma 4 31B on Bedrock: 0.64 × $0.14 plus 0.064 × $0.40, about $0.12 a month.
  • gpt-6-luna: 0.64 × $0.10 plus 0.064 × $0.50, about $0.10 a month.
  • A local model: no per-token cost, but a 10,700-token document is far past the 4k-token default context Ollama uses on machines with under 24GB of graphics memory, so it needs a better-specified machine and changed settings.

With the running cost between ten cents and two dollars a month, price can't decide this. Three other questions do.

  1. Do client contracts allow cloud processing? The practice reads its standard appointment terms and finds nothing barring a processor with no-training terms. One public-sector client's tender does forbid third-party processing, so those documents go down a separate route.
  2. Which option is most accurate on these documents? The practice runs 12 past tenders with known requirement lists (180 requirements in total) through four candidates.
  3. Who maintains it? Nobody in the practice wants to look after a server.

Illustrative results from that test, with the models anonymised because your documents will rank them differently:

CandidateRequirements found (of 180)Invented itemsDocuments cut short
A: frontier paid model via API17220
B: cheapest paid tier via API15850
C: large open model on a cloud host16530
D: small open model on an existing 16GB laptop12945

The practice picks candidate A for everyday tenders: at about $2 a month, catching seven more requirements than the next-best candidate is worth far more than the price difference, since one missed deliverable can cost a bid. For the client that forbids third-party processing, it borrows the D setup on a machine with more memory and splits each document by section, accepting slower, less complete results and a longer manual check. Two routes, each chosen on evidence rather than on the label.

Run the same test on both before you commit

The test that settled the architect practice's choice takes an afternoon. Pick 10-20 past documents where you already know the right answer, run the identical prompt through each candidate, and count what was found, missed and invented. This is the prompt it used:

Read the tender document below and list every requirement the bidder must meet.
Return a table: Ref | Requirement | Deadline or date | Page or section
Quote the document's own words in the Requirement column.
If a requirement has no deadline, write "none stated".
Leave out background and the client's general aspirations.
Finish with the line: "Requirements found: N".

DOCUMENT:
[paste the tender text here]

An illustrative answer from the small open model, candidate D:

Ref | Requirement | Deadline or date | Page or section
1 | Submit design proposals as a single PDF | 14 March | 2.1
2 | Hold professional indemnity cover of 5 million | none stated | 4.3
3 | Attend the site visit | 2 March | 2.4
4 | Support the client's sustainability aspirations | none stated | 1.2
Requirements found: 4

What the reviewer marked:

  • Line 1 lost detail. The document gave a time as well as a date ("by 12 noon"). A deadline without the time is a deadline you can miss.
  • Line 2 paraphrased instead of quoting. The tender specified the currency and that the cover applied "for each and every claim", which changes the insurance the practice needs. The prompt said quote; small models often don't.
  • Line 4 isn't a requirement. It's exactly the kind of aspiration the prompt said to leave out.
  • Four found, fifteen in the document. Everything from section 5 onwards was missing because the text overflowed the context window. The count line exposed it at a glance.

Candidate A's answer to the same document wasn't flawless either: it listed a requirement under section 6.2, and the document has no section 6.2. Every candidate needs the same human check against the source; the test tells you how much fixing each one needs, which is the number that matters.

Mistakes that surface months after the choice

Choosing an open model "for privacy", then using the publisher's app. DeepSeek's model weights are MIT-licensed, but the DeepSeek app and website are a cloud service with their own privacy policy. Downloading the weights and running them yourself is a completely different arrangement from chatting on the website, as whether DeepSeek is safe with business data explains.

Building on a non-commercial licence. Take a kitchen fitter whose developer picks a free voice model to read out appointment reminders on the phone line, because it sounded the most natural in testing. If that model carries a CC BY-NC 4.0 licence, as Mistral's Voxtral TTS does, a business using it for customer calls is outside the licence, and the fix is a rebuild on a different model. Ask for the licence name of every model in anything a developer hands over.

Assuming a paid model stays the same. Vendors update the models behind familiar product names and retire old ones. An automation that sorted invoices perfectly in March can drift by September with nobody changing a setting. Keep your test set, and rerun it whenever the vendor announces a model change.

Comparing unlike with unlike. Owners test a large paid model against a tiny open model on an old laptop, see the gap, and conclude open models are useless. Compare like with like: a hosted large open model against the paid option, or a local model against the job it's actually meant for.

Forgetting the running costs you can't see. A first-year accountant might say "it's free" about a local model, but somebody installs it, updates it, and checks its output. If nobody is named for that, the setup decays quietly, and a tool nobody maintains ends up worse than the paid seat it replaced.

The mixed setup most small firms end up with

For most small businesses the answer is both, split by job: paid business seats for everyday writing and research, the cheap paid API tiers or a hosted open model for automations, and an open-weight model on your own hardware only for the set of documents that genuinely can't leave the building. What stops that becoming a mess is a one-page record. Here's a filled-in example for a six-person accountancy practice:

AI MODEL RECORD: [practice name]              Reviewed: [month, year]
Everyday assistant:  Claude Team, 6 Standard seats, annual billing
Confidential set:    payroll files and client identity documents
  Route:             gpt-oss-20b in LM Studio on the partners' 32GB desktop
  Licence:           Apache 2.0 (checked on the publisher's page, [date])
Automation:          bank-statement categoriser via the OpenAI API
  Model:             gpt-5-nano; version as shown in the API console
  Test set:          20 past statements with known answers; rerun on any change
Exit route:          all prompts kept as text files on the shared drive
Owner:               [first name], reviews this page every 6 months

That record answers the questions that matter when something changes: which model does what, under which licence, how you'd know if it got worse, and how you'd move away. With it in place, the open-versus-paid question stops being a one-off bet and becomes a routine you revisit twice a year. If an API is the piece you're weighing up, what the ChatGPT API costs for a business automation works through the per-token maths for common jobs.

Open models and paid models: follow-up questions

Can I use an open-source AI model commercially without paying?

It depends on the licence of the exact version you download. Apache 2.0 and MIT models can be used commercially without a fee; Meta's Llama 4 licence allows commercial use with conditions such as attribution; a CC BY-NC licence rules commercial use out. The model may be free, but the machine or cloud host that runs it and the time to maintain it are not.

Are open models less safe to use than paid ones?

A model file run on your own machine sends nothing anywhere, so on privacy it can be safer. The risks sit elsewhere: community-modified copies can have their safety training stripped out, so download only from the publisher's official account, and nobody patches or supports your setup unless you arrange it. A paid vendor handles those jobs but holds your text.

If we build on a paid API now, can we switch to an open model later?

Yes, if you plan for it. Keep your prompts as plain text files, keep a test set of past documents with known correct answers, and store outputs in your own systems rather than only inside the vendor's app. Switching then means rerunning the test set on the new model and comparing scores, not rebuilding from memory.

Further reads

Sources: Hugging Face model cards for OpenAI gpt-oss-20b and gpt-oss-120b; Google's Gemma 4 announcement, Gemma releases page and Gemma Terms of Use; Meta's Llama 4 Community License Agreement; Mistral's models overview; DeepSeek's R1 and V3 repositories and the V4.1 Flash model card; the Open Source Initiative's Open Source AI Definition 1.0; Amazon Bedrock pricing; OpenAI and Anthropic API pricing as of September 2026.

Unsure which model route suits your workload?

On a 1:1 call we'll look at the jobs you want AI for, which of them touch documents that can't leave your control, and whether a paid plan, an API or an open model fits each one.

Book a 1:1 call with me