Most small businesses should start with paid models, because they arrive as finished products: a chat app, admin controls, support and business terms for $20-$25 a seat a month. Open models make sense in three cases: documents that must stay on your own hardware, automations where you want one model version fixed for years, and software you're building yourself.
The word "open" hides more variety than most owners expect. Nearly all of these models are open-weight: you can download the trained model and run it anywhere, but you don't get the training data or training code, which the Open Source Initiative's definition of open-source AI requires. Each also carries its own licence, from the permissive Apache 2.0, through Meta's Llama 4 Community License with its attribution rules, to licences that forbid commercial use outright. And "open" says nothing about privacy. An open model used through someone else's app or website is as much a cloud service as any paid one.
Four ways to reach a model, and who holds your text in each
Owners often frame this as "ChatGPT or an open-source model", but the real decision is how you reach the model and who ends up holding what your staff type. There are four routes.
| Route | How you pay | Examples, September 2026 | Who holds what you type |
|---|---|---|---|
| Paid model in a finished app | Per seat, monthly or annual | ChatGPT Business, Claude Team, Gemini in Google Workspace | The vendor, under business terms that exclude training by default |
| Paid model through an API | Per token, billed separately from any subscription | OpenAI API, Anthropic API | The vendor, under API terms; OpenAI doesn't train on API data by default |
| Open-weight model on a cloud host | Per token, at the host's prices | Gemma 4, gpt-oss, Mistral or DeepSeek models on Amazon Bedrock | The host, under its own terms |
| Open-weight model on your own machine | Hardware and staff time; the software can be free | A downloaded model running in Ollama or LM Studio | You, and nobody else |
Only the last row keeps text entirely in your hands, and it's also the row that turns you into the IT department. Running AI on your own computers covers the hardware side. Whichever route you use, the model family you pick brings its own licence, price and trade-offs.
Licences that decide what a free model lets you do
A downloaded model comes with a licence, and the licence, not the word "open", decides what you're allowed to do with it. These are the terms the publishers state on their own pages for families a small business is likely to meet, as of September 2026:
| Model family | Licence | What it means for a small business |
|---|---|---|
| OpenAI gpt-oss (20b and 120b) | Apache 2.0 | Use it, change it and build products on it commercially; it comes "as is", with no warranty |
| Google Gemma 4 (released March 2026) | Apache 2.0 | The same freedoms. Earlier Gemma versions (1 to 3n) sit under Google's Gemma Terms of Use instead, which add a prohibited-use policy |
| Meta Llama 4 | Llama 4 Community License Agreement | Commercial use allowed with conditions: follow Meta's acceptable use policy, show "Built with Llama" if you make a product containing it available, start any derived model's name with "Llama", and get a separate licence above 700 million monthly users |
| Mistral Small 4, Mistral Large 3, Ministral 3 | Apache 2.0 | Permissive, like gpt-oss |
| Mistral Voxtral TTS (a text-to-speech model) | CC BY-NC 4.0 | Non-commercial only, so not for anything customer-facing in a business |
| DeepSeek V4.1 Flash | MIT | Permissive, commercial use allowed; at 763 billion parameters it's far too big for office hardware, so businesses reach it through a host |
Two details in that table catch people out. First, licences change between versions of the same family: a developer who built on Gemma 3 is bound by different terms from one who built on Gemma 4, so always record the exact version. Second, the 700-million-user clause in Llama's licence is irrelevant to a small firm, but it's the reason the licence isn't open source in the strict sense. The parts that do affect you are attribution and the acceptable use policy.
Here's how that plays out. A painter and decorator's web developer builds a colour-advice chatbot for the website on a Llama 4 model. Customers use it, so the business is making a service that contains Llama available to others. Under the licence, the site should prominently display "Built with Llama", and the bot has to stay within Meta's acceptable use policy. Nobody mentioned either at handover. Neither costs much to fix, but both belong on the list of questions you ask a developer before signing off, and a solicitor is worth a short call if the product is central to your business.
Where paid models win, and where open ones do
Once the licence is acceptable, the choice comes down to a handful of criteria. This is the table I'd put in front of a small team before it commits either way.
| What matters to you | Paid model (app or API) | Open-weight model |
|---|---|---|
| Up and running this week | Wins: sign up, invite staff, done | Needs a host or hardware, plus setup time |
| Text never leaves your machines | Not possible; business terms limit what the vendor does with it | Wins, when you run it yourself |
| User list, single sign-on, data export | Built into business plans | You assemble them, for example with a self-hosted front end |
| Hard reasoning and very long documents | Frontier paid models are built for this; test on your own work | Large open models compete, but the ones that fit office hardware are small |
| The same model for years | Vendors retire products: OpenAI closed Sora's apps in April 2026 and its API in September 2026 | Wins: a downloaded file doesn't change or vanish |
| Cost at light use | $20-$25 a seat covers most staff | The software is free; the hardware and your time aren't |
| Cost at high volume | Cheap tiers cost cents per million tokens | Hosted open models cost about the same; local is a fixed cost |
| Help when it breaks | Vendor support and status pages | Forums, or whoever set it up |
| Rights in the output | OpenAI and Anthropic assign you their rights in outputs "if any" | Governed by the licence; Apache 2.0 and MIT place few limits |
The version-stability row deserves more weight than it usually gets. Paid vendors change products constantly: OpenAI's custom GPTs stop running on 11 December 2026, and when the calendar tool Clockwise shut down in March 2026 it deleted user data rather than transferring it. If an automation depends on a model behaving exactly as it did when you tested it, an open-weight file you control is the one thing nobody can withdraw. Keeping your data and prompts portable is the cheaper insurance for everything else.
The price gap is smaller than "free" suggests
Open means cheaper only when you run the model on hardware you already own. When a cloud host runs an open model for you, it charges per token (roughly three-quarters of an English word), and the cheapest paid models now sit in the same range. Prices per million tokens, input then output:
| Model | Type | Input / output per million tokens |
|---|---|---|
| gpt-5-nano (OpenAI API) | Paid | $0.05 / $0.40 |
| gpt-6-luna (OpenAI API) | Paid | $0.10 / $0.50 |
| Gemma 4 31B on Amazon Bedrock | Open, hosted | $0.14 / $0.40 |
| Mistral Large 3 on Amazon Bedrock | Open, hosted | $0.50 / $1.50 |
| Claude Haiku 4.5 (Anthropic API) | Paid | $1 / $5 |
| Claude Sonnet 5 (Anthropic API) | Paid | $2 / $10 |
The Bedrock figures are its listed on-demand prices in September 2026 and vary by region; Anthropic's Batch API halves its prices for work that can wait. At small-business volumes, the bill is rarely what separates the options, as the next example shows. For the staff-time costs hiding behind "free", the true cost of open-source AI goes line by line.
A worked choice: an architect practice reading tender documents
Consider a nine-person architect practice that wants every incoming tender document and client brief read, with the requirements pulled into a checklist: deliverables, deadlines, submission format and any standards the client names. It receives about 60 documents a month, averaging 8,000 words each.
The volume sum. 8,000 words is roughly 10,700 tokens, so 60 documents is about 640,000 input tokens a month. A checklist of around 800 words is about 1,070 tokens, so output adds roughly 64,000 tokens. At the listed prices:
- Claude Sonnet 5: 0.64 million × $2 plus 0.064 million × $10, about $1.92 a month.
- Gemma 4 31B on Bedrock: 0.64 × $0.14 plus 0.064 × $0.40, about $0.12 a month.
- gpt-6-luna: 0.64 × $0.10 plus 0.064 × $0.50, about $0.10 a month.
- A local model: no per-token cost, but a 10,700-token document is far past the 4k-token default context Ollama uses on machines with under 24GB of graphics memory, so it needs a better-specified machine and changed settings.
With the running cost between ten cents and two dollars a month, price can't decide this. Three other questions do.
- Do client contracts allow cloud processing? The practice reads its standard appointment terms and finds nothing barring a processor with no-training terms. One public-sector client's tender does forbid third-party processing, so those documents go down a separate route.
- Which option is most accurate on these documents? The practice runs 12 past tenders with known requirement lists (180 requirements in total) through four candidates.
- Who maintains it? Nobody in the practice wants to look after a server.
Illustrative results from that test, with the models anonymised because your documents will rank them differently:
| Candidate | Requirements found (of 180) | Invented items | Documents cut short |
|---|---|---|---|
| A: frontier paid model via API | 172 | 2 | 0 |
| B: cheapest paid tier via API | 158 | 5 | 0 |
| C: large open model on a cloud host | 165 | 3 | 0 |
| D: small open model on an existing 16GB laptop | 129 | 4 | 5 |
The practice picks candidate A for everyday tenders: at about $2 a month, catching seven more requirements than the next-best candidate is worth far more than the price difference, since one missed deliverable can cost a bid. For the client that forbids third-party processing, it borrows the D setup on a machine with more memory and splits each document by section, accepting slower, less complete results and a longer manual check. Two routes, each chosen on evidence rather than on the label.
Run the same test on both before you commit
The test that settled the architect practice's choice takes an afternoon. Pick 10-20 past documents where you already know the right answer, run the identical prompt through each candidate, and count what was found, missed and invented. This is the prompt it used:
Read the tender document below and list every requirement the bidder must meet.
Return a table: Ref | Requirement | Deadline or date | Page or section
Quote the document's own words in the Requirement column.
If a requirement has no deadline, write "none stated".
Leave out background and the client's general aspirations.
Finish with the line: "Requirements found: N".
DOCUMENT:
[paste the tender text here]
An illustrative answer from the small open model, candidate D:
Ref | Requirement | Deadline or date | Page or section
1 | Submit design proposals as a single PDF | 14 March | 2.1
2 | Hold professional indemnity cover of 5 million | none stated | 4.3
3 | Attend the site visit | 2 March | 2.4
4 | Support the client's sustainability aspirations | none stated | 1.2
Requirements found: 4
What the reviewer marked:
- Line 1 lost detail. The document gave a time as well as a date ("by 12 noon"). A deadline without the time is a deadline you can miss.
- Line 2 paraphrased instead of quoting. The tender specified the currency and that the cover applied "for each and every claim", which changes the insurance the practice needs. The prompt said quote; small models often don't.
- Line 4 isn't a requirement. It's exactly the kind of aspiration the prompt said to leave out.
- Four found, fifteen in the document. Everything from section 5 onwards was missing because the text overflowed the context window. The count line exposed it at a glance.
Candidate A's answer to the same document wasn't flawless either: it listed a requirement under section 6.2, and the document has no section 6.2. Every candidate needs the same human check against the source; the test tells you how much fixing each one needs, which is the number that matters.
Mistakes that surface months after the choice
Choosing an open model "for privacy", then using the publisher's app. DeepSeek's model weights are MIT-licensed, but the DeepSeek app and website are a cloud service with their own privacy policy. Downloading the weights and running them yourself is a completely different arrangement from chatting on the website, as whether DeepSeek is safe with business data explains.
Building on a non-commercial licence. Take a kitchen fitter whose developer picks a free voice model to read out appointment reminders on the phone line, because it sounded the most natural in testing. If that model carries a CC BY-NC 4.0 licence, as Mistral's Voxtral TTS does, a business using it for customer calls is outside the licence, and the fix is a rebuild on a different model. Ask for the licence name of every model in anything a developer hands over.
Assuming a paid model stays the same. Vendors update the models behind familiar product names and retire old ones. An automation that sorted invoices perfectly in March can drift by September with nobody changing a setting. Keep your test set, and rerun it whenever the vendor announces a model change.
Comparing unlike with unlike. Owners test a large paid model against a tiny open model on an old laptop, see the gap, and conclude open models are useless. Compare like with like: a hosted large open model against the paid option, or a local model against the job it's actually meant for.
Forgetting the running costs you can't see. A first-year accountant might say "it's free" about a local model, but somebody installs it, updates it, and checks its output. If nobody is named for that, the setup decays quietly, and a tool nobody maintains ends up worse than the paid seat it replaced.
The mixed setup most small firms end up with
For most small businesses the answer is both, split by job: paid business seats for everyday writing and research, the cheap paid API tiers or a hosted open model for automations, and an open-weight model on your own hardware only for the set of documents that genuinely can't leave the building. What stops that becoming a mess is a one-page record. Here's a filled-in example for a six-person accountancy practice:
AI MODEL RECORD: [practice name] Reviewed: [month, year]
Everyday assistant: Claude Team, 6 Standard seats, annual billing
Confidential set: payroll files and client identity documents
Route: gpt-oss-20b in LM Studio on the partners' 32GB desktop
Licence: Apache 2.0 (checked on the publisher's page, [date])
Automation: bank-statement categoriser via the OpenAI API
Model: gpt-5-nano; version as shown in the API console
Test set: 20 past statements with known answers; rerun on any change
Exit route: all prompts kept as text files on the shared drive
Owner: [first name], reviews this page every 6 months
That record answers the questions that matter when something changes: which model does what, under which licence, how you'd know if it got worse, and how you'd move away. With it in place, the open-versus-paid question stops being a one-off bet and becomes a routine you revisit twice a year. If an API is the piece you're weighing up, what the ChatGPT API costs for a business automation works through the per-token maths for common jobs.
Open models and paid models: follow-up questions
Can I use an open-source AI model commercially without paying?
It depends on the licence of the exact version you download. Apache 2.0 and MIT models can be used commercially without a fee; Meta's Llama 4 licence allows commercial use with conditions such as attribution; a CC BY-NC licence rules commercial use out. The model may be free, but the machine or cloud host that runs it and the time to maintain it are not.
Are open models less safe to use than paid ones?
A model file run on your own machine sends nothing anywhere, so on privacy it can be safer. The risks sit elsewhere: community-modified copies can have their safety training stripped out, so download only from the publisher's official account, and nobody patches or supports your setup unless you arrange it. A paid vendor handles those jobs but holds your text.
If we build on a paid API now, can we switch to an open model later?
Yes, if you plan for it. Keep your prompts as plain text files, keep a test set of past documents with known correct answers, and store outputs in your own systems rather than only inside the vendor's app. Switching then means rerunning the test set on the new model and comparing scores, not rebuilding from memory.
Further reads
- Cloud AI vs On-Device AI: Which Is Safer for Business Data? — Where cloud and on-device AI actually keep your business data.
- What If Your AI Vendor Shuts Down? Checks Before You Commit — The supplier-risk checks behind the 'version stability' row.
- How to Evaluate an AI Software Vendor: A Small Business Scorecard — A scorecard for the paid vendors on your shortlist.
- What Is an API? Why It Matters When You Buy Software — Plain-English background on APIs before pricing per token.
- What to Check in an AI Vendor's Data Processing Agreement — What to check in the data terms of whichever route you pick.
- Who in Your Team Actually Needs a Paid AI Licence? — Work out how many paid seats your team really needs.
- AI Tools and AI Development: The Complete 2026 Guide — the AI hub, including every tutorial in the AI-for-business series.
Sources: Hugging Face model cards for OpenAI gpt-oss-20b and gpt-oss-120b; Google's Gemma 4 announcement, Gemma releases page and Gemma Terms of Use; Meta's Llama 4 Community License Agreement; Mistral's models overview; DeepSeek's R1 and V3 repositories and the V4.1 Flash model card; the Open Source Initiative's Open Source AI Definition 1.0; Amazon Bedrock pricing; OpenAI and Anthropic API pricing as of September 2026.