Is Open-Source AI Really Free? The True Cost for a Small Business

Coding Liquids tutorial cover featuring Sagnik Bhattacharya for Is Open-Source AI Really Free? The True Cost for a Small Business.
Coding Liquids tutorial cover featuring Sagnik Bhattacharya for Is Open-Source AI Really Free? The True Cost for a Small Business.

No. Open-source AI models are free to download, not to run. You pay for the computer that runs them, either your own or a rented server (from about $0.34 an hour for a graphics card to $2,500 a month for an always-on large model), plus setup and upkeep time. For most small businesses that costs more than a $20 subscription.

The exception is when privacy, offline use or very high volume is the reason for looking. "Open-source" here usually means open-weight: the company publishes the trained model's files under a licence that lets you run them yourself. OpenAI's gpt-oss models, Meta's Llama family, Google's Gemma, and models from Mistral, Qwen and DeepSeek all fall into this group. The licence is free. Everything around it isn't.

Follow me on Instagram@sagnikteaches

Three ways to run an open model, and what each costs

How you run the model decides both the bill and whether you get the privacy benefit people usually want.

Connect on LinkedInSagnik Bhattacharya
RouteWhat you payDoes data leave your premises?Best for
On your own computer, using a free runner such as LM Studio or OllamaNothing extra if the machine is capable; otherwise new hardware, typically over $1,000 for a machine with 32GB of memoryNoPrivate drafting and occasional batch jobs
A rented graphics card (GPU) by the hourRunpod lists an RTX 4090 at $0.34 an hour on its community cloud or $0.74 on its secure cloud; an 80GB A100 at $1.19 to $1.59; an H100 at $1.99 to $3.49Yes, to the rental providerOccasional large batches, or larger models than your hardware can hold
A hosted API for an open modelPer token: Together AI lists gpt-oss-120b at $0.15 per million input tokens and $0.60 per million outputYes, to the API providerCheap high-volume processing when privacy isn't the concern

Notice that the cheapest route by far, the hosted API, gives up the privacy advantage. You're then comparing one cloud AI provider with another, and the same questions about training, retention and processing agreements apply. Cloud AI vs on-device AI goes through that safety comparison.

Subscribe on YouTube@codingliquids

The itemised bill: one-off, monthly and your time

Here's a complete list of costs for running an open model yourself. Most owners count the first two lines and stop.

CostTypeTypical amount for a small business
Model licenceOne-off$0, subject to the licence terms
Runner software (Ollama is MIT-licensed; LM Studio is free for work use)One-off$0
HardwareOne-off$0 on a capable existing machine; over $1,000 for new hardware with 32GB of memory; far more for a server with an 80GB GPU
Rented GPU timeMonthlyFrom a few dollars for occasional batches to about $250 a month for a 4090 running all the time on a community cloud
ElectricityMonthlySmall for occasional use; noticeable for a machine working all day
Setup: choosing a model, installing, testingInternal time, one-offRoughly 4 to 12 hours for someone comfortable with software
Upkeep: updates, re-testing, fixing breakagesInternal time, monthly1 to 3 hours
Checking outputs from a weaker modelInternal time, ongoingDepends on the job; often more than with a frontier model
Integration with your other toolsInternal time or contractorNo built-in connectors to email or accounts; any link has to be built

Time estimates are illustrative, for a small business with one reasonably technical person; they'll be higher if everything is new to you.

A café prices one job four ways

An illustration with real list prices. A café wants AI to sort 600 online reviews and 400 supplier emails a month: tag each review by topic (food, service, wait time, price) and pull the price, item and delivery date out of each supplier email. That's about 1,000 items a month, averaging roughly 800 tokens in and 200 out, so 800,000 input tokens and 200,000 output tokens.

Route 1: ChatGPT Plus, pasted in batches
  $20/month + about 3 hours of copy-paste                 = $20 + 3 h

Route 2: API, small closed model (gpt-5.6-luna, $0.20 / $1.20 per M)
  0.8 x $0.20 + 0.2 x $1.20                               = $0.40/month
  + automation plan and a few hours of setup

Route 3: Hosted open model (gpt-oss-120b, $0.15 / $0.60 per M)
  0.8 x $0.15 + 0.2 x $0.60                               = $0.24/month
  + automation plan and a few hours of setup

Route 4: gpt-oss-20b on the owner's 16GB laptop
  $0 cash; about 8 h setup and testing, 1-2 h a month upkeep;
  runs overnight because the laptop is slow at this size

The token cost of the hosted open model is 24 cents a month, cheaper than the closed model by 16 cents. At this volume the difference is meaningless; the costs that matter are the automation around it and the owner's hours. Route 4 is the only one where the reviews and supplier prices never leave the café's own computer, and it costs no money, but it costs the most time and the laptop is tied up overnight.

The café chose Route 2 for the reviews, which are public anyway, and kept supplier emails on its business email suite's built-in AI, which already had a processing agreement. The open model lost not because it was worse but because the privacy it offered wasn't needed for public reviews, and the time it cost was real.

Cheap, middle and heavy set-ups, priced

Cheap: a laptop you already own

A photography studio owner with a recent 16GB laptop installs LM Studio and downloads gpt-oss-20b, which OpenAI says fits in 16GB of memory. Cost: $0 and an afternoon. It works for private drafting, such as rewording client contracts with names removed, but it's slow. Reported speeds on Apple laptops range from about 5 to 30 tokens a second depending on the chip and memory. A 300-word reply is about 400 tokens, so it takes somewhere between 13 seconds and over a minute. Fine for occasional use; frustrating as an everyday assistant.

Middle: a rented GPU for a monthly batch

A driving school wants to summarise a month of instructor lesson notes into progress reports once a month without sending them to a consumer AI service. Someone rents a GPU for three hours on a secure cloud at $0.74 an hour, runs the batch, and shuts it down. Cash cost: about $2.22 a month. Time: a day to build the first time, an hour a month after. The data does go to the rental provider, but onto a machine the school controls and deletes afterwards, which some owners are more comfortable with. Check the provider's own terms either way.

Heavy: an always-on server for the whole team

A ten-person business wants a private chat assistant everyone can use all day, running gpt-oss-120b, which needs a single 80GB GPU. Running an 80GB A100 around the clock (about 730 hours a month) at Runpod's listed rates costs about $869 to $1,161 a month, and an H100 about $1,453 to $2,548. Add setup, a login system, backups and someone responsible for keeping it running. Ten Claude Team or ChatGPT Business seats cost $200 to $250 a month. The private server costs roughly 3.5 to 13 times as much before counting anyone's time.

Costs people forget until they've started

  • The quality gap. Open models have improved fast, and OpenAI describes gpt-oss-120b as close to its o4-mini model on core reasoning tests. Smaller models that fit on ordinary hardware still make more mistakes than the paid frontier models. That's fine for tagging and drafting; less fine for anything customers see unchecked.
  • Your time is the biggest line. If the person setting it up costs the business $40 an hour, eight hours of setup is $320, which is sixteen months of a $20 subscription, before any upkeep.
  • Licence terms. gpt-oss uses the permissive Apache 2.0 licence. Meta's Llama 4 licence allows commercial use but needs a separate licence from Meta for services above 700 million monthly users and includes an acceptable use policy. Other models have their own terms; read them for each model and version.
  • Security is now your job. A model server reachable from the internet without a login is an open door to your data and your compute bill. A cloud provider handles that for you; self-hosting doesn't.
  • No support line. When an update breaks something, the fix is a forum post, not a helpdesk.
  • No connectors. Paid assistants now link to email, calendars and file storage. A self-hosted model links to nothing until you build it.

Here's how the quality gap tends to show up. An illustrative output from a small local model asked to tag a review and return JSON:

Review: "Lovely flat white but waited 25 minutes for two toasties,
staff seemed rushed. Won't rush back at those prices."

Output:
{"topics": ["food", "service"], "sentiment": "positive",
 "wait_time_minutes": 25}

It missed "price" as a topic and labelled a clearly negative review as positive. A larger model usually gets both right. The fix is a stricter prompt with two worked examples and a rule ("if the reviewer says they won't return, sentiment is negative"), then checking a sample of 20 outputs a week. That checking time belongs in your cost estimate.

A first-year comparison for a five-person team

Monthly figures hide the setup cost, so compare a full year. This illustration assumes a five-person team that wants an everyday assistant, and values internal time at $40 an hour; use your own figure.

OptionCash, year oneInternal time, year oneTotal at $40 an hour
Five individual $20 subscriptions$1,200About 2 hours to set upabout $1,280
Team plan, five seats at $20 annual$1,200About 3 hours, including admin settingsabout $1,320
Small open model on one existing capable machine, shared$010 hours setup plus 2 hours a month upkeep: 34 hoursabout $1,360, and only one person can use it at a time
Rented RTX 4090, community cloud, always onAbout $2,98012 hours setup plus 2 hours a month: 36 hoursabout $4,420

The "free" option costs about the same as the subscriptions once time is counted, and delivers less: a slower, smaller model on one machine, with no connectors. The rented server costs more than three times as much. The picture only flips when the subscription route isn't allowed for the data involved, or when volumes are far higher than a five-person team generates.

One more cost belongs on any rented-GPU budget: the one you forget to switch off. A realistic mistake, for illustration: a bicycle repair shop owner rents an 80GB A100 on a secure cloud for a Friday afternoon batch job, finishes early, and leaves it running over a long weekend. Four days at $1.59 an hour is about $153, for a job that needed three hours and under $5. Set a spending limit or an automatic shutdown with the provider before the first run, and check the dashboard when you finish.

Where open models genuinely come out ahead

There are real cases where running an open model yourself is the right call:

  • Data that must not leave your premises, where your contracts, your clients or your professional rules forbid sending it to a cloud AI provider at all, even under a processing agreement. This is the strongest reason, and the one to be honest about: if a business-plan agreement would satisfy the rule, a subscription is cheaper.
  • Offline or unreliable internet, such as a workshop or site with poor connectivity.
  • Very high volume. A rough break-even: a community-cloud A100 running all month costs about $869. At a blended rate of about $0.24 per million tokens for a hosted gpt-oss-120b (four parts input to one part output), $869 buys about 3.6 billion tokens, or roughly 120 million tokens a day. Few small businesses come anywhere near that; below it, paying per token is cheaper than renting a server all month.
  • Customisation, such as fine-tuning a model on thousands of your own examples, which some licences permit and closed models restrict.
  • Protection against a vendor changing terms. Weights you've downloaded can't be withdrawn or repriced. After a year in which several AI products were retired, that's a fair consideration, and it's covered in open-source vs paid AI models.

Trying it cheaply, if you still want to

  1. Pick one job with non-sensitive test data, such as tagging last month's public reviews.
  2. Install a free runner on a machine with at least 16GB of memory. LM Studio suits people who want a desktop app; Ollama suits anyone scripting.
  3. Start with a small model such as gpt-oss-20b and time a typical reply.
  4. Run 20 real examples through it and through the paid assistant you already use. Score both on accuracy and on how much editing each output needed before you'd use it.
  5. Log your hours from the first download onwards, including the time spent searching forums when something doesn't work. That number is usually the surprise.

After a week you'll know three things: whether the quality is good enough, how slow it feels, and what it cost in time. If you're weighing whether your own machines can cope, running AI on your own computers covers the hardware side in more depth, and n8n self-hosted vs cloud is a good comparison of the same "free software, paid upkeep" trade-off in automation.

Checking whether it's paying off after three months

Compare three numbers against the subscription or API you'd otherwise use:

  • Total cost including hours: hardware or rental plus upkeep time at a realistic hourly figure.
  • Error rate: the share of a weekly 20-item sample you had to correct, next to the paid alternative's.
  • Reason still valid: is the privacy, offline or volume reason that started this still true?

If the cost including hours is higher, the error rate is worse and the reason was mainly "it's free", switch back without guilt. If the reason was a hard privacy requirement, the higher cost may simply be the price of meeting it, and that's a legitimate business decision rather than a failed experiment.

Open-source AI costs: follow-up questions

Is a hosted open model more private than ChatGPT or Claude?

Not automatically. When a cloud provider runs the open model for you, your data still goes to that provider, so the same questions apply: training use, retention and a processing agreement. The privacy advantage of open models only appears when the model runs on hardware you control, such as your own computer or a server you manage.

Can an open-source model run on an ordinary office laptop?

Small ones can if the laptop has enough memory. OpenAI's gpt-oss-20b is designed to run within 16GB, though on a 16GB machine it nearly fills memory and replies arrive slowly. Older laptops with 8GB will struggle with anything useful. Try a small model first with a free runner such as LM Studio before buying hardware.

Do open-source licences let me use models commercially?

Many do, but terms differ by model. gpt-oss uses the permissive Apache 2.0 licence. Meta's Llama community licence allows commercial use but requires a separate licence from Meta for services above 700 million monthly users and includes an acceptable use policy. Read the licence for the specific model and version before building on it.

Further reads

Sources: OpenAI gpt-oss announcement and model cards; Meta Llama 4 Community License; Runpod GPU pricing page; Together AI serverless pricing page; Ollama and LM Studio documentation and licence terms; OpenAI API pricing page; checked September 2026.

Wondering if an open model makes sense for you?

On a 1:1 call we'll look at the job you want AI for, the data involved and your volumes, and work out whether a subscription, an API or a model on your own hardware is the sensible route.

Book a 1:1 call with me