Can a Small Business Run AI on Its Own Computers?

Coding Liquids tutorial cover featuring Sagnik Bhattacharya for Can a Small Business Run AI on Its Own Computers?
Coding Liquids tutorial cover featuring Sagnik Bhattacharya for Can a Small Business Run AI on Its Own Computers?

Yes. Free tools such as Ollama and LM Studio run open-weight AI models on an ordinary office computer, and nothing typed into a local model leaves the machine. With 16GB of memory you can run smaller models for drafting, summarising and sorting text; long documents, several users or bigger models need 32-64GB or a powerful graphics card.

Local doesn't automatically mean private or cheap, though. Ollama also offers cloud models, whose names end in :cloud, and those run on Ollama's servers, so one careless pick sends prompts out of the building. Expect a small local model to make more mistakes than a paid cloud assistant on the same job, with no web access, and expect to become the person who installs updates, controls access and checks the answers. For most small firms the sensible split is a business plan of a cloud assistant for everyday work, since those don't train on business content by default, plus a local model for the few documents that must never leave your network.

Follow me on Instagram@sagnikteaches

Jobs a model on one office PC handles, and where it falls over

A local model is a file, anything from a few gigabytes to tens of gigabytes, that Ollama or LM Studio loads into the computer's memory. It answers from what it learned in training plus whatever you paste into the conversation. It can't look anything up online, and it knows nothing about your business beyond what you give it. That shapes what it's good for.

Connect on LinkedInSagnik Bhattacharya
JobOn a 16GB machine?What to watch
Summarise a two-page site or meeting note into bullet pointsYes, reliablyFigures copied across wrongly; check every measurement
Draft a standard letter from five facts and an example letterYesStiff tone unless you paste a past letter for it to copy
Sort 40 email subject lines into five categoriesYes, in batchesItems at the end of an over-long batch quietly disappear
Pull names, dates and addresses out of a form into a tableUsuallyBlank fields filled with plausible guesses
Answer questions about a 40-page specificationPoorlyThe default context window is too small (see the memory section)
Anything needing current prices, rules or newsNoNo web access; answers come from training data of unknown age
Five staff using it at the same momentSlowlyTest with two people typing at once before promising it to everyone

The pattern is simple: short, self-contained text jobs work, and anything that needs a long document, fresh facts or reasoning across many pages is where the gap with paid cloud assistants shows. If the privacy question matters more to you than the practical one, the comparison of cloud AI and on-device AI for business data sets out where each option actually keeps your information.

Subscribe on YouTube@codingliquids

Memory decides which models will load, not processor speed

The whole model has to sit in memory while it runs. On a Windows PC the fastest place for it is the graphics card's own memory (VRAM); when the model doesn't fit there, part of it runs on the main processor instead and answers slow down badly. On Apple Silicon Macs the processor and graphics share one pool of "unified memory", so most of the machine's memory is available to the model. Apple's current Mac mini, for instance, can be configured with anything from 16GB to 64GB.

A rule of thumb that matches the published figures: the computer needs more memory than the model's download size, plus a few gigabytes for the conversation and the operating system. The Ollama library page for OpenAI's gpt-oss says the smaller model, gpt-oss-20b, a 14GB download, runs on systems with as little as 16GB of memory, while the larger gpt-oss-120b is a 65GB download designed to fit on a single 80GB graphics card.

Memory in the machineModels that fit (Ollama library, September 2026)Sensible use
8GBOnly the smallest. LM Studio says 8GB Macs can work if you "stick to smaller models and modest context sizes"Curiosity, not daily work
16GBDownloads up to roughly 10-14GB, such as Gemma 4's e4b version (9.6GB) or gpt-oss-20b (14GB)One person, short documents
32GBDownloads around 19-20GB, such as Gemma 4's 26b and 31b versionsBetter drafting, longer notes, light sharing
64GB or moreLarger models, or medium models with a much longer contextA shared office machine with someone looking after it

Memory also sets how much text the model can consider at once, called its context window and measured in tokens (fragments of words). Ollama's documentation sets the default by graphics memory: 4k tokens below 24GB, 32k from 24GB to 48GB, and 256k at 48GB or more. A quick sum shows why that matters. 4,000 tokens is roughly 3,000 English words, and that allowance has to hold your instructions, the document and the answer together. A 5,000-word inspection report won't fit, and the model won't necessarily tell you it only read part of it. You can raise the limit in the settings, but only as far as the memory allows.

The minimum specifications on the vendors' own pages are modest. LM Studio needs an Apple Silicon Mac on macOS 14 or newer (Intel Macs aren't supported), or a Windows PC whose processor supports AVX2; on Windows it recommends at least 16GB of RAM and 4GB of dedicated graphics memory. Ollama on Windows needs Windows 10 22H2 or newer and 4GB of disk for the program itself, and its documentation warns that models "can be tens to hundreds of GB in size", so check free disk space before downloading anything. Buying a new laptop purely for this is rarely the first step; the tutorial on whether you need a Copilot+ PC or AI laptop explains what those badges do and don't buy you.

Ollama or LM Studio for your first test

Both are free to run on your own hardware, both work offline once a model is downloaded, and both can serve a model to other software. They differ in who they suit.

LM StudioOllama
Cost for business useFree for use at work since July 2025; paid Teams and Enterprise plans add sharing and single sign-onFree and MIT-licensed; its pricing page says running on your own hardware "is always unlimited"
How you use itDesktop app with a chat window, a model catalogue and document chatIts own app or the command line; many other tools plug into it
Where your text goesIts docs say nothing entered in chats leaves the device; the internet is used for model search, downloads and updatesLocal models stay on the machine; models ending :cloud run on Ollama's servers
Sharing with colleaguesCan run a local server on your networkListens only on the computer itself unless you change the OLLAMA_HOST setting
Best first userA non-technical owner testing on one laptopWhoever in the office is comfortable with settings, or feeding an automation tool

For a non-technical owner, LM Studio is the easier first test: install it, pick a model from its catalogue, and chat. Ollama suits the person who'll wire the model into other software later, such as an automation platform or a self-hosted chat front end. Ollama's paid plans ($20 and $100 a month for individuals, $500 a month for a team) buy usage of its cloud models, not permission to run models locally, so nobody needs a subscription to try this.

An afternoon trial on a surveying firm's spare laptop

Here's how a five-person building surveying firm might test the idea before spending anything. The goal: turn surveyors' typed site notes into a first draft of the defects section of a report, without client addresses going to any outside service. The spare laptop has 16GB of memory and no separate graphics card.

  1. Check the machine (10 minutes). On Windows, Settings, System, About shows installed RAM; on a Mac, About This Mac shows memory. Check free disk space too. Around 30GB spare leaves room for a couple of models.
  2. Install LM Studio (15 minutes). It installs like any other desktop app.
  3. Download one model sized for the machine (20-40 minutes, mostly waiting). With 16GB, choose a download under about 10GB so there's room for the notes and the answer. Write down the exact model name and version so you can compare against it later.
  4. Pick ten old site notes that already have finished reports (20 minutes). Old jobs matter because you already know the right answer, so you can score each draft honestly.
  5. Run the same prompt on all ten and score every draft (2 hours). Record defects found against defects in the finished report, anything invented, and the minutes needed to fix the draft.
  6. Write down a verdict (15 minutes). Keep it, drop it, or test a bigger model on a machine with more memory.

Illustrative results from a trial like that:

  • Eight of the ten notes were under 2,000 words. On those, the model found 58 of the 61 defects in the finished reports, invented one, and each draft took about 7 minutes to fix, against roughly 25 minutes to write the section from scratch.
  • The two long notes, at 3,400 and 4,100 words, overflowed the 4k-token default once the instructions were added. Both drafts stopped listing defects about two-thirds of the way round the building.
  • Each draft took one to two minutes to appear on the laptop, where a cloud assistant answers in seconds.

The sum that decides it: saving 18 minutes on each of about 30 reports a month is 9 hours a month, from a free app on a laptop the firm already owns. The two failures mark the limit. Longer notes need either a machine with more memory and a raised context setting, or notes split floor by floor before they go in.

A prompt for local drafting, and what the model got wrong

Small models follow tight instructions far better than open-ended ones. This is the kind of prompt the surveying firm would use, followed by an illustrative draft from a small local model with the problems marked.

You are drafting the defects section of a building survey report.
Use ONLY the site notes below. For each defect, give one line:
Number | Location | Defect | Words from the notes | Suggested next step
If the notes don't say something, write "not stated".
Do not add defects that are not in the notes.
At the end, write: "Defects listed: N".

SITE NOTES:
[paste the notes here]

Illustrative output:

1 | Rear elevation | Hairline crack in render below kitchen window | "hairline crack approx 400mm below kitchen cill" | Monitor, re-inspect in 6 months
2 | Loft | Insulation missing at eaves | "insulation patchy at eaves, both sides" | Top up to current standard
3 | Chimney stack | Lead flashing lifted | "flashing lifted NE corner" | Replace lintel
4 | Bathroom | Extractor fan not working | not stated | Repair or replace fan
Defects listed: 4

What a reviewer should catch:

  • Line 3's next step is nonsense. Nothing in the notes mentions a lintel, and a lintel has nothing to do with flashing. The model produced a plausible-sounding survey phrase.
  • Line 4 is invented outright. The evidence column says "not stated", which means the notes never mention a bathroom fan. Keeping that column is what exposed it; delete the line.
  • Line 2's recommendation belongs to the surveyor. "Current standard" is a professional judgement the model can't make from the notes.
  • Four listed, nine in the notes. The count line makes the gap obvious in seconds. Here the notes had been cut off by the context window, so the last five defects were never read.

The fixed version of the process: split long notes, raise the context length only as far as memory allows, keep the evidence column, and drop "Suggested next step" from the prompt so the surveyor writes that part. The draft then does the transcription work and the professional does the judging.

Letting the whole office use one machine without opening a hole

Once one person likes it, others will want it, and sharing is where most of the security work sits, because the defaults assume one user on one computer.

  • Know what the network setting does. Ollama listens only on the computer itself by default (address 127.0.0.1, port 11434). Setting OLLAMA_HOST to 0.0.0.0 makes it answer anyone on the network, and Ollama's own documentation states that the local API "does not require authentication". Anyone on the office Wi-Fi, a visitor included, could use it.
  • Put logins in front of it. If several people will use one model, add a front end with its own accounts. Open WebUI is one self-hosted option that works with Ollama and can run entirely offline; LM Studio's Enterprise plan adds single sign-on and controls over which models staff can load.
  • Never expose it to the internet. Don't forward the port on your router. If staff need it off-site, route them through a VPN rather than opening the machine up.
  • Treat the chat history as client data. Conversations are stored on the machine, so switch on disk encryption, require a sign-in, and decide how long history is kept.

A short written rule sheet stops the setup drifting. Here's a filled-in version for an eight-person engineering consultancy:

LOCAL AI HOUSE RULES: [consultancy name]
Machine:        office desktop "AI-01", 64GB memory, locked comms cupboard
Looked after by: [first name], 1 hour on the first Monday of each month
Software:       Ollama with Open WebUI; staff sign in with their own account
Allowed:        project specifications, client correspondence and calc
                summaries for projects under NDA
Not allowed:    personal data beyond names and job titles; anything
                marked "client eyes only"
Models:         [model name, version, date downloaded]
                No model whose name ends in ":cloud"
Chat history:   deleted after 90 days; export first if it belongs on file
Checks:         every output used in a deliverable is reviewed by the
                engineer who signs that deliverable
Quarterly:      rerun the 10 sample documents on any newer model

The sum: one shared AI machine against five cloud seats

Cost is where local AI surprises people. Here's the comparison for a five-person architect practice. The licence figures are list prices; the hardware and time figures are assumptions to replace with your own quotes.

Cloud business planShared local machine
Licences5 seats at $20 a month billed annually (ChatGPT Business Standard or Claude Team Standard): $1,200 a year$0; LM Studio and Ollama are free for work
HardwareNoneAssume $2,500 for a desktop with 64GB of memory
SetupAbout an hour to invite users and set sharingAssume 8 hours of your most technical person at $50 an hour: $400
UpkeepMinutes a month1 hour a month: $600 a year
First-year total$1,200$3,500
Three-year total$3,600$4,700

On those assumptions the local machine costs more over three years and gives the practice a less capable assistant. Monthly billing narrows the gap (at $25 a seat the cloud plan is $1,500 a year), and a practice that already owns a powerful rendering workstation saves the hardware line entirely. Even so, cost alone rarely justifies local AI for a small team.

High volume usually doesn't change that. Suppose the practice wanted 10,000 archived documents of about 1,000 words each sorted into project types. That's roughly 13 million tokens of input. On OpenAI's API, gpt-5-nano is listed at $0.05 per million input tokens and $0.40 per million output tokens, so the input costs about 67 cents and a short label for each document adds about 40 cents: around a dollar for the lot. API use is billed separately from any ChatGPT subscription, and OpenAI doesn't train on API data by default. The real reason to run that job locally is confidentiality, not the bill. The true cost of open-source AI for a small business goes through the staff-time side in more detail.

Engineering firm, accountancy, kitchen fitter: who should bother

Strong case: the engineering consultancy. A structural engineering consultancy whose NDAs forbid sharing drawings and specifications with third parties has a genuine reason to keep AI in-house, because a cloud AI provider is a third party even on a business plan. Check that reading with whoever drafted the NDA, then set up one well-specified machine for NDA projects only, used for summarising specifications and drafting responses to requests for information (RFIs). Everything else can stay on a normal cloud plan.

Narrow case: the accountancy practice. A six-person accountancy practice that already keeps client records in cloud bookkeeping software adds less new exposure with a business AI plan than it first seems, as long as that plan doesn't train on client content. Local AI makes sense for a defined set of files the partners have decided never go to an outside AI service, not as the practice's main assistant. Deciding what's in that set is a data-classification job, and how to classify business data before using AI tools walks through sorting files by sensitivity.

No case: the kitchen fitter. A two-person kitchen fitting business writing quotes, customer updates and social posts has nothing sensitive enough to justify an afternoon of setup and a slower, weaker assistant. An individual cloud plan, with model training switched off in its privacy settings, does the job.

Signs a local setup is quietly letting you down

  • Lists that stop early, or summaries that ignore the last pages. The document is bigger than the context window. Split it, or raise the limit if memory allows.
  • Answers that take minutes. In Ollama, run ollama ps: its Processor column shows whether the model loaded as "100% GPU", "100% CPU" or a split between them. A big CPU share means the model is too large for the graphics card; try a smaller one.
  • A model name ending in :cloud in someone's history. That conversation ran on Ollama's servers. Ollama says it doesn't store cloud prompts or train on them, but the text left the building, which breaks any rule that said those documents stay in the office.
  • Nobody has touched the setup in a year. The app needs updates like any other software, and newer models may do your job better.

One realistic way the context problem shows up: a trainee at a small accountancy practice asks a local model to summarise a 30-page engagement letter. The summary reads well and leaves out the limitation-of-liability clause near the end, because the model only ever saw the first two-thirds of the text. Nobody notices until a partner compares it with the original. The fix is procedural as much as technical. Long documents get split before they go in, and any summary of a contract is checked against the contract's list of clauses before anyone relies on it.

Four questions to answer before buying any hardware

  1. Is there a document type you're barred from putting into a cloud AI service, by contract or by your own policy, even on a business plan that doesn't train on your content? If not, a business cloud plan is the simpler answer and you can stop here.
  2. Are the jobs short, under about 3,000 words at a time? Then test LM Studio this week on a 16GB machine you already own, using the ten-document trial above.
  3. Will several people use it daily, or will documents run past 20 pages? Then budget for a 32-64GB machine, a front end with logins and an hour a month of upkeep, and rerun the cost sum with your own quotes.
  4. Can you name the person who'll maintain it? If not, don't start. An unmaintained AI box with no logins on the office network is worse than no box at all.

If the first answer is yes, the next choice is which model family to trust with those documents, and the licences differ more than most owners expect; open-source versus paid AI models compares them side by side.

Local AI questions owners ask before installing anything

Does a local model learn from what my staff type into it?

No. Chatting with a model doesn't change the model file; it answers from its training plus whatever is in the conversation. The chats themselves are saved on the computer by the app, though, so treat that machine like any other that holds client files: sign-in required, disk encryption on, and a rule for how long chat history is kept before someone deletes it.

Will a local AI model work with no internet connection?

Yes, once the model is downloaded. LM Studio's documentation says chatting and chatting with documents both work offline; you only need a connection to search for and download models, fetch runtimes and check for app updates. That makes local AI usable on a site with poor signal, as long as the laptop already has the model on it.

Can a local model read our PDFs and spreadsheets?

LM Studio can chat with documents you attach, and its documentation says that processing happens on your machine. Scanned PDFs with no text layer and complex tables are where small models most often slip, so test a few real files first. For spreadsheets, paste the relevant rows as text rather than handing over the whole workbook.

What happens to our local AI when a better model comes out?

Nothing, until someone downloads it. Local models don't update themselves: the file stays exactly as it was. Put a quarterly reminder in the calendar to check the model library in Ollama or LM Studio, rerun your ten sample documents on any promising new model, and switch only if it scores better on your own work.

Further reads

Sources: LM Studio documentation (system requirements, offline operation) and its July 2025 announcement on free use at work; Ollama documentation (Windows requirements, FAQ, context length, cloud models, API authentication), Ollama pricing page and its model library pages for Gemma 4 and gpt-oss; Apple's Mac mini product page; Open WebUI documentation; ChatGPT Business, Claude Team and OpenAI API pricing as of September 2026.

Not sure local AI is worth it for your firm?

On a 1:1 call we'll pick out the documents that genuinely can't go to a cloud AI, check whether a local model handles them well enough, and compare that with a business plan.

Book a 1:1 call with me