Yes. Free tools such as Ollama and LM Studio run open-weight AI models on an ordinary office computer, and nothing typed into a local model leaves the machine. With 16GB of memory you can run smaller models for drafting, summarising and sorting text; long documents, several users or bigger models need 32-64GB or a powerful graphics card.
Local doesn't automatically mean private or cheap, though. Ollama also offers cloud models, whose names end in :cloud, and those run on Ollama's servers, so one careless pick sends prompts out of the building. Expect a small local model to make more mistakes than a paid cloud assistant on the same job, with no web access, and expect to become the person who installs updates, controls access and checks the answers. For most small firms the sensible split is a business plan of a cloud assistant for everyday work, since those don't train on business content by default, plus a local model for the few documents that must never leave your network.
Jobs a model on one office PC handles, and where it falls over
A local model is a file, anything from a few gigabytes to tens of gigabytes, that Ollama or LM Studio loads into the computer's memory. It answers from what it learned in training plus whatever you paste into the conversation. It can't look anything up online, and it knows nothing about your business beyond what you give it. That shapes what it's good for.
| Job | On a 16GB machine? | What to watch |
|---|---|---|
| Summarise a two-page site or meeting note into bullet points | Yes, reliably | Figures copied across wrongly; check every measurement |
| Draft a standard letter from five facts and an example letter | Yes | Stiff tone unless you paste a past letter for it to copy |
| Sort 40 email subject lines into five categories | Yes, in batches | Items at the end of an over-long batch quietly disappear |
| Pull names, dates and addresses out of a form into a table | Usually | Blank fields filled with plausible guesses |
| Answer questions about a 40-page specification | Poorly | The default context window is too small (see the memory section) |
| Anything needing current prices, rules or news | No | No web access; answers come from training data of unknown age |
| Five staff using it at the same moment | Slowly | Test with two people typing at once before promising it to everyone |
The pattern is simple: short, self-contained text jobs work, and anything that needs a long document, fresh facts or reasoning across many pages is where the gap with paid cloud assistants shows. If the privacy question matters more to you than the practical one, the comparison of cloud AI and on-device AI for business data sets out where each option actually keeps your information.
Memory decides which models will load, not processor speed
The whole model has to sit in memory while it runs. On a Windows PC the fastest place for it is the graphics card's own memory (VRAM); when the model doesn't fit there, part of it runs on the main processor instead and answers slow down badly. On Apple Silicon Macs the processor and graphics share one pool of "unified memory", so most of the machine's memory is available to the model. Apple's current Mac mini, for instance, can be configured with anything from 16GB to 64GB.
A rule of thumb that matches the published figures: the computer needs more memory than the model's download size, plus a few gigabytes for the conversation and the operating system. The Ollama library page for OpenAI's gpt-oss says the smaller model, gpt-oss-20b, a 14GB download, runs on systems with as little as 16GB of memory, while the larger gpt-oss-120b is a 65GB download designed to fit on a single 80GB graphics card.
| Memory in the machine | Models that fit (Ollama library, September 2026) | Sensible use |
|---|---|---|
| 8GB | Only the smallest. LM Studio says 8GB Macs can work if you "stick to smaller models and modest context sizes" | Curiosity, not daily work |
| 16GB | Downloads up to roughly 10-14GB, such as Gemma 4's e4b version (9.6GB) or gpt-oss-20b (14GB) | One person, short documents |
| 32GB | Downloads around 19-20GB, such as Gemma 4's 26b and 31b versions | Better drafting, longer notes, light sharing |
| 64GB or more | Larger models, or medium models with a much longer context | A shared office machine with someone looking after it |
Memory also sets how much text the model can consider at once, called its context window and measured in tokens (fragments of words). Ollama's documentation sets the default by graphics memory: 4k tokens below 24GB, 32k from 24GB to 48GB, and 256k at 48GB or more. A quick sum shows why that matters. 4,000 tokens is roughly 3,000 English words, and that allowance has to hold your instructions, the document and the answer together. A 5,000-word inspection report won't fit, and the model won't necessarily tell you it only read part of it. You can raise the limit in the settings, but only as far as the memory allows.
The minimum specifications on the vendors' own pages are modest. LM Studio needs an Apple Silicon Mac on macOS 14 or newer (Intel Macs aren't supported), or a Windows PC whose processor supports AVX2; on Windows it recommends at least 16GB of RAM and 4GB of dedicated graphics memory. Ollama on Windows needs Windows 10 22H2 or newer and 4GB of disk for the program itself, and its documentation warns that models "can be tens to hundreds of GB in size", so check free disk space before downloading anything. Buying a new laptop purely for this is rarely the first step; the tutorial on whether you need a Copilot+ PC or AI laptop explains what those badges do and don't buy you.
Ollama or LM Studio for your first test
Both are free to run on your own hardware, both work offline once a model is downloaded, and both can serve a model to other software. They differ in who they suit.
| LM Studio | Ollama | |
|---|---|---|
| Cost for business use | Free for use at work since July 2025; paid Teams and Enterprise plans add sharing and single sign-on | Free and MIT-licensed; its pricing page says running on your own hardware "is always unlimited" |
| How you use it | Desktop app with a chat window, a model catalogue and document chat | Its own app or the command line; many other tools plug into it |
| Where your text goes | Its docs say nothing entered in chats leaves the device; the internet is used for model search, downloads and updates | Local models stay on the machine; models ending :cloud run on Ollama's servers |
| Sharing with colleagues | Can run a local server on your network | Listens only on the computer itself unless you change the OLLAMA_HOST setting |
| Best first user | A non-technical owner testing on one laptop | Whoever in the office is comfortable with settings, or feeding an automation tool |
For a non-technical owner, LM Studio is the easier first test: install it, pick a model from its catalogue, and chat. Ollama suits the person who'll wire the model into other software later, such as an automation platform or a self-hosted chat front end. Ollama's paid plans ($20 and $100 a month for individuals, $500 a month for a team) buy usage of its cloud models, not permission to run models locally, so nobody needs a subscription to try this.
An afternoon trial on a surveying firm's spare laptop
Here's how a five-person building surveying firm might test the idea before spending anything. The goal: turn surveyors' typed site notes into a first draft of the defects section of a report, without client addresses going to any outside service. The spare laptop has 16GB of memory and no separate graphics card.
- Check the machine (10 minutes). On Windows, Settings, System, About shows installed RAM; on a Mac, About This Mac shows memory. Check free disk space too. Around 30GB spare leaves room for a couple of models.
- Install LM Studio (15 minutes). It installs like any other desktop app.
- Download one model sized for the machine (20-40 minutes, mostly waiting). With 16GB, choose a download under about 10GB so there's room for the notes and the answer. Write down the exact model name and version so you can compare against it later.
- Pick ten old site notes that already have finished reports (20 minutes). Old jobs matter because you already know the right answer, so you can score each draft honestly.
- Run the same prompt on all ten and score every draft (2 hours). Record defects found against defects in the finished report, anything invented, and the minutes needed to fix the draft.
- Write down a verdict (15 minutes). Keep it, drop it, or test a bigger model on a machine with more memory.
Illustrative results from a trial like that:
- Eight of the ten notes were under 2,000 words. On those, the model found 58 of the 61 defects in the finished reports, invented one, and each draft took about 7 minutes to fix, against roughly 25 minutes to write the section from scratch.
- The two long notes, at 3,400 and 4,100 words, overflowed the 4k-token default once the instructions were added. Both drafts stopped listing defects about two-thirds of the way round the building.
- Each draft took one to two minutes to appear on the laptop, where a cloud assistant answers in seconds.
The sum that decides it: saving 18 minutes on each of about 30 reports a month is 9 hours a month, from a free app on a laptop the firm already owns. The two failures mark the limit. Longer notes need either a machine with more memory and a raised context setting, or notes split floor by floor before they go in.
A prompt for local drafting, and what the model got wrong
Small models follow tight instructions far better than open-ended ones. This is the kind of prompt the surveying firm would use, followed by an illustrative draft from a small local model with the problems marked.
You are drafting the defects section of a building survey report.
Use ONLY the site notes below. For each defect, give one line:
Number | Location | Defect | Words from the notes | Suggested next step
If the notes don't say something, write "not stated".
Do not add defects that are not in the notes.
At the end, write: "Defects listed: N".
SITE NOTES:
[paste the notes here]
Illustrative output:
1 | Rear elevation | Hairline crack in render below kitchen window | "hairline crack approx 400mm below kitchen cill" | Monitor, re-inspect in 6 months
2 | Loft | Insulation missing at eaves | "insulation patchy at eaves, both sides" | Top up to current standard
3 | Chimney stack | Lead flashing lifted | "flashing lifted NE corner" | Replace lintel
4 | Bathroom | Extractor fan not working | not stated | Repair or replace fan
Defects listed: 4
What a reviewer should catch:
- Line 3's next step is nonsense. Nothing in the notes mentions a lintel, and a lintel has nothing to do with flashing. The model produced a plausible-sounding survey phrase.
- Line 4 is invented outright. The evidence column says "not stated", which means the notes never mention a bathroom fan. Keeping that column is what exposed it; delete the line.
- Line 2's recommendation belongs to the surveyor. "Current standard" is a professional judgement the model can't make from the notes.
- Four listed, nine in the notes. The count line makes the gap obvious in seconds. Here the notes had been cut off by the context window, so the last five defects were never read.
The fixed version of the process: split long notes, raise the context length only as far as memory allows, keep the evidence column, and drop "Suggested next step" from the prompt so the surveyor writes that part. The draft then does the transcription work and the professional does the judging.
Letting the whole office use one machine without opening a hole
Once one person likes it, others will want it, and sharing is where most of the security work sits, because the defaults assume one user on one computer.
- Know what the network setting does. Ollama listens only on the computer itself by default (address 127.0.0.1, port 11434). Setting OLLAMA_HOST to 0.0.0.0 makes it answer anyone on the network, and Ollama's own documentation states that the local API "does not require authentication". Anyone on the office Wi-Fi, a visitor included, could use it.
- Put logins in front of it. If several people will use one model, add a front end with its own accounts. Open WebUI is one self-hosted option that works with Ollama and can run entirely offline; LM Studio's Enterprise plan adds single sign-on and controls over which models staff can load.
- Never expose it to the internet. Don't forward the port on your router. If staff need it off-site, route them through a VPN rather than opening the machine up.
- Treat the chat history as client data. Conversations are stored on the machine, so switch on disk encryption, require a sign-in, and decide how long history is kept.
A short written rule sheet stops the setup drifting. Here's a filled-in version for an eight-person engineering consultancy:
LOCAL AI HOUSE RULES: [consultancy name]
Machine: office desktop "AI-01", 64GB memory, locked comms cupboard
Looked after by: [first name], 1 hour on the first Monday of each month
Software: Ollama with Open WebUI; staff sign in with their own account
Allowed: project specifications, client correspondence and calc
summaries for projects under NDA
Not allowed: personal data beyond names and job titles; anything
marked "client eyes only"
Models: [model name, version, date downloaded]
No model whose name ends in ":cloud"
Chat history: deleted after 90 days; export first if it belongs on file
Checks: every output used in a deliverable is reviewed by the
engineer who signs that deliverable
Quarterly: rerun the 10 sample documents on any newer model
The sum: one shared AI machine against five cloud seats
Cost is where local AI surprises people. Here's the comparison for a five-person architect practice. The licence figures are list prices; the hardware and time figures are assumptions to replace with your own quotes.
| Cloud business plan | Shared local machine | |
|---|---|---|
| Licences | 5 seats at $20 a month billed annually (ChatGPT Business Standard or Claude Team Standard): $1,200 a year | $0; LM Studio and Ollama are free for work |
| Hardware | None | Assume $2,500 for a desktop with 64GB of memory |
| Setup | About an hour to invite users and set sharing | Assume 8 hours of your most technical person at $50 an hour: $400 |
| Upkeep | Minutes a month | 1 hour a month: $600 a year |
| First-year total | $1,200 | $3,500 |
| Three-year total | $3,600 | $4,700 |
On those assumptions the local machine costs more over three years and gives the practice a less capable assistant. Monthly billing narrows the gap (at $25 a seat the cloud plan is $1,500 a year), and a practice that already owns a powerful rendering workstation saves the hardware line entirely. Even so, cost alone rarely justifies local AI for a small team.
High volume usually doesn't change that. Suppose the practice wanted 10,000 archived documents of about 1,000 words each sorted into project types. That's roughly 13 million tokens of input. On OpenAI's API, gpt-5-nano is listed at $0.05 per million input tokens and $0.40 per million output tokens, so the input costs about 67 cents and a short label for each document adds about 40 cents: around a dollar for the lot. API use is billed separately from any ChatGPT subscription, and OpenAI doesn't train on API data by default. The real reason to run that job locally is confidentiality, not the bill. The true cost of open-source AI for a small business goes through the staff-time side in more detail.
Engineering firm, accountancy, kitchen fitter: who should bother
Strong case: the engineering consultancy. A structural engineering consultancy whose NDAs forbid sharing drawings and specifications with third parties has a genuine reason to keep AI in-house, because a cloud AI provider is a third party even on a business plan. Check that reading with whoever drafted the NDA, then set up one well-specified machine for NDA projects only, used for summarising specifications and drafting responses to requests for information (RFIs). Everything else can stay on a normal cloud plan.
Narrow case: the accountancy practice. A six-person accountancy practice that already keeps client records in cloud bookkeeping software adds less new exposure with a business AI plan than it first seems, as long as that plan doesn't train on client content. Local AI makes sense for a defined set of files the partners have decided never go to an outside AI service, not as the practice's main assistant. Deciding what's in that set is a data-classification job, and how to classify business data before using AI tools walks through sorting files by sensitivity.
No case: the kitchen fitter. A two-person kitchen fitting business writing quotes, customer updates and social posts has nothing sensitive enough to justify an afternoon of setup and a slower, weaker assistant. An individual cloud plan, with model training switched off in its privacy settings, does the job.
Signs a local setup is quietly letting you down
- Lists that stop early, or summaries that ignore the last pages. The document is bigger than the context window. Split it, or raise the limit if memory allows.
- Answers that take minutes. In Ollama, run
ollama ps: its Processor column shows whether the model loaded as "100% GPU", "100% CPU" or a split between them. A big CPU share means the model is too large for the graphics card; try a smaller one. - A model name ending in
:cloudin someone's history. That conversation ran on Ollama's servers. Ollama says it doesn't store cloud prompts or train on them, but the text left the building, which breaks any rule that said those documents stay in the office. - Nobody has touched the setup in a year. The app needs updates like any other software, and newer models may do your job better.
One realistic way the context problem shows up: a trainee at a small accountancy practice asks a local model to summarise a 30-page engagement letter. The summary reads well and leaves out the limitation-of-liability clause near the end, because the model only ever saw the first two-thirds of the text. Nobody notices until a partner compares it with the original. The fix is procedural as much as technical. Long documents get split before they go in, and any summary of a contract is checked against the contract's list of clauses before anyone relies on it.
Four questions to answer before buying any hardware
- Is there a document type you're barred from putting into a cloud AI service, by contract or by your own policy, even on a business plan that doesn't train on your content? If not, a business cloud plan is the simpler answer and you can stop here.
- Are the jobs short, under about 3,000 words at a time? Then test LM Studio this week on a 16GB machine you already own, using the ten-document trial above.
- Will several people use it daily, or will documents run past 20 pages? Then budget for a 32-64GB machine, a front end with logins and an hour a month of upkeep, and rerun the cost sum with your own quotes.
- Can you name the person who'll maintain it? If not, don't start. An unmaintained AI box with no logins on the office network is worse than no box at all.
If the first answer is yes, the next choice is which model family to trust with those documents, and the licences differ more than most owners expect; open-source versus paid AI models compares them side by side.
Local AI questions owners ask before installing anything
Does a local model learn from what my staff type into it?
No. Chatting with a model doesn't change the model file; it answers from its training plus whatever is in the conversation. The chats themselves are saved on the computer by the app, though, so treat that machine like any other that holds client files: sign-in required, disk encryption on, and a rule for how long chat history is kept before someone deletes it.
Will a local AI model work with no internet connection?
Yes, once the model is downloaded. LM Studio's documentation says chatting and chatting with documents both work offline; you only need a connection to search for and download models, fetch runtimes and check for app updates. That makes local AI usable on a site with poor signal, as long as the laptop already has the model on it.
Can a local model read our PDFs and spreadsheets?
LM Studio can chat with documents you attach, and its documentation says that processing happens on your machine. Scanned PDFs with no text layer and complex tables are where small models most often slip, so test a few real files first. For spreadsheets, paste the relevant rows as text rather than handing over the whole workbook.
What happens to our local AI when a better model comes out?
Nothing, until someone downloads it. Local models don't update themselves: the file stays exactly as it was. Put a quarterly reminder in the calendar to check the model library in Ollama or LM Studio, rerun your ten sample documents on any promising new model, and switch only if it scores better on your own work.
Further reads
- Is DeepSeek Safe to Use With Business Data? — DeepSeek's open models can run locally; its app is a different story.
- How to Stop AI Tools Training on Your Business Data — If you stay in the cloud, switch model training off first.
- Shadow AI: How to Stop Staff Pasting Client Data Into Free Tools — Stop staff pasting client files into free tools meanwhile.
- How to Run a Two-Week AI Tool Trial Before You Commit — A two-week trial plan for any AI tool, local or cloud.
- Which AI Is Best for a Small Business? — Choosing the everyday cloud assistant most firms still need.
- How to Keep Customer Data Private When Your Team Uses AI — Wider rules for keeping customer data private when staff use AI.
- How Much Does AI Visual Inspection Cost a Small Manufacturer? — Real price bands for one AI inspection station, the line items quotes leave out, and a worked payback sum for a small moulding shop.
- AI Tools and AI Development: The Complete 2026 Guide — the AI hub, including every tutorial in the AI-for-business series.
Sources: LM Studio documentation (system requirements, offline operation) and its July 2025 announcement on free use at work; Ollama documentation (Windows requirements, FAQ, context length, cloud models, API authentication), Ollama pricing page and its model library pages for Gemma 4 and gpt-oss; Apple's Mac mini product page; Open WebUI documentation; ChatGPT Business, Claude Team and OpenAI API pricing as of September 2026.