Should a Small Business Build a Custom AI Tool or Buy One?

Coding Liquids tutorial cover featuring Sagnik Bhattacharya for Should a Small Business Build a Custom AI Tool or Buy One?
Coding Liquids tutorial cover featuring Sagnik Bhattacharya for Should a Small Business Build a Custom AI Tool or Buy One?

Buy, in most cases. A small business should buy an AI tool, or switch on the AI in software it already has, unless three things are true: the task is central to how it earns money, no product handles it after a trial on its own data, and there's budget to maintain a build, not only to make it.

Most build-or-buy debates also skip the option that usually wins: configure. That means buying general tools and shaping them to your process with shared instructions, a Project in ChatGPT or Claude that holds your templates, or a no-code automation in Zapier or Make. You get most of a custom tool's fit, you own the instructions, and someone else maintains the software underneath.

Follow me on Instagram@sagnikteaches

Buy, configure or build: what each one means in practice

The three routes differ less in what they can do than in who carries the work after launch. That's the part owners underestimate.

Connect on LinkedInSagnik Bhattacharya
RouteWhat you pay forWho keeps it runningTime to first useWhat you own at the end
BuyA product with AI inside: a helpdesk with an AI agent, a note-taker, sector software that drafts documentsThe vendorDaysYour data, if the contract says so; not the tool
ConfigureGeneral AI tools plus your instructions, templates, example outputs and no-code automationsYou or a consultant, on the vendor's platformDays to a few weeksInstructions, templates and workflow designs
BuildSoftware written for you that calls an AI model through its API (the connection developers use to send it work)You, through a developerWeeks to monthsThe code, if the contract assigns it to you

One billing detail catches people out on the build route. A ChatGPT Plus or Claude Pro subscription doesn't cover API use; the API is billed separately, per token (roughly, per chunk of words sent and received), on the vendor's developer platform. So a custom build adds a new bill even if the whole team already has chat subscriptions.

Subscribe on YouTube@codingliquids

Seven questions that settle build versus buy

Answer these for the specific task, not for "AI" in general. The same business can sensibly buy one tool and build another.

QuestionAnswer that points to buying or configuringAnswer that points to building
1. Is the task done much the same way in every business like yours?Yes: receipts, meeting notes, customer FAQs, schedulingNo: your method is unusual and deliberate
2. Does doing it your way win you work, funding or members?No, customers wouldn't noticeYes, it's why people choose you
3. Did a product pass a two-week trial on your own data?At least one did, with edits you can live withEvery candidate failed on your real cases
4. How many of your systems must it connect to?One or two, with ready-made integrationsSeveral, including an in-house database nothing else connects to
5. Who fixes it at 8am when it breaks?Nobody in-houseA developer on a retainer or a capable staff member
6. Where is the data allowed to go?A business-plan product with clear terms is acceptableIt must stay inside systems you control
7. How many people or runs a month?A handful of users, modest volumeDozens of users or thousands of runs, where per-seat fees balloon

Count the answers in the right-hand column. Fewer than four, and building is a hobby project paid for with your money. Question 5 works as a veto: if nobody can fix the tool within a working day, don't build, whatever the other six say.

A community interest company weighs all three for funder reports

Consider a community interest company (a social enterprise whose profits go back into its community purpose); the figures are invented but realistic. It has six staff and runs job-readiness programmes for about 180 people a year. Three funders each want a quarterly report in their own format: starts, completions and outcomes, plus three or four anonymised participant stories. Two coordinators spend about 60 hours a quarter between them reading case notes, pulling numbers from a spreadsheet and writing. The directors value staff time at $20 an hour.

The three options on the table

  • Buy: a case-management platform whose AI drafts funder reports. Its quote is $45 a user a month for all six staff, $270 a month, and moving three years of spreadsheets and notes across would take about 40 hours.
  • Configure: ChatGPT Business for the two coordinators at $25 a seat on monthly billing, with a shared Project holding each funder's template, guidance and two past reports, plus a Make scenario (from about $9 a month) that tidies the monthly sign-up form into a sheet. Setup takes about 25 hours, and upkeep about 2 hours a month.
  • Build: a developer quotes $9,000 for a small web app that reads the case-notes database, removes names from the stories and drafts each funder's report through a model's API. Hosting is about $30 a month. The developer quotes two maintenance days a year at a total of $1,000.

The build's usage cost is tiny, which surprises people. A quarter's notes for three funders might mean around 450,000 tokens in and 25,000 out. At Claude Sonnet 5's API price of $2 per million input tokens and $10 per million output, that's about $1.15 a quarter. The money in a build goes on people, not on the model.

Two years of costs and benefits

Cost line over 24 monthsBuyConfigureBuild
Upfront payment$0$0$9,000
Staff setup time at $20/hour40 hours: $80025 hours: $50010 hours of testing: $200
Subscriptions or hosting$6,480$1,416$720
Model usageIncludedIncludedAbout $10
UpkeepVendor's job48 hours: $960$2,000
Two-year cost$7,280$2,876$11,930
Report hours saved per year160144176
Two-year benefit at $20/hour$6,400$5,760$7,040
Net after two years-$880+$2,884-$4,890

The build fits best, saving the most hours, and loses the most money. The bought platform nearly breaks even on reports alone, and it would come out ahead if the company needs proper case management anyway, which is a records decision rather than an AI decision. Configuring wins clearly: it recovers its 25 hours of setup in about three and a half months and costs well under half of buying.

The data question shaped the configure option too. Case notes hold personal and sometimes sensitive details, so the coordinators remove names before anything goes into the Project, the raw notes stay in the company's own systems, and they use a business plan, which doesn't train on business content by default. A consumer chat account wouldn't have passed question 6.

What a custom build costs after launch day

A build quote covers making the thing. Keeping it working is a separate, open-ended cost, because the models and services underneath keep changing. A few real changes from 2026 show what a maintenance budget has to absorb:

  • Settings stop working. Anthropic is deprecating the temperature setting (a dial for how varied answers are) on Claude 4.7 and later models, and OpenAI's reasoning models don't support it. Code that relied on it needs rework when you move to a newer model.
  • Services close. OpenAI shut the Sora API on 24 September 2026. Anything built on it had to be replaced, at the owner's expense.
  • Promotional prices end. OpenAI lists gpt-5.6-sol at $4 per million input tokens and $20 per million output as a promotional price, at least until 21 November 2026. A cost forecast built on a promotion needs a second line for the full price.
  • Retention terms move. Since 9 June 2026, Anthropic keeps prompts and outputs on its most capable "covered" models for 30 days even for customers with zero-data-retention agreements, with an application route back to zero retention added in September. If your build promised a funder or client that nothing is stored, someone has to notice changes like this.
  • Your other systems change. The database or CRM the tool reads from renames a field or updates its own API, and the tool quietly starts pulling blanks.

Before accepting any build quote, get a separate maintenance quote in days per year, and make sure the code, the hosting account and the API keys sit in accounts your business owns. Who holds what at the end is worth settling in writing; who owns the AI workflows a consultant builds for you covers the clauses. For the longer-range comparison, per-seat fees against one custom tool over five years runs the sums across a longer horizon.

Building on someone else's platform is still renting

The configure route has its own version of this risk, and custom GPTs are the clearest example. In this made-up case, a charity shop set up a custom GPT in 2025 to answer volunteers' questions from its handbook: pricing guidance, what can't be sold, how to log donations. It worked well. Then OpenAI announced that custom GPTs stop running on 11 December 2026. OpenAI's migration turns each GPT into a plugin, with its instructions becoming a skill and its knowledge files becoming reference files, and the shop now has to move it and test the answers again.

The shop that kept a master copy of its instructions, its handbook files and ten test questions in its own shared drive can recreate the assistant as a ChatGPT Project, a Claude Project or a Gemini Gem (becoming a skill from November 2026) in an afternoon. The shop that only ever edited the GPT inside ChatGPT has to reconstruct what it wrote a year ago.

Automation platforms change too, and configured workflows need someone who reads the alerts:

  • Make renamed its error handlers in September 2026 (Ignore became Skip, Break became Retry), so anyone following older notes needs to know the new names.
  • Zapier pauses a Zap automatically when 95% of its runs error over seven days.
  • Power Automate disables a flow once it has failed for 14 consecutive days.

None of that is a reason to avoid configuring. It's a reason to name an owner for every configured tool and to keep its instructions somewhere you control.

Test the bought tools with your own data first

Question 3 in the table only means something if the trial uses your real material. Vendor demos run on tidy sample data; your records are messier. For the community interest company, a fair bake-off takes ten real cases from last quarter, with names removed, including three awkward ones: a participant who left and came back, one whose notes are two lines long, and one with outcomes recorded in two different places.

Here's the test sheet after the first round (illustrative scores out of 5 for accuracy, with editing time in minutes):

CaseBought platformConfigured ProjectNotes
Straightforward completer5, 3 min5, 4 minBoth fine
Left and returned2, 15 min4, 6 minPlatform counted the person twice
Two-line notes3, 8 min3, 9 minBoth padded the story; flag for more notes
Outcome in two places4, 5 min2, 12 minProject missed the job start in the second sheet
Average across all ten3.9, 7 min4.1, 6 minClose; configure wins on cost

Scores this close mean the decision falls back on cost and upkeep. Before paying for a specialist product, it's also worth checking how much of it is a thin layer over the same models you could use directly; how to tell whether an AI tool is a ChatGPT wrapper explains the checks. A wrapper can still be worth buying if its workflow saves you time, but you should know what the margin pays for.

The trial also exposes the failure you most need to catch. A prompt the coordinators used in the configured Project:

Using the Funder B template in this Project, draft the Q3 progress
section from the anonymised notes and the attached figures sheet.
Use only figures that appear in the notes or the sheet. Where a figure
is missing, write [MISSING: what is needed] instead of estimating.

What came back, trimmed (illustrative):

In Q3 the programme supported 47 participants, 31 of whom completed all modules. 85% of completers reported increased confidence in interviews. Participant C secured a warehouse role three weeks after finishing and has since been promoted to team leader.

Two of those sentences are fine: 47 and 31 are in the sheet. The 85% appears nowhere in the data; the model filled a gap it had been told to flag. And "promoted to team leader" came from a different participant's notes. So the configured route kept a human step: a coordinator ticks every figure and every story detail against the source before a report leaves the building. That check would be needed whichever of the three routes they chose.

Turning "we want our own AI" into a spec a developer can price

If the seven questions do point to building, the next risk is a vague brief, which produces a vague quote and a long argument later. Compare two versions of the same request.

Before: "We want an AI that writes our funder reports automatically."

After:

  • Inputs: case notes from our database, about 600 entries a quarter, plus one figures sheet.
  • Outputs: three reports a quarter in the funders' templates, each about two pages.
  • Accuracy rule: every figure traceable to the sheet; no outcome that isn't in the notes; missing data flagged, never estimated.
  • Human step: a coordinator approves each report before it's sent.
  • Data rules: names removed before anything reaches the model; nothing kept by the model provider beyond what its business or API terms state, checked at launch and yearly.
  • Acceptance test: regenerate last year's four quarters; each report needs no more than 15 minutes of edits.
  • Ownership: code in our repository, hosting and API keys in our accounts.
  • Maintenance: quoted separately, in days per year.

The second version can be priced, tested and argued about fairly. For the full briefing process, including what a developer will ask you, see how to brief a developer on a custom AI workflow.

When building genuinely wins

Building earns its cost when several of these line up: the process is how you win work, volume is high enough that per-seat pricing becomes silly, the data can't leave systems you control, the tool must reach an in-house system nothing off the shelf connects to, and you have a developer relationship for the years after launch.

Per-seat pricing is where the sums most often flip, so run them carefully. Say a small charity has 40 volunteer advisers who all need answers from the same guidance library. On a standard business chat plan at $25 a seat a month, that's $1,000 a month, or $24,000 over two years. A narrow build (one question-and-answer page over the guidance, running on a cheap model such as gpt-6-luna at $0.10 per million input tokens and $0.50 per million output) might be quoted at $8,000 plus $100 a month, about $10,400 over two years. On those figures, building wins.

Then check the discounts before commissioning code. OpenAI for Nonprofits prices ChatGPT Business seats at $8 a user a month on annual billing, which makes 40 seats $320 a month, or $7,680 over two years, below the build and with nothing to maintain. Claude for Nonprofits offers Team at $8 a user a month too. Eligibility rules vary, so confirm the charity qualifies, but the lesson holds: the case for building often rests on a price nobody has negotiated yet.

If the tool you're weighing is a customer-service chatbot in particular, build or buy an AI chatbot for customer service goes through that narrower decision, including hand-offs to staff and per-resolution pricing.

Further reads

Sources: OpenAI help articles (custom GPT retirement and migration), ChatGPT Business and nonprofit pricing, OpenAI API pricing; Anthropic API pricing and model documentation; Make and Zapier help pages on error handling and Zap pausing; Microsoft Power Automate documentation.

Weighing a custom AI build against a product?

On a 1:1 call we'll put your buy, configure and build options side by side with your own volumes, and check whether tools you already have can do the job before anyone writes code.

Book a 1:1 call with me