Should a Small Business Hire a Prompt Engineer?

Coding Liquids tutorial cover featuring Sagnik Bhattacharya for Should a Small Business Hire a Prompt Engineer?
Coding Liquids tutorial cover featuring Sagnik Bhattacharya for Should a Small Business Hire a Prompt Engineer?

Usually not. Few small businesses have enough prompt-writing work to justify a hire: current AI models follow plain instructions well, and the harder work is choosing which tasks to automate, connecting your systems and checking the output. Hire an automation freelancer for defined projects instead, and make one existing employee the owner of your prompts.

The answer flips when AI talks to customers at volume, or when its output is what you sell. At a thousand or more AI conversations a week, one changed instruction can move accuracy or cost noticeably, and you need someone who tests before changing anything. Anthropic's own prompting guide assumes you already have clear success criteria, a way to test against them and a first draft. That testing discipline, not clever wording, is the skill worth paying for.

Follow me on Instagram@sagnikteaches

Why prompt writing rarely fills a job in a small firm

A prompt is the instruction you give an AI model. Writing good ones used to involve a lot of trial and error with oddly specific phrasing, which is where the job title came from. Three things have eroded that as a standalone role.

Connect on LinkedInSagnik Bhattacharya

First, the instructions themselves got easier. Current models respond well to the same things a capable new employee needs: context, a clear task, the format you want and an example or two. Both big model makers publish their methods free. Anthropic's prompt engineering overview links to an interactive tutorial on GitHub and a lighter version as a Google Sheet, and OpenAI keeps a prompt engineering guide in its API documentation. OpenAI's Playground will even draft a prompt from a plain description of the task.

Subscribe on YouTube@codingliquids

Second, several old tricks have stopped working. For years the standard advice for consistent answers was to set the temperature, a randomness setting, to zero. Anthropic has deprecated that parameter on its Claude 4.7 and later models, and OpenAI's reasoning models don't support it. Consistency now comes from structured output formats, fixed instructions with examples, and a review step.

Third, prompts are short-lived assets tied to platforms that change. OpenAI is retiring custom GPTs, which stop running on 11 December 2026, and is migrating them into plugins, where a GPT's instructions become a skill and its knowledge files become reference files. Anyone whose value was building GPTs has had to relearn the tooling within a couple of years. What survives each platform change is a written description of the process, the reference material and the tests.

What the work has turned into: context, tests and plumbing

The work that remains splits into four parts, and only one of them is wording.

  • Context: deciding what information the AI sees on each run, such as your price list, stock data, policies and tone guide, and keeping it current. Shared projects in ChatGPT and Claude, or Gems (which become skills from November 2026) in Gemini, hold that context for a team.
  • Plumbing: connecting the AI to the systems where the work lives, whether that's a stock system, an inbox, a CRM or an automation tool such as Zapier or Make.
  • Tests: a set of real inputs with known good answers that you rerun whenever anything changes, whether the prompt, the model or the source data.
  • Guardrails: what the AI must never do, who reviews what before it goes out, and how to handle hostile inputs. What prompt injection means for a small business covers the last of these.

Anthropic's documentation also makes a point that deflates the prompt-engineer idea: some goals, such as speed and cost, can sometimes be met more easily by choosing a different model than by rewriting the prompt.

Here is how the wording part plays out in practice, using used-car adverts at a dealership. The first attempt:

Write an advert for this car: 2021 hatchback, 1.0 petrol, 42,000 miles,
one owner, full service history, grey.

An illustrative reply opened: "This stunning 2021 hatchback pairs efficiency with style. Heated seats, a panoramic roof and smartphone mirroring make it perfect for families..." None of those three features was in the spec. For a dealership that's a customer complaint in waiting, and possibly a misdescription problem. The fix isn't a magic phrase. It's a rule, a format and a source of truth:

You write used-car adverts for our dealership website.
Use ONLY the features listed in the spec below. If a feature is not
listed, do not mention it. Never state fuel economy, insurance group
or warranty terms unless they appear in the spec.
Format: a headline under 70 characters, then three short paragraphs:
condition and history; practical features; how to book a viewing.
Tone: plain and factual. No exclamation marks.
Spec:
{paste the vehicle record from the stock system}

The illustrative second version was accurate but called the car "economical", a claim the spec didn't support. That came to light because the dealership had built a test set of twenty past vehicles, each with the features an advert may mention. "Economical" went on a banned-words list and the test set gained a case for it. That loop of spotting a failure, adding a rule and adding a test is the whole craft, and a sales administrator can learn it.

Five thresholds that would justify a specialist

These are rules of thumb rather than industry standards. If your business crosses two or more, specialist prompt and testing skills start to pay for themselves, whether you hire or contract them.

  1. Volume. One workflow runs more than about 1,000 times a week. At that scale a small drop in accuracy means dozens of bad outputs a day.
  2. No human in the loop. The AI's output reaches customers directly, as in a website chat assistant answering finance or warranty questions without review.
  3. Costly errors. A single wrong output can cost more than about $500: a mispriced quote, a wrong contract term, a compliance statement.
  4. Interlocking workflows. Five or more live AI workflows feed each other, so a change in one breaks another.
  5. AI output is the product. You sell reports, translations, listings or analyses that the AI produces, so quality is revenue.

Most small firms cross none or one. A business that crosses one should buy the skill for the specific project, not hire for it.

A car dealership's year of AI work, counted in hours

Take an illustrative independent dealership with 38 staff across sales, service and parts. It wants AI for five jobs: used-car adverts for about 60 vehicles a month, replies to service-booking emails, replies to online reviews, follow-up emails after test drives, and a weekly sales summary for the owner. Here is the work involved in year one, estimated task by task:

TaskHours
Build five workflows: prompts, context documents, links to the stock system and inbox (5 x 8 hours)40
Test sets: 20 to 30 real examples per workflow with expected answers (5 x 4 hours)20
Weekly review: sample outputs, update price lists and policies, fix drift (2 hours x 48 weeks)96
Platform changes: moving off a retired feature, retesting after a model update12
Total, year one168

A full-time role is roughly 1,700 working hours a year (37.5 hours a week over about 46 working weeks). The dealership's prompt work is about a tenth of that in year one, and less in year two once the build is done. A dedicated prompt engineer would be idle most of the time or drift into other work they weren't hired for.

What the dealership did instead: it contracted a freelancer for the 60 hours of build and test-set work, with a written handover, and made the sales administrator, who already wrote most adverts, the owner of the weekly review. At an illustrative internal cost of $28 an hour, her 96 review hours come to about $2,700 a year of existing staff time, much of it replacing the advert writing she was already doing. The dealership crossed none of the five thresholds, so a hire would have bought capacity it couldn't use.

When the answer flips to yes

A wholesaler answering trade customers at scale

Picture a wholesaler with 9,000 product lines that lets trade customers ask a chat assistant about stock, pack sizes and delivery cut-offs, around 3,000 conversations a week with no human reviewing each answer. It crosses the volume, no-human-in-the-loop and interlocking-workflow thresholds, since the chat assistant draws on the same product data that writes its web descriptions. Here a part-time or contracted specialist who owns the test sets, monitors wrong answers and manages changes is money well spent. The job description should stress testing and data quality, not phrasing.

An estate agency with an out-of-hours enquiry assistant

An estate agency's assistant might handle 400 conversations a month, well under the volume threshold. But it answers questions about fees and tenancy terms, where one wrong answer can cost real money, so it crosses the costly-errors line. That points to a short contract to build a thorough test set and a monthly review routine, not a hire. The agency's lettings manager can run the monthly review once the tests exist.

What to hire instead, and how to brief them

OptionFits whenYou pay forWatch for
Train an internal ownerOne to five workflows, no thresholds crossedA few staff hours a weekKnowledge leaving with one person; write everything down
Freelancer for a defined buildA clear project of roughly 20 to 80 hoursA fixed quote or agreed hoursPrompts living in the freelancer's accounts
AI automation specialist hireMany workflows across several systemsA salaryHiring for phrasing when you need integration skills
Consultant to set prioritiesYou don't yet know which jobs to automateA few days of adviceAdvice with no handover or ownership plan

For the freelancer route, how to hire an AI automation freelancer covers finding and vetting one. If the work spans several systems and will keep growing, hiring your first AI automation specialist is the better model, and AI consultant versus in-house hire compares the ongoing costs. I don't publish rates for any of these here; get two or three quotes against the same written brief and compare what each includes.

Whoever you use, the brief should name the deliverables so the work stays yours. A short example for the dealership:

Deliverables for the AI advert and email workflows
1. Final prompts and context documents, stored in our shared drive
2. A test set per workflow: 20+ real inputs, expected outputs,
   and a pass/fail note for each
3. A run log showing test results before and after each change
4. All automations built in accounts we own, not the contractor's
5. A 60-minute handover to our sales administrator, recorded
6. A one-page guide: what to check weekly and what to do on failure

Point 4 is the one most often missed. If the automations live in a contractor's Zapier or Make account, you don't own your workflows. Keeping the finished prompts in a shared prompt library makes the handover stick.

The cost of skipping it shows up late. Consider an illustrative two-site storage business that paid a contractor to build a customer-reply assistant as a custom GPT inside the contractor's own ChatGPT account, shared with staff by link. Staff could use it but nobody at the business could open its instructions. When OpenAI announced that custom GPTs stop running on 11 December 2026, the business had nothing to migrate: it had to rebuild the assistant in a ChatGPT Project it owned, reverse-engineering the rules from months of old replies. A single line in the brief would have prevented it.

Before paying the final invoice, check the work yourself in three ways. Rerun the test set on your own account and confirm the pass rate matches the contractor's report. Feed in five fresh inputs the contractor has never seen, including one awkward case. Then open every prompt, context file and automation while logged in as your business, not as a guest.

Growing an internal prompt owner in four weeks

For most small firms the prompt owner already works there. The job needs judgement about the business far more than technical skill: knowing what a good service reply or advert looks like, and noticing when an answer is subtly wrong. A realistic plan for one workflow, at a few hours a week:

  1. Week 1, about 3 hours: collect 20 to 30 real examples of the task alongside the answer you'd have been happy to send, and write down the rules experienced staff apply without thinking, such as "never promise a collection date before the controller confirms".
  2. Week 2, about 3 hours: write the instruction with context, format and two worked examples, and store it in a shared project so everyone uses the same version rather than private copies.
  3. Week 3, about 2 hours: run the test set, log pass or fail for each example, and fix the three most common failures, either with a rule or by correcting the source data.
  4. Week 4 onwards, about 30 minutes a week: sample ten real outputs, rerun the tests after any change to the prompt, the model or the price list, and date every change in a simple log.

Name a backup person who sits in on week 3, so the knowledge doesn't leave with one employee. After two or three workflows, the owner will be faster than most outside prompt specialists at the only thing that matters for your business: knowing when an answer is wrong.

If you do recruit: a skills test that beats a CV

If you cross the thresholds and decide to hire, skip interview questions about prompting tips. Give shortlisted candidates a paid work sample instead: twenty anonymised real inputs, your current prompt and two hours. Ask them to find where the prompt fails, write five new test cases, improve the prompt and report the pass rate before and after.

A filled-in scorecard from one illustrative candidate for the wholesaler's role:

CriterionWhat good looks likeScore (1-5)Note
Found real failuresNamed specific inputs that broke the prompt5Spotted wrong pack sizes on 3 of 20 inputs
Test casesNew cases cover the failure types found4Good, missed discontinued lines
Measured changePass rate before and after, honestly reported514/20 to 19/20, flagged the remaining miss
Knew when not to promptSuggested data fixes or a different model where wording couldn't help4Said stock data, not prompt, caused 2 errors
Explained it plainlyA non-technical manager could follow the write-up3Too much jargon in the summary

Warning signs in candidates or contractors alike: talk of secret prompt formulas or bought prompt packs, no mention of testing, and no instinct to check whether bad output comes from bad source data. The best answer in a hiring conversation for this role is often "the prompt isn't the problem here".

Prompt engineering questions from owners

Is prompt engineering still worth learning for my staff?

Yes, as part of existing jobs rather than as a job of its own. Staff who can give an AI clear context, a format and examples get usable drafts faster. The more valuable habit is testing: keeping a handful of real examples with known good answers and rerunning them whenever an instruction changes. That habit protects you far more than clever phrasing does.

Can an AI tool write our prompts for us?

It can write a decent first draft. OpenAI's Playground has a generate option that drafts a prompt from a plain description of the task, and Anthropic's documentation points to a metaprompt notebook that does the same for Claude. Neither knows your business, so the draft still needs your rules, your real examples and a test run before anyone relies on it.

What job title should we use instead of prompt engineer?

Describe the work, not the buzzword. If the job is connecting systems and building automated workflows, AI automation specialist or operations analyst fits. If it is mostly marketing content, hire a marketer who uses AI well. Titles built around prompt engineering attract candidates who sell phrasing tricks, while small firms usually need integration and testing skills.

How much prompt work should a small business budget for?

Count it in hours before thinking about money. A handful of workflows usually needs a few days of build and testing, then an hour or two a week of review. Multiply those hours by quotes from two or three freelancers, or by an existing employee's hourly cost, and compare the totals with the value of the time the workflows save.

Further reads

Sources: Anthropic Claude documentation (prompt engineering overview and prompting best practices); OpenAI API documentation (prompt engineering and prompt generation guides); OpenAI help centre on custom GPT retirement; vendor notes on the temperature setting. Checked September 2026.

Not sure whether your AI work needs a hire?

On a 1:1 call we'll count the AI work you actually have, decide whether it needs a hire, a freelancer or an internal owner, and sketch a handover so the prompts and tests stay yours.

Book a 1:1 call with me