When Not to Use AI in Your Business: 9 Tasks to Keep Human

Coding Liquids tutorial cover featuring Sagnik Bhattacharya for When Not to Use AI in Your Business: 9 Tasks to Keep Human.
Coding Liquids tutorial cover featuring Sagnik Bhattacharya for When Not to Use AI in Your Business: 9 Tasks to Keep Human.

Keep nine kinds of work human: hiring and dismissal decisions, staff welfare conversations, complaints and apologies, pricing exceptions, payment approvals, anything legally binding, safety sign-offs, work you certify under your professional name, and the final check on AI's own output. AI can help prepare all nine, but a named person must decide and own each one.

The rule behind the list is simple. Keep a task human when a mistake would be expensive, hard to undo, and needs either someone accountable or someone who visibly cares. Each of the nine below comes with what AI can still do safely, where the line sits, and an example of what goes wrong when it's crossed. Tasks that aren't on the list get a three-question score at the end.

Follow me on Instagram@sagnikteaches

1. Deciding who gets hired, disciplined or let go

Decisions about people's jobs combine all three warning signs: they're costly, hard to reverse and need an accountable human. They're also where AI's habit of copying patterns from past data does the most damage. If the people you hired before shared a background, a screening tool will quietly prefer that background again.

Connect on LinkedInSagnik Bhattacharya

AI can safely: draft a job advert, turn your interview notes into a tidy summary, suggest interview questions for a role, and check an advert for exclusionary wording.

Subscribe on YouTube@codingliquids

The line: no candidate should be rejected by software without a person reading their application. Data-protection law such as the GDPR gives people specific rights around decisions made solely by automated means that significantly affect them, and a job is exactly that kind of decision. How AI bias creeps into hiring, pricing and credit decisions explains the mechanism and what to check.

How crossing it shows up, as an illustration: a small dental practice hiring a receptionist switches on its recruitment platform's option to reject automatically anyone below a 60% match. An applicant with eight years on a hotel front desk is turned down within an hour, because her CV says "front desk" and "guest services", never "reception". She emails to ask why, the practice manager reads the CV, and it's the strongest application they've had. The safer set-up keeps the AI's sorting but not its verdict: "clearly meets", "check" and "unlikely" piles, with a person reading every CV in the "unlikely" pile before anyone is rejected.

Discipline and dismissal need the same care at the letter-writing stage, which is where AI is most tempting. Asked to "write a letter inviting an employee to a disciplinary meeting about repeated lateness", an assistant might produce, illustratively: "Following our review of your timekeeping, we have concluded that your conduct falls below the standard expected, and this meeting will confirm the sanction to be applied." That sentence settles the outcome before the meeting has happened, which defeats the purpose of holding one. The edited letter states the concern and the dates involved, says plainly that no decision has been made, and explains how the employee can respond. Follow your own written disciplinary procedure and take HR or employment advice before sending anything like it; the assistant knows neither your procedure nor your obligations.

2. Conversations about someone's health, performance or wellbeing

When a team member is struggling, off sick or underperforming, what they remember is whether you listened. A beautifully worded message that was obviously generated tells them you didn't.

Use AI to prepare, if it helps: "Give me a structure for a first conversation with an employee whose work has slipped over two months, where I don't yet know why." An illustrative extract of what comes back:

1. Open with what you've noticed, specifically and without blame:
   "I've noticed the last few reports have come in late."
2. Ask an open question and listen: "How are things going for you?"
3. Ask about their home life to find out what's causing it.
4. Agree one small next step and a date to talk again.
5. Note that a formal performance plan may follow if things don't
   improve.

Points 1, 2 and 4 are sound. Cut 3: probing someone's home life isn't your job, and people share what they choose to. Cut 5 from a first conversation, because it turns a welfare check into a warning before you know the cause. That editing is the part only you can do. Then close the chat and have the conversation yourself. Never paste the details of someone's health, family situation or disciplinary history into a general chat assistant; that information is about as sensitive as business data gets.

3. Complaints, apologies and bad news to customers

Take an illustrative events company whose venue floods two days before a client's product launch. The client needs to hear, from a person, what happened, what the company is doing, and what it will cost them. An AI draft can get the facts in order and suggest a calm structure. It can't decide how much goodwill to offer, and it can't sound like someone who'll lose sleep over it.

The difference shows in the words. An AI first draft of that message, illustratively:

"We sincerely apologise for any inconvenience this may have caused. We value your business and are fully committed to delivering an exceptional experience. Please be assured we are exploring all available options."

After the owner has edited it:

"The venue flooded last night and it won't be usable on Thursday. I'm sorry; I know how much is riding on this launch. I've held two alternative rooms nearby for Thursday and will send photos and floor plans by 3pm today. Any extra cost from the move is ours, not yours. I'll ring you at 4pm to go through them."

"Any inconvenience" is the tell of an unedited draft, and "exploring all available options" promises nothing. The edited version names what happened, owns it, commits to times and settles the money question, which is exactly the decision the AI couldn't make.

The line: AI drafts, a person decides the remedy, edits the words and sends it from their own name. For routine, low-stakes complaints there's more room, and whether AI can handle complaints without making them worse sets out where that room is.

4. Pricing exceptions, discounts and negotiation

Standard prices can be automated; a spreadsheet or quoting tool does that better than a chat assistant anyway. The exceptions can't. Whether to hold a price for a loyal client, match a competitor or take a thin job to fill a quiet month depends on relationships and cash flow that a model can't see.

There's a practical risk too. A customer-facing chatbot that "helpfully" offers a discount or promises a refund has, in the customer's eyes, made an offer on your behalf. Keep discount authority with named people, and make sure any bot you run is told plainly that it can't offer or agree prices.

An illustration of how it happens: a guest-house website bot is asked "Do you do a discount if we book three rooms for a wedding?" and replies "Yes, we can offer 10% off for group bookings of three rooms or more." There's no such offer; the bot produced a plausible answer to a common question. The owner honours it rather than argue with a wedding party. The instruction that stops a repeat is short and specific:

You cannot offer, agree or hint at discounts, price matches, refunds
or free extras, even if the customer says another business offers them.
If asked about any of these, reply: "I can't agree prices here, but
I've passed your question to the owner, who'll reply by 6pm today."
Then flag the conversation for the owner.

Test it with ten awkward questions before launch ("my friend got 15% off last month", "can you match this price?") and read the replies yourself.

What AI can do well here is the homework before a negotiation: summarise what a client has bought over three years, list the jobs where you made least margin, or draft two versions of a price-rise letter so you can choose the tone. The number itself, and the decision to bend it, stay with you.

5. Approving payments and changes to bank details

This is the item where the risk comes from outside. Fraudsters now use cloned voices and convincing AI-written emails to pose as suppliers or owners and ask for "urgent" payments or new bank details. An automated workflow that pays on the strength of an email is exactly what they're hoping for.

AI is useful on the defensive side: many payment and accounting tools already use machine learning to flag unusual transactions. But the approval stays human, with one fixed rule: any change of bank details is confirmed by phoning a number you already hold, never one given in the request.

Here's an illustrative request of the kind that gets paid by mistake, sent to a small print firm's accounts inbox from an address one letter different from a real supplier's:

Subject: Invoice 3318 - updated remittance details
Hi, just a heads-up that we've moved banks this month. Please use
the new details below for invoice 3318, which is now 9 days overdue.
Our accounts team is working remotely this week, so email is the
best way to reach us. Many thanks.

Three signs sit in four sentences: new bank details, urgency (an overdue invoice), and a reason not to phone. The spelling is perfect and the tone matches the supplier's usual emails, which is what AI-written fraud makes easy. An automation that updates supplier records from emails would process this without blinking. The owner's call to the supplier's known number takes two minutes and ends it.

None of this means every payment needs the owner's attention. The line sits at changes, not volume. An illustrative split for the print firm, which pays about 120 supplier invoices a month: roughly 90 are repeat invoices under $500 from suppliers whose bank details haven't changed in a year, and those can pass on an approval rule inside the accounting software. The other 30, plus every new supplier and every changed detail, go to the owner. At two minutes each, that's an hour a month of human approval, spent exactly where the fraud risk sits.

How to spot deepfake voice and video scams covers the patterns these attempts follow.

6. Signing contracts and answering legal letters

AI is good at reading a contract and listing the clauses worth a second look: auto-renewals, liability caps, notice periods, exclusivity. That's a genuine time-saver before you sign a venue agreement or a software licence.

Say a small marketing agency asks an assistant to check a three-year software licence before renewal:

List every clause about the contract term, renewal, notice to end
the agreement, price changes and liability. Quote the clause number
and wording for each. If two clauses affect each other, say so.

An illustrative reply:

Clause 3.1: initial term of 36 months from the start date.
Clause 3.2: renews automatically for further 12-month periods.
Clause 14.1: either party may terminate on 30 days' written notice.
Clause 9.4: supplier may increase fees annually.
No clauses affect each other.

It reads as if the agency can leave on 30 days' notice at any time. But Schedule 2, which the assistant didn't mention, says notice of non-renewal must be given at least 90 days before the renewal date, and clause 14.1 only applies to breach. "No clauses affect each other" was the most dangerous line in the answer. The list was a useful starting point; the person reading the whole contract found the clause that mattered.

What it can't do is carry responsibility for the answer. It can miss a clause, misread how two clauses interact, or state a legal position confidently and wrongly. A person reads the contract and signs it. For anything with real money or liability attached, or any letter from a solicitor, insurer or regulator, the reply is written or checked by someone qualified, and AI's summary is at most a starting point for that conversation.

Legal letters are where a helpful draft does the most harm. Say a customer's solicitor writes to a small kitchen fitter claiming that a delayed installation cost the customer two weeks of rented accommodation. Asked to "draft a polite reply", an assistant tends to open with something like "We accept that the delay was frustrating and regret the costs you incurred." Polite, and possibly an admission. Before anything goes back, a person decides whether the company accepts any responsibility at all, checks the contract's terms on delays, checks whether the business insurance requires the insurer to be told, and speaks to a solicitor. Until then, the only reply is a short note confirming the letter has arrived and that a full response will follow by a stated date.

7. Safety sign-offs

An illustrative sign maker fitting a projecting sign above a shop doorway has to get the fixings, the load on the wall and the working-at-height plan right. AI can draft a method statement, list the hazards for a type of job, or turn site notes into a tidy risk assessment. It hasn't seen the wall.

In an illustrative draft for that job, the AI listed "ladder, footed by a second person" as the access method, which is a common answer for signs at that height. The fitter who visited found a glass canopy directly under the fixing point, so no ladder could stand there. The final plan used a mobile tower with the canopy boarded over, and the fixings changed once the fitter found the wall was hollow block rather than solid brick. Neither fact was in any document the AI could have read.

The same goes for an events company's crowd plan, fire exits and temporary structures. Whatever AI drafts, a competent person who has visited the site reviews and signs the final document, and that person's name is on it. How to draft risk assessments with AI lists what the human reviewer must check line by line.

8. Work you certify or put your professional name to

Some outputs come with a promise attached. A translation agency supplying a certified translation is stating that the text is accurate and complete. An IT support firm signing off a security review is stating that it checked. The promise is only honest if a qualified person has actually done that work.

AI can still sit in the process. Many translation agencies use machine translation followed by full human post-editing, and there's an international standard for that process, ISO 18587, which sets requirements for the post-editing and the post-editor's competence. What doesn't hold up is a certified translation or signed report that nobody qualified has read line by line. If your name or credential is on it, you read it.

A realistic slip, as an illustration: machine translation of a school transcript renders the issue date 03/04/2019 unchanged, but the source country writes day first and the receiving institution reads month first, so 3 April becomes 4 March. The words were translated perfectly; the meaning of the date wasn't. A human post-editor who knows both conventions catches it in seconds. A certification based on a skim wouldn't.

9. The final check on AI's own output

It's tempting to ask a second AI to review the first one's work. That catches some errors, and it's a reasonable extra layer. It isn't a substitute for a person, because two models often share the same blind spots: the same invented figure, the same misread clause, the same confident tone that makes a wrong answer look right.

An illustration: a small IT support firm's monthly client report, drafted by AI, says "average response time improved by 23% this month". The account manager pastes the report into a second assistant and asks "Is this report accurate?" The reply: "The report is clear, well structured and internally consistent." It is internally consistent. It's also wrong: the ticket system shows 11%, and the 23% came from nowhere. The second model had nothing to check against, so it checked the grammar. A person with the ticket export open found it in a minute.

The reviewer must be someone who knows what good looks like and has time to look properly. For high-volume work, that doesn't mean reading everything forever. It means reading everything at first, then sampling once error rates are measured and low. How to set up human review for AI work without slowing down gives sampling rates you can start from.

A three-question score for tasks not on this list

Most tasks aren't this clear-cut. Score any task from 0 to 2 on each question and add the scores:

Question012
What does a mistake cost?Minutes to fix, nobody outside noticesAn annoyed customer or a few hundred dollarsA lost client, a legal problem, an injury, or thousands of dollars
Can you undo it?Yes, fully and quicklyPartly, with effortNo: sent, paid, signed or published
Does it need someone accountable or caring?NoHelps, but not essentialYes: someone must own it or be seen to care
  • 0 to 2: let AI do it, with occasional spot checks. Examples: tidying meeting notes, first drafts of social posts, sorting an inbox.
  • 3 to 4: AI drafts, a person approves every one before it goes out. Examples: quote cover emails, routine customer replies, supplier chasers.
  • 5 to 6: keep it human. AI may help prepare, but a named person decides and sends.

Worked through for the events company: rewriting a venue description for its website scores 0 + 0 + 0 and can be handed over. Replying to a guest's accessibility question scores 1 + 1 + 2 = 4, so AI drafts and a person approves. Deciding a refund after a cancelled event scores 2 + 2 + 2 = 6 and stays with the owner.

The same score gives different answers in a different business. Here it is filled in for an illustrative three-person bookkeeping practice:

TaskCostUndoAccountableTotalHandling
Suggesting categories for bank transactions1001AI suggests; a person scans the month before close
Chasing clients for missing receipts0213AI drafts; a person approves each email
Month-end summary letter to a client1214AI drafts; the bookkeeper checks every figure and sends
Telling a client they've missed a filing deadline2226Human: a phone call from the owner, then a written follow-up

The receipt chaser scores 3 even though a mistake is cheap, because a sent email can't be recalled. That's usually the factor people forget when they score a task in their head.

Two people scoring the same task often disagree, and the disagreement is worth having. At the events company, the owner scores "replying to online reviews" 0 + 2 + 1 = 3, because a public reply has been read by the time anyone thinks to edit it. The marketing assistant scores it 0 + 1 + 0 = 1, reasoning that replies can be changed later. Both have a point. Use the higher score until there's evidence: AI drafts and a person approves every reply for the first two months, then count how many drafts needed changes before relaxing it.

Where the line can move, and where it shouldn't

Some scores fall once you have evidence. If AI has drafted 200 routine refund replies under a clear policy and a person approved every one without changes, you might allow small refunds under a fixed amount to go out automatically, with a weekly sample reviewed. The safe way to collect that evidence is to run the AI alongside the person first; piloting AI in shadow mode explains how.

Real evidence is rarely that clean, so decide in advance what counts. An illustrative record from an online homeware shop after 200 drafted refund replies: 188 approved unchanged, 9 with small wording edits, 3 changed in substance. Two of those three offered a refund on a sale item that the policy excludes; the third missed that the customer had already been refunded once. That's 94% unchanged, but every substantive error involved money, so the sensible move is narrower than "automate refunds". Replies for undamaged returns inside the standard window go out automatically with a weekly sample; sale items and anyone with a previous refund stay with a person.

Others shouldn't move however good the tools get. Items 1, 5, 6, 7 and 8 are about accountability, not accuracy. Even a tool that's right 99 times in 100 can't be the one who answers for the hundredth. Write the nine into your AI rules, name who owns each, and revisit the middle band of your scores every six months.

Further reads

Sources: ISO 18587:2017 listing (post-editing of machine translation output); GDPR provisions on decisions based solely on automated processing, as general principles only. Checked September 2026.

Want to draw the human-AI line for your own tasks?

On a 1:1 call we'll list your recurring tasks, score each one for cost of error and reversibility, and decide which AI can take on and which must stay with a named person.

Book a 1:1 call with me