AI bias is when a tool's outputs consistently favour or penalise groups of people, usually because it learned from skewed past decisions or relies on stand-ins such as address, name or career gaps. In hiring it can screen out qualified candidates; in pricing and credit it can charge or refuse people unfairly, with legal and reputational consequences.
Small firms rarely build these tools, but they increasingly use them: CV screening inside recruitment software, AI-suggested quotes, automated checks that decide who gets payment terms. Whatever the software did, the decision remains yours, legally and commercially. Knowing how bias gets in makes it much easier to spot, and the test at the end can be run on any AI-assisted decision in an afternoon.
Four ways bias gets into an AI tool
1. It learns from past decisions
A tool trained or tuned on your history will reproduce the patterns in it, including unfair ones. The best-known example: Reuters reported in 2018 that Amazon had scrapped an experimental recruiting tool trained on ten years of CVs, most of them from men. The system learned to mark down CVs that included the word "women's", as in "women's chess club captain". Nobody programmed that. It came from the history.
2. It uses stand-ins, called proxies
A tool can be kept away from protected characteristics such as sex, age, ethnicity or disability and still reach them indirectly. Common proxies in small-business data:
- Address or neighbourhood, which can track ethnicity and income.
- Graduation year or years of experience, which track age.
- Career gaps, which often reflect caring responsibilities, illness or disability.
- Names, which signal gender and ethnic background.
- Writing style, which can reflect whether someone is writing in a second language, or has dyslexia.
3. You teach it with skewed examples
This one is common with chat assistants. If you paste in three CVs of "the kind of person we want", all from people with similar backgrounds, the AI generalises from them. The bias is in your prompt, not the model.
An illustration from a small accountancy practice hiring a trainee. The first prompt looked harmless:
Here are the CVs of our three best trainees from recent years.
Rank the 40 applicants below by how similar they are to them.
An illustrative extract of the reply:
1. Applicant 17: very similar. Same university as Trainee B,
joined straight from graduation, rowing club.
2. Applicant 03: similar. Straight from graduation, no gaps.
...
38. Applicant 29: low similarity. Six years in retail management
before retraining; 2023 part-time qualification.
The reasons give it away: university, "straight from graduation" and a club are standing in for age and background, and the career changer drops to the bottom for having a history. The rewrite scores against written criteria only:
Score each applicant 0, 1 or 2 against these five criteria only:
accounting qualification progress; spreadsheet skills shown in
work or study; written communication; client-facing experience of
any kind; evidence of meeting deadlines.
Ignore names, institutions, dates and hobbies. For each score,
quote the evidence from the CV. If there is none, write NO EVIDENCE.
Under the second prompt, the career changer scores 2 for client-facing experience and deadlines, from the retail management years, and lands in the top ten. The practice still reads every CV that meets the essentials; the prompt just stops the list being sorted by resemblance.
4. Generative models carry stereotypes
Assistants trained on huge amounts of public writing absorb its assumptions. Left unchecked, they produce job adverts with gender-coded wording, or assume a nurse is a woman and an engineer is a man. Spotting bias and stereotypes in AI-written content covers this for everyday writing.
A busy café hiring a shift supervisor shows how quickly this happens, as an illustration. The owner asks an assistant for an advert and gets back: "We're after a young, energetic team player, ideal for a recent graduate, who's a native English speaker and can keep up with our fast-moving crew." Three phrases there filter people out for reasons unrelated to the job. "Young" and "recent graduate" signal age. "Native English speaker" excludes fluent speakers who learned later. The corrected line keeps the real requirements and drops the proxies: "You'll run a team of four on busy weekend shifts, handle the till and stock count, and speak confidently with customers in English." What the job needs is fluent spoken English for serving customers; where someone learned it is irrelevant.
Hiring: where bias enters a small firm's recruitment
An illustration. Say a twelve-person marketing agency is hiring a junior account executive and receives 180 applications. It uses its recruitment software's AI ranking to shortlist 30. Bias can enter at four points:
- The advert. Phrases like "digital native" or "recent graduate" signal age; some wording puts off particular groups. Writing job adverts with ChatGPT without biased language shows how to check the draft.
- The ranking criteria. If the tool ranks against the agency's current team, it favours people like the current team.
- How it treats unusual routes. Career gaps, career changers and qualifications from institutions the tool doesn't recognise can be marked down without anyone deciding they should be.
- Interview notes and summaries. Automatic transcription tends to be less accurate for some accents, and an AI summary built on a poor transcript can make a strong candidate read as muddled.
The transcript problem is the easiest of the four to miss, because the damage surfaces two steps later. In an illustrative case, a candidate with a strong accent says "I briefed the copywriters and chased sign-off with the client", and the transcript records "I brief the copy rights and chase sign off with the client". The AI summary built on that reads "seemed unclear about copyright responsibilities". Nobody in the room would have written that. The check is cheap: for any candidate the summary rates below your shortlist line, open the transcript beside the recording and read the two or three answers the summary criticises. If the words don't match what was said, discard the summary and score from the interviewers' own notes.
Checking the outcome. The agency invites candidates to fill in a voluntary equal-opportunities form, and 150 do. Of those, 60 men applied and 14 were shortlisted, a rate of 23%; 90 women applied and 13 were shortlisted, a rate of about 14%. A widely used rule of thumb in recruitment auditing, the four-fifths rule, says to investigate when one group's selection rate is below 80% of the highest group's. Here it's about 62%.
That doesn't prove bias on its own; with numbers this small, chance plays a part. But it's a clear prompt to look. Say the investigation finds two causes: the tool was marking down CVs with gaps of more than a year, and it had been set up with three example CVs from the agency's current senior staff. The fix: remove the examples, score only against the written criteria, and have a person read every rejected CV that meets the essential requirements. Bias checks every recruiter should run on AI CV screening goes further into the tests.
Pricing: when personalised quotes become unfair
Pricing differently for different jobs is normal business. The problem starts when price varies with who the customer is, through characteristics or proxies that have nothing to do with the work.
An illustration. Say a web design studio adds an AI step that reads each enquiry and suggests a quote band, tuned on its past quotes. Six months later, a founder notices a pattern: enquiries written in less fluent English get suggested the higher "extra support" band more often, even for small, simple sites. The tool had learned that from a period when one project manager padded quotes for clients he expected to need more hand-holding. The price was tracking the writing, not the scope.
Controls that prevent this:
- Price on the job, not the person. Number of pages, features, integrations, content support, deadline. Write the factors down.
- Feed the tool a structured scope, not the raw enquiry. If the AI only sees "8 pages, booking integration, client supplies copy, six-week deadline", it can't price on name, address or writing style.
- Run paired tests. The same scope sent with different names, locations and writing styles should get the same band.
- Keep a person on the final quote, and log how often they change the AI's suggestion and why.
If you sell to consumers in the EU, there's also a disclosure rule: consumer law requires traders to tell consumers, in distance and off-premises sales, when a price has been personalised on the basis of automated decision-making.
Credit and payment terms: deciding who pays up front
Small firms make credit decisions all the time without calling them that: who gets 30-day terms, who pays a deposit, who pays in full before work starts. AI shows up here in credit-check tools and in accounting software that predicts late payers. Both can be useful; predicting late payers with AI covers the legitimate use.
An illustration. Say a video production company takes on a mix of business clients and private individuals commissioning films. It considers a tool that scores new clients and recommends either 30-day terms or a 50% deposit. The risks:
- Proxies again. Features such as location or how recently someone started trading can stand in for characteristics the decision shouldn't depend on.
- Fully automated refusals. Data-protection law such as the GDPR gives individuals the right not to be subject to a decision based solely on automated processing that has legal or similarly significant effects on them, subject to exceptions; where a business relies on the contract or explicit-consent exceptions, people are still entitled to a human review, to give their view and to contest the decision.
- No explanation. If a client asks why they need to pay a deposit and the only answer is "the score said so", that's a poor answer commercially and potentially a legal problem.
That last point is easy to fix with a prepared reply. An illustrative version the video company could send when an individual client asks about a deposit:
"Thanks for asking. For new clients whose project is over $8,000, and who haven't worked with us before, we ask for 50% up front. That's the rule for everyone in that position, and it's about the size of the job and the lack of payment history with us, not about you personally. After one project paid on time, we move clients to 30-day terms. If you think your situation is different, reply and one of the directors will look at it."
It names the criteria, says a person can review it, and gives a route to better terms. If you can't write a reply like that for a decision your tool makes, you don't yet know what the tool is deciding on.
The practical design: the tool recommends, a person decides for every individual client who gets the stricter option, the criteria are written down (for example, payment history with you and the size of the job), and clients are told they can ask for the decision to be reviewed. For the manual side, checking a new customer's credit before offering terms sets out a sensible process.
One edge case needs its own rule: the client the tool knows nothing about. A sole trader who started four months ago has no filed accounts and little credit history, and many scoring tools treat missing data as high risk. That says something about the tool's information, not about the client. For the video company, the fairer rule is that a "not enough data" score always goes to a director, who can offer staged payments instead of a flat deposit. On an illustrative $12,000 film, 30% at booking ($3,600), 40% at first edit ($4,800) and 30% on delivery ($3,600) asks the client for $2,400 less up front than a 50% deposit, and the finished film still isn't handed over until 70% has been paid.
What the law expects, in general terms
This isn't legal advice, and the details depend on where you and your customers are. The broad picture:
- Discrimination law applies to the outcome. It covers who gets shortlisted, what they're charged and what terms they get, however the decision was reached. "The software did it" is unlikely to be a defence.
- Data-protection law restricts solely automated decisions with significant effects on individuals, as above.
- The EU AI Act lists both areas as high-risk. AI used for recruitment and selection, and AI used to evaluate the creditworthiness of individuals, are in its Annex III. Obligations for these stand-alone high-risk systems were deferred and now apply from 2 December 2027, which gives time to get your checks in place now.
- Consumer law may require disclosure of personalised pricing, as above.
If you use AI in any of these three decisions and aren't sure where you stand, talk to an employment solicitor for hiring and your data-protection adviser for credit and pricing. The AI compliance checklist for small businesses helps you prepare for that conversation.
An afternoon bias test for any AI-assisted decision
- Write down what the decision should depend on. Five to eight criteria, in plain words.
- List everything the tool can see. Strip anything not on the criteria list: names, addresses, dates of birth, photos, graduation years, and raw free text where a structured summary would do.
- Run paired tests. Make ten pairs of inputs that are identical except for one proxy, and record the outcomes in a log like the one below. Outcomes should match.
- Check real outcomes by group where you have voluntary data. Use the four-fifths rule of thumb as a trigger to investigate, not as proof, and remember that small numbers are noisy.
- Put a person on every negative outcome for an individual: every rejection, every stricter payment term, every quote the AI pushed up.
- Record the test and repeat it whenever the tool, its settings or your inputs change, and at least once a year.
Pair | What differs | Result A | Result B | Same? | Note
1 | Name | | | |
2 | Address or neighbourhood | | | |
3 | 18-month career gap | | | |
4 | Graduation year (+15 years) | | | |
5 | Written in a second language | | | |
6 | Sole trader vs limited company | | | |
Here is the same log filled in for the web design studio's quote-band step from the pricing section. Each pair uses one fixed scope (an eight-page site with a booking integration) and changes one thing in the enquiry. The results are illustrative:
Pair | What differs | Result A | Result B | Same? | Note
1 | Name | Band B | Band B | Yes |
2 | Address or neighbourhood | Band B | Band B | Yes |
3 | Written in a second language | Band B | Band C | NO | "may need extra support"
4 | Sole trader vs limited company | Band B | Band B | Yes |
5 | Short enquiry vs long enquiry | Band B | Band B | Yes |
6 | Says "I'm not very techy" | Band B | Band C | NO | same reason as pair 3
Two mismatches, both with the same stated reason, point straight at the cause: the tool is guessing at support needs from how people write. The fix is the structured scope described earlier, so the AI never sees the raw enquiry. Run the six pairs again after the change; they should all read "Yes" before the step goes back into use.
Paired tests show the tool can behave; live numbers show whether it does. A quarter after the change, the studio can pull every enquiry for a site of ten pages or fewer and count how many were put in Band C. In an illustrative before-and-after, 11 of 40 small-site enquiries (28%) landed in Band C in the quarter before the fix, against 3 of 38 (8%) in the quarter after, and each of those three had a reason written in its scope, such as "client needs all copy written". A drop like that, with a scope-based reason for every exception, is what a working fix produces. If the rate barely moves, something other than writing style is still steering the band.
A test like this won't catch everything, but it catches the common failures, and having run it puts you in a far better position if a candidate or customer ever asks how the decision was made.
Questions about bias and AI decisions
Is it legal to use AI to screen CVs?
Generally yes, but the responsibility for the outcome stays with you. Discrimination law applies to who gets shortlisted whether a person or software did the sorting, and data-protection law such as the GDPR limits decisions made solely by automated means. If you recruit in the EU, recruitment AI is also listed as high-risk under the EU AI Act. Keep a person in the decision and ask an employment adviser if unsure.
Can I just tell the AI not to be biased?
An instruction helps with wording, such as neutral language in a job advert, but it doesn't remove bias that comes from the information the tool sees. If names, addresses or career gaps are in the input, they can still sway the output. Remove what the decision shouldn't depend on, then test with paired examples to see whether results change.
Do these concerns apply if my customers are businesses, not individuals?
Partly. Data-protection rules and the EU AI Act's credit-scoring category focus on individuals, so a limited company applying for payment terms is treated differently. But sole traders and partnerships are people, and small business owners can still be affected by proxies such as location. Treat any decision about a sole trader as a decision about a person.
Further reads
- AI Ethics for Small Businesses: A Practical Checklist — A wider checklist for fairness, transparency and accountability.
- How to Set Up an AI-Assisted Hiring Process for a Small Team — Build a hiring process with AI that keeps people in charge.
- Can AI Tell You What to Charge? The Limits of AI Pricing Research — The limits of AI when setting prices.
- GDPR and AI Tools: What a Small Business Must Do — Data-protection duties when AI touches personal data.
- Questions to Ask a Recruitment Agency About Its AI Screening — What to ask if an agency screens candidates for you.
- What Are the Risks of Using AI in My Small Business? — Bias in context with the other risks of using AI.
- Do I Need an AI Consultant? 8 Signs It's Time to Get Help — Eight signs a small business needs outside AI help, each with a test you can run this week, plus the signs that it doesn't need a consultant yet.
- How to Test Job Candidates' AI Skills in an Interview — An AI skills test for small-firm interviews: one realistic task with a planted error, a scoring rubric, and follow-up questions that reveal judgement.
- The Limits of AI: What It Still Gets Wrong in a Small Business — Eight things AI still gets wrong in a small firm, where each one bites, the warning signs, and the specific check that catches it.
- When Not to Use AI in Your Business: 9 Tasks to Keep Human — Nine tasks a person should always decide, what AI can safely prepare for each, and a three-question score for any task that isn't on the list.
- What an AI Implementation Plan Looks Like for a Small Restaurant — One illustrative 48-cover restaurant's AI plan, start to finish: a week of tracking, three jobs chosen, four ruled out, and the day-90 numbers.
- Is AI Dynamic Pricing Worth It for a Small Hotel? — When AI rate-setting pays for a small hotel and when it doesn't: five deciding factors, three example properties, break-even maths and the guardrails to set.
- How to Build a Restaurant Staff Rota With AI Demand Forecasts — Turn a covers forecast into an hourly staffing plan and a fair rota, with a worked Saturday that shows the labour-cost difference in dollars and hours.
- Can AI Help a Small Shop Set Prices and Promotions? — How a small shop can use AI to test promotions against its own margins before running them, with the break-even maths and a pharmacy example worked through.
- How to Evaluate the AI Features in Your Recruitment Software — Six tests to run on the AI already in your recruitment software, using past placements as the answer key, with pass marks and a filled-in scorecard.
- AI CV Screening for Recruitment Agencies: Setup and Safeguards — Five set-up steps and the safeguards to have in place before an agency lets AI score CVs, with a rubric, prompt, sample output and back-test.
- AI Tools and AI Development: The Complete 2026 Guide — the AI hub, including every tutorial in the AI-for-business series.
Sources: GDPR Article 22 text; EU AI Act Annex III and Digital Omnibus timetable; Consumer Rights Directive as amended by Directive (EU) 2019/2161; Reuters reporting (October 2018) on a scrapped recruiting tool.