Ask the vendor to name what its AI is (a language model, a trained prediction model or fixed rules), then make it perform your task live on your own data. Software that oversells its AI blurs that line, quotes accuracy without saying how it was measured, and can't explain what happens when the AI gets something wrong.
AI washing, labelling ordinary software as AI or exaggerating what its AI does, is more than a marketing irritation. Regulators have already fined firms for claiming to use AI they weren't using, and acted against a subscription service sold as an "AI lawyer" whose output had never been tested against a real lawyer's work. For a small buyer the risks are simpler: paying an AI premium for a rules engine, or trusting an unchecked "smart" decision with your customers or stock. A rules engine can be excellent. It should be sold, and priced, as one.
What AI washing sounds like on a sales page
Four kinds of technology get called AI, and all four are legitimate when described honestly. Rules are if-then instructions written by people, such as "send the nearest free engineer". Prediction models are trained on past data to estimate something, such as next month's demand. Generative models, the large language models behind ChatGPT and Claude, produce text or images. Agents are models that take actions through other software, such as booking an appointment. Washing happens when the label on the page doesn't match the thing underneath.
| Phrase on the page | What it may really be | Question that exposes it |
|---|---|---|
| "AI-powered smart scheduling" | A nearest-available rule | Could you write the rule down for me in one sentence? |
| "Predictive demand forecasting" | A moving average of recent sales | How does it compare with last year's sales for the same weeks? |
| "Learns from your data" | Remembers your settings | What exactly changes after it learns, and how often? |
| "Proprietary AI" | A rented language model with a prompt | Which model providers are on your sub-processor list? |
| "AI agent" | A scripted chatbot with fixed buttons | What can it do without a human clicking anything? |
| "95% accurate" | Measured on the vendor's tidy demo data | Measured on what data, how many cases, and what counted as correct? |
| "Fully autonomous" | A team of people checking outputs behind the scenes | Does anyone on your staff review our results before we see them? |
None of the right-hand answers is disqualifying on its own. A rented model is how most AI products work, as how to tell if a tool is a ChatGPT wrapper explains, and human review is often a good thing. The problem is being told one thing and sold another.
Checklist, part one: the claims in writing
Do these before any demo. Most take minutes and need only the vendor's own website and help pages.
- The AI is named and described. Why: "AI-powered" with no method attached is the commonest wash. Verify: look for help pages or documentation explaining how the feature works; if there are none, ask for a one-paragraph description of the method in writing.
- Model providers are disclosed. Why: "proprietary AI" often means a rented model. Verify: read the sub-processor list linked from the privacy policy and search it for model providers.
- AI features are separated from ordinary ones. Why: vendors often stretch one AI feature into an "AI platform". Verify: ask for a list of which features use AI and which don't. A vendor that can't produce one hasn't thought about it or doesn't want you to.
- Release status is clear. Why: beta features change or vanish. Verify: check release notes and the plan comparison page for labels such as beta, preview or "available to selected users".
Checklist, part two: the live demonstration
A scripted demo on the vendor's sample data proves almost nothing, because sample data is chosen to work. How to run an AI software demo covers the logistics; these are the AI-specific tests.
- Your data, not theirs. Why: tidy demo data flatters any method. Verify: send 20 to 30 anonymised real records a few days ahead and watch them processed live.
- An unscripted awkward case. Why: real AI copes, imperfectly, with inputs it hasn't seen; rules break or ignore them. Verify: bring one messy example on the day, such as a handwritten note, a misspelt product or an enquiry that mixes two requests.
- A visible failure. Why: honest products show uncertainty and hand off. Verify: ask "show me what happens when it's wrong or unsure". Look for confidence flags, review queues or escalation to a person.
- The same job with AI switched off. Why: if the output barely changes, the AI isn't doing much. Verify: ask to see the feature's result with the AI option disabled. Not every product allows this; QuickBooks Online, for instance, doesn't currently let you switch its AI features off one by one.
Checklist, part three: evidence and accuracy numbers
- Accuracy comes with its test. Why: a percentage without a method is marketing. Verify: ask what data it was measured on, how many cases, what counted as correct, and when.
- A simple baseline is beaten. Why: many "AI" forecasts do no better than last year's figures or a moving average. Verify: compare the tool's output on your data with the simplest method your staff could run in a spreadsheet.
- References match your size and trade. Why: a result at a 500-person company says little about a 12-person one. Verify: speak to two customers of similar size; how to check a vendor's customer references has the questions.
- Some evidence isn't written by the vendor. Why: case studies are sales material. Verify: look for reviews, user forums or a trial you run yourself.
Checklist, part four: price, contract and what happens when it's wrong
- The AI premium is itemised. Why: you can only judge the AI if you know what it costs. Verify: get prices for the plan with and without the AI features or add-on.
- Usage limits are stated. Why: "AI included" often means a monthly allowance of credits. Verify: ask what happens when you hit the allowance and what extra usage costs.
- The contract describes the feature and its withdrawal. Why: marketing claims you can't point to in the contract are hard to hold anyone to, and AI features do get retired. HubSpot, for example, stopped new installs of five Breeze agents on 23 July 2026, though existing installs kept working. Verify: check the order form or service description names the AI feature, and ask what notice you get if it's withdrawn.
- Someone owns the mistakes. Why: an AI that acts without review can cost you money. Verify: ask whether you can review outputs before they reach customers or stock systems, and who is responsible when the AI is wrong.
Running the checklist on a car dealership's pricing tool
Here is an illustrative independent dealership selling about 60 used cars a month. A stock-management vendor offers an "AI Pricing" tier at $300 a month more than its listings-only tier, claiming "AI sets the optimal retail price for every car".
Part one went badly. The help pages described the feature as "machine-learning pricing" with no further detail. The sub-processor list named no model providers, which is plausible for a pricing tool that doesn't generate text. The vendor couldn't say which parts of the tier were AI. On the demo call, the product specialist explained the method: the tool finds comparable cars listed for sale nearby, by make, model, age and mileage, and adjusts for mileage and options. That's a comparables search with fixed adjustments, useful but closer to rules than to anything that learns.
So the dealership ran the test that matters, on its own data. It gave the vendor 30 recently sold cars and compared three prices for each with the actual selling price: the tool's suggestion, the price its own manager had set using the dealership's rule of thumb, and the final sale price.
| Measure (30 cars, illustrative) | Tool's price | Manager's rule |
|---|---|---|
| Within 3% of the final sale price | 19 cars | 15 cars |
| Priced more than 5% too high (sat on the forecourt) | 3 cars | 7 cars |
| Time for the manager to price one car | 2 minutes to check | 10 minutes |
The claim was washed; the product still performed. Four fewer overpriced cars in 30 means roughly eight fewer a month at 60 sales. If each overpriced car costs the dealership around $200 in extra days on the forecourt and a later price cut (the dealership's own estimate), that's about $1,600 a month of value against a $300 premium, plus about eight hours a month of the manager's time. The dealership bought it, on two conditions: the order form describes the feature as comparables-based pricing rather than "AI", and the vendor commits to notice if the method changes. That's the balanced lesson of AI washing. Judge the result on your data, and make the contract describe what the product actually does.
Three shorter cases, and an assistant to help
A wholesaler's "predictive" forecasting add-on
A wholesaler with 2,400 product lines was offered an AI forecasting add-on at an illustrative $400 a month. It compared the add-on's forecasts for 50 lines over 12 weeks with a 12-week moving average built in a spreadsheet in an afternoon. The add-on's average error was 31%; the moving average's was 33%. Two points of accuracy weren't worth $4,800 a year, especially since the add-on couldn't explain its seasonal adjustments. The wholesaler kept the spreadsheet and set a reminder to repeat the test in a year.
A property maintenance firm's "smart dispatch"
A field-service platform sold "AI dispatch" to a property maintenance firm. Asked to write the rule down, the vendor's answer was: assign the nearest engineer with the right trade who has a free slot. The firm's scheduler could configure exactly that rule in its existing job-management software. Nothing wrong with the rule, but it wasn't worth an upgrade.
An "AI" invoice reader that needed templates
Here is a mistake and how it surfaced. A courier firm bought a document-capture tool described as AI-powered. It worked well on the firm's six regular fuel and vehicle suppliers. Then a new tyre supplier's invoices started arriving blank in the accounts system. The tool relied on a template set up per supplier; the AI label covered a character-recognition step, not reading unfamiliar layouts. Demo test 6, one unfamiliar document on the day, would have caught it.
Letting an assistant read the sales page first
A general assistant can do the tedious first pass. Paste in the vendor's feature page and pricing page, and ask:
Below is a software vendor's feature page and pricing page.
1. List every claim about AI, machine learning, prediction or automation.
2. For each claim, quote any evidence the page gives (method, data,
accuracy figures, customer results). If none, write "no evidence".
3. For each claim, suggest one question to ask the vendor that would
show whether the claim is true.
Do not judge the product; only extract and question.
An illustrative extract of the reply for the dealership tool:
Claim: "AI sets the optimal retail price for every car"
Evidence: "trusted by hundreds of dealers" (customer count, not accuracy)
Question: How often is your suggested price within 3% of the final sale
price, on how many cars, and can we test it on 30 of ours?
Claim: "Learns from your sales"
Evidence: none
Question: What changes in the pricing after a car sells, and when?
One fix was needed. In an earlier run the assistant had accepted "learns from your sales" as supported because the page mentioned "continuous improvement". Marketing wording isn't evidence, so the prompt now says to quote evidence only if it includes a method, data or a number. The assistant is good at finding claims; deciding what counts as proof stays with you.
Scoring the checklist and what to do with a low score
Score each of the 16 items 2 for a clear yes, 1 for partial and 0 for no, for a maximum of 32. As a rule of thumb:
- 24 or more: the claims hold up. Decide on price and fit as normal.
- 14 to 23: the product may be fine but the AI story is inflated. Negotiate the price towards what the tool is worth without the AI label, and get the real method written into the contract.
- Under 14: walk away, or trial it on a monthly plan with a written exit plan and no customer-facing use until it proves itself on your data.
The dealership's pricing tool scored 18: weak on written claims and evidence, strong on the demo and the numbers. Its filled-in sheet:
| Items | Scores | What decided them |
|---|---|---|
| 1-4: claims in writing | 0, 1, 0, 1 = 2 | No method in the docs; no model providers needed or named; no AI feature list; release status shown |
| 5-8: live demonstration | 2, 2, 1, 1 = 6 | Ran on 30 of the dealer's cars; coped with a rare import; showed a low-confidence flag; no AI-off comparison available, but the manager's rule served as one |
| 9-12: evidence | 0, 2, 1, 1 = 4 | No published accuracy method; beat the manager's rule; one similar-sized reference; a few independent dealer reviews |
| 13-16: price and contract | 2, 2, 1, 1 = 6 | Premium itemised at $300; no usage limits; method written in only after asking; the manager still approves every price |
It sat squarely in the middle band, which is where its buying decision ended up too. If you want the questions that go beyond AI claims, such as security, support and data handling, the small-business AI vendor scorecard covers those, and red flags in AI marketing software has more examples from one crowded category.
When you ask for evidence, a short written request keeps the conversation on facts. This one worked for the dealership:
Subject: Evidence for the AI Pricing tier before we decide
Thanks for the demo. Before we commit, please send:
1. A written description of how the price suggestion is calculated.
2. Any accuracy figures you publish, with the data and period used.
3. Prices for our plan with and without the AI Pricing tier.
4. Confirmation of the notice you give if the method changes.
We'd also like to run the suggestion against 30 of our recent sales.
A vendor with a real product answers that email quickly. One selling a label tends to reply with another case study.
Finally, keep a record of what you were sold. Save a PDF of the feature page, the pricing page and the vendor's written answers on the day you sign, and file them with the contract. AI features get renamed, reworked and retired, and sales pages change without notice. If the tool stops doing what it claimed a year from now, the saved pages turn "I'm sure it used to do that" into a dated document you can put in front of the vendor at renewal.
Further reads
- Questions to Ask an AI Vendor Before You Sign Anything — The broader question list for any AI purchase.
- AI Agent vs Chatbot vs Automation: Which Does Your Business Need? — Plain definitions of the labels vendors stretch most.
- Did Your AI Pilot Work? How to Set Success Criteria That Hold Up — Set pass marks before a trial, not after.
- AI Features Already in Your Software: What to Switch On First — Check what your current tools already do first.
- What Uptime and Support Should an AI Vendor Promise You? — What support to demand once you do buy.
- What to Do Before You Buy Any AI Tool: A 10-Point Checklist — Ten checks to run before paying for any AI tool, from pricing the job it replaces to reading the exit terms, with a kitchen-fitter worked example.
- How to Read AI Software Pricing: Seats, Credits, and Usage Fees — The six pricing units AI software uses, the fine print that changes the total, and a sports shop turning three pricing pages into one monthly figure.
- How to Choose an SEO Agency That Understands AI Search — Gather your own AI search baseline, ask agencies the questions that expose GEO hype, score two proposals and know what honest monthly reporting looks like.
- AI Tools and AI Development: The Complete 2026 Guide — the AI hub, including every tutorial in the AI-for-business series.
Sources: published regulatory actions on misleading AI claims (2024); HubSpot product notices on retired Breeze agents; Intuit notes on QuickBooks AI settings. Checked September 2026.