Score every vendor on one weighted scorecard: fit in a trial on your own data, data handling, two-year cost, how easily you could leave, how likely the product is to last, and support when it's wrong. Knock out any vendor that fails on data handling, however well it scores elsewhere.
The trial matters most because demos mislead in a predictable way: they run on tidy sample data and on the questions the product answers best. A vendor that scores well on your own messy records and averagely on everything else usually beats one with a flawless demo and vague answers about where your data goes. A scorecard makes that trade-off visible instead of leaving it to whoever gave the best presentation.
Knock-out questions to answer before any scoring
Ask these first, in writing. A "no" to any of them ends the evaluation, which saves you scoring a vendor you could never sign.
- Will you use our data to train or improve AI models, and can we have a business-level opt-out in the contract? Business plans from the large AI providers don't train on business content by default, so a smaller vendor building on them has no good reason to insist.
- Can we export all our data in a standard format, such as CSV or JSON, without asking permission? If leaving needs the vendor's help, you can't leave on your own terms.
- Will you provide a data processing agreement and a current list of sub-processors? Sub-processors are the other companies that handle your data for the vendor, including any AI model provider behind the product.
- Did the product complete our most important task in a trial on our own data? Not a demo, not a sample account: your records.
- For anything customers see: can we switch the AI off straight away if it starts giving wrong answers? Not every product lets you. Intuit says QuickBooks Online's AI features can't currently be switched off one by one, which is tolerable in bookkeeping and wouldn't be in a customer chat.
The knock-outs deliberately cover data and control rather than features. Missing features show up in the scores; a data problem shows up later, after you've moved everything across.
The scorecard: six groups weighted for a small business
The weights add up to 100. Score each item from 1 to 5, average the items in a group, then multiply by the group's weight divided by 5. The weights below suit most small businesses; move up to 10 points between groups if your case is unusual (a customer-facing tool might take points from cost and give them to support), but set the weights before you see any demos, or you'll bend them towards the vendor you already like.
Fit on your own work: 30 points
- Completes your top three tasks in a trial. Why: this is what you're paying for. Verify: run ten real cases through it over two weeks, including three awkward ones, and time the edits each output needs.
- Shows its uncertainty. Why: a tool that says "I'm not sure" is safer than one that guesses fluently. Verify: include two questions your records can't answer and see whether it admits it.
- Connects to the systems you use now. Why: integrations on a slide often turn out to be "on the roadmap". Verify: see the integration running with your own account during the trial.
- Usable without help after an hour. Why: a tool only the keenest person can use saves one person's time. Verify: ask your least technical colleague to try the top task alone.
For getting a demo that shows the real product rather than a rehearsed path, see how to run an AI software demo so you see the real thing.
Data handling: 20 points
- Training and improvement use. Why: "we don't sell your data" is not the same as "we don't train on it". Verify: find the clause in the terms and check which plan it covers.
- Retention and deletion. Why: prompts and outputs kept for years are a liability. Verify: ask how long each is stored and how deletion works when you cancel.
- Sub-processors, including the AI model provider. Why: your data's real journey is the list of companies it passes through. Verify: the published sub-processor list, and whether you're told when it changes.
- Security evidence. Why: independent audits beat assurances. Verify: ask for a current SOC 2 Type II report (it tests whether controls actually worked over a period of months, where a Type I checks their design on a single date) or an ISO/IEC 27001 certificate, and check that its scope includes the AI features. ISO/IEC 42001, the standard for AI management systems, was only published in December 2023, so treat a certificate as a plus rather than marking a small vendor down for lacking one.
The agreement itself deserves a careful read; what to check in an AI vendor's data processing agreement goes clause by clause, and SOC 2 and ISO 27001 explained covers what the reports do and don't prove.
Two-year cost: 15 points
- Price after any promotion. Why: first-year discounts are common in business software. Verify: ask for the list price in writing, and the price in month 13.
- How usage is counted. Why: seats, credits, resolutions and overage charges behave very differently as you grow. Verify: model your busiest month, not your average one.
- Minimums and price-rise terms. Why: minimum seats or usage commitments set a floor you pay even in quiet months, and rises happen. Microsoft 365 Business Standard went from $12.50 to $14 a user a month at renewals from 1 July 2026. Verify: the notice period and any cap on rises in the contract.
Exit and portability: 15 points
- Self-service export of everything. Why: records, settings and AI conversation logs all have value. Verify: do a full export during the trial and open it.
- Renewal and notice terms. Why: auto-renewal with a long notice period can lock you in for another year. Verify: the renewal clause and the date you'd have to give notice by.
- What happens to your data if the vendor closes. Why: Clockwise, an AI calendar tool, shut down on 27 March 2026 and deleted its users' data rather than transferring it. Verify: ask what notice and export window customers would get, and look for it in the contract.
Vendor durability: 10 points
- Recent retirements and how they were handled. Verify: ask which features were retired in the past year and how much notice customers had.
- Changes to privacy defaults. Verify: ask whether any default setting affecting customer data changed in the past year.
- Customers like you. Verify: ask how many customers of your size and type use the AI features, and speak to two of them.
Support and accountability: 10 points
- A route for wrong answers. Why: every AI product gets things wrong; what matters is whether reports get fixed. Verify: report a wrong answer during the trial and see what happens.
- Response times and a named contact. Verify: the support terms for your plan, not the top tier's.
- Transparency duties if customers talk to it. Why: if you sell to customers in the EU, Article 50 of the EU AI Act has required chatbots to tell people they're talking to AI since 2 August 2026. Verify: the product can show that notice at the start of a chat.
What a 1, a 3 and a 5 look like
Two people scoring the same vendor will disagree unless the numbers mean something specific. These anchors keep a 3 from becoming a polite 4.
| Item | Scores 1 | Scores 3 | Scores 5 |
|---|---|---|---|
| Trial on your top task | Wrong or heavily rewritten on most cases | Right on most; edits needed on the awkward ones | Right on nearly all ten, and flags what it can't answer |
| Training on your data | On by default; opt-out by email request | Off on your plan, but terms allow "aggregated" use to improve services | Off by default, in the contract, and passed down to AI sub-processors |
| Export | Only through a support request | CSV of main records; no conversation logs | Self-service export of everything, including AI logs and settings |
| Two-year price | Year-two price unknown | List price known; rises allowed with 30 days' notice | Price fixed for the term; rises capped with 60 or more days' notice |
| Security evidence | "We take security seriously" | ISO/IEC 27001 certificate or SOC 2 Type I | Current SOC 2 Type II or ISO/IEC 27001 whose scope covers the AI features |
| Support when it's wrong | Email only, no named contact | Chat support in business hours | Named contact, response-time commitment, and wrong answers get fixed |
Evidence to ask for, and answers too vague to score
Send every shortlisted vendor the same request, so the replies can be compared. A version you can adapt:
Subject: Information needed before we decide
We're comparing three systems and will decide by [date]. Please send:
1. Your data processing agreement and current sub-processor list,
including any AI model providers.
2. Whether customer data is used to train or improve models on the
plan we'd buy, and the contract clause that says so.
3. Your latest SOC 2 Type II report or ISO/IEC 27001 certificate,
with its date and scope (we can sign an NDA).
4. The list price for [number] members/users in year one and year two.
5. How we export all our data, and in what formats.
6. Any features retired in the past 12 months and the notice given.
Thank you. Replies after [date] can't be included in our comparison.
An illustrative reply of the kind that comes back more often than it should:
Thanks for reaching out! Security is our top priority. We're SOC 2 compliant, all data is encrypted in transit and at rest, and we will never sell your data. Our AI is powered by industry-leading models. Pricing is tailored to each club, so let's jump on a call!
Score what's there, not what's implied. "SOC 2 compliant" doesn't say whether it's Type I or Type II, which period it covers, or whether the AI features were in scope. "Never sell" says nothing about training. "Industry-leading models" hides the sub-processor you asked about. And a price available only on a call is a price you can't compare. That reply earns 1s on security evidence and training and leaves cost blank until the vendor puts a figure in writing. A second, polite request that quotes your numbered questions usually gets better answers; if it doesn't, that tells you something about support too. For a longer list of questions to add, see questions to ask an AI vendor before you sign anything.
A members' club scores three membership systems
Three invented vendors, one realistic decision: a sports and social members' club with about 600 members runs on a spreadsheet, a shared inbox and a part-time secretary. The committee wants a membership system with an AI assistant that answers routine member questions (guest rules, fees, booking times) and drafts the monthly newsletter. Three vendors reach the shortlist:
- Vendor A: an established membership platform, eight years old, which added an AI assistant last year. Quote: $160 a month for up to 750 members, AI included. It sent a SOC 2 Type II report, a data processing agreement, and a sub-processor list naming its AI model provider.
- Vendor B: a general helpdesk with an AI agent charged per resolved question, used alongside the club's existing spreadsheet. About $150 a month at the club's volume. Excellent data terms, but it doesn't handle renewals or bookings.
- Vendor C: an AI-first startup, 18 months old, with the best demo of the three: members could renew by chatting. A "founding price" of $99 a month for 12 months, with the year-two price "to be confirmed". Export by support request. It trains on "anonymised usage data" by default and agreed an opt-out by letter.
All three survived the knock-outs, C narrowly, because its opt-out letter counted as written. The committee's two scorers worked separately and then agreed these scores out of 5:
| Group (weight) | Vendor A | Vendor B | Vendor C |
|---|---|---|---|
| Fit on own work (30) | 4 = 24 | 3 = 18 | 5 = 30 |
| Data handling (20) | 4 = 16 | 5 = 20 | 2 = 8 |
| Two-year cost (15) | 4 = 12 | 3 = 9 | 2 = 6 |
| Exit and portability (15) | 4 = 12 | 3 = 9 | 1 = 3 |
| Vendor durability (10) | 4 = 8 | 5 = 10 | 2 = 4 |
| Support (10) | 3 = 6 | 4 = 8 | 4 = 8 |
| Total out of 100 | 78 | 74 | 59 |
The two-year costs explain part of the order. Vendor A costs $3,840 in fees plus about 20 hours of migration by the secretary ($360 at $18 an hour), so $4,200. Vendor B's fees come to about $3,600, but renewals stay in the spreadsheet at roughly four hours a month, another $1,728, so $5,328 in practice. Vendor C costs $1,188 in year one and an unknown amount in year two; if the price merely doubles, the two-year figure is $3,576, and nobody could tell the committee whether it would.
The best demo came third. Vendor C's fit score was the only 5 on the sheet, and it still lost, because a club that can't export its own member list without asking has handed a startup the keys. Vendor A won by four points, which is close enough to deserve a tie-break: the committee re-ran the top task (answering guest-policy questions from the club's own rules) on A and B, and A's handling of renewals, which B can't do at all, settled it.
Durability signals from products that changed or vanished in 2026
The durability group is the hardest to score, because you're judging the future. The best evidence is how a vendor, and the market around it, has behaved recently. A few 2026 events show what to look for:
- Features withdrawn from sale. HubSpot sunset five of its Breeze agents on 23 July 2026. Existing installs keep working, but they can't be newly installed. Ask vendors what they've withdrawn and whether current customers kept what they had.
- Products closing on a date. Kajabi's Creator Studio shuts down on 9 November 2026, and Microsoft removed its Lens scanning app from app stores in February 2026 before disabling scanning in March. A dated shutdown with an export route is the good version of this; a sudden one is the bad version.
- Privacy defaults that move. Since 16 June 2026, SimplePractice has opted new users of its note-taker in to keeping de-identified transcripts by default. Nothing was hidden, but anyone who evaluated the product before June and didn't re-check would have a wrong picture of it.
- Whole companies disappearing. Clockwise's closure, with data deleted, is the case every exit clause is written for.
So ask two direct questions and score the answers: "What have you retired or changed in the past twelve months, and how much notice did customers get?" and "Have any default settings affecting customer data changed in that time?" A vendor with a clear, dated answer scores higher than one that claims nothing ever changes. Nothing standing still is itself unlikely in this market. For the checks that protect you if the worst happens, see what to check in case an AI vendor shuts down.
Reading a vendor's terms with AI, then checking it yourself
Terms of service and data agreements run to many pages, and an AI assistant can find the relevant clauses quickly. It can also get them wrong in a convincing way. A prompt that makes checking easier:
Below are [vendor]'s terms of service and data processing agreement.
For each question, quote the exact clause that answers it, with its
section number and the plan or tier it applies to. If a question isn't
answered, write "not stated". Don't summarise beyond the quote.
1. Is customer content used to train or improve models? Can we opt out?
2. How long are prompts and outputs kept?
3. Which sub-processors receive our data, including AI model providers?
4. How do we export our data, and in which formats?
5. What notice is given before price changes?
6. What happens to our data when we cancel?
Its answer to question 1 (illustrative):
Section 7.2: "Customer Content will not be used to train or fine-tune machine learning models." Applies to: all plans. Opt-out: not needed, as training is off.
The committee's secretary opened section 7 to check the quote and found that 7.2 sat under a heading for the Enterprise plan. The club's plan was covered by 7.3, which allowed "aggregated and de-identified usage data" to be used "to improve the Services". The assistant quoted a real clause and put the wrong plan against it. That's the typical failure: accurate words, wrong context. So the rule is to use the assistant to find clauses and never to certify them. Open every quoted section in the original document, read the heading above it, and score from the document.
Running the scorecard in one week
For a purchase of this size, a week is enough if you prepare the trial material first:
- Day 1: send the evidence request and the knock-out questions to every shortlisted vendor. Set the weights and agree who scores.
- Days 2 and 3: run the ten real cases through each trial account. Record accuracy and editing minutes as you go, not from memory afterwards.
- Day 4: read the evidence that came back, check the terms, and do a full data export from each trial.
- Day 5: two people score separately, then compare. Any item where they differ by two or more points gets discussed until they agree what the evidence shows.
Keep the filled-in scorecard. At the 90-day review, score the winner again on the items you can now see properly (fit, support, real cost), and you'll know whether the evaluation predicted the experience or whether the demo still won the argument.
Further reads
- How to Check an AI Software Vendor's Customer References — How to get honest answers from a vendor's existing customers.
- AI Software Contracts: Auto-Renewals, Price Rises, Notice Periods — The contract terms that decide your two-year cost.
- AI Washing: How to Spot Software That Oversells Its AI — Spot software that oversells how much AI it contains.
- What Uptime and Support Should an AI Vendor Promise You? — What uptime and support promises are reasonable to ask for.
- AI Vendor Lock-In: How to Keep Your Data and Prompts Portable — Keep your data and prompts portable after you sign.
- Should a Small Business Build a Custom AI Tool or Buy One? — Check that buying is the right route before comparing vendors.
- Pet Grooming Software With AI: Features and Prices Compared — What the AI in MoeGo, Teddy, Groomify, DaySmart Pet and Gingr actually does, what each costs a year, and a 14-day trial script for groomers.
- Salon and Spa Software With Built-In AI: What to Compare — Five tests that separate useful salon and spa AI from a feature list, with September 2026 prices for Vagaro, GlossGenius, Boulevard, Mangomint and more.
- Gym Retention Software With AI: What It Costs and Delivers — What gym retention AI costs in 2026, what it really delivers (a ranked call list), and a 60-day holdout test to see if it works at your gym.
- Best AI Inventory Tools for Small Retailers Compared — AI inventory tools for small shops compared: Prediko, Inventory Planner, Square Plus and Shopify Sidekick, plus what to do for weighed and perishable stock.
- Choosing a Mortgage CRM With AI Built In: What Actually Matters — A demo-ready checklist for mortgage CRMs with AI: document extraction, chasing, client-bank alerts, compliance logs, integrations and exit terms.
- Questions to Ask Before Buying AI That Touches Client Data — Nineteen questions to put to any AI vendor before client files go in, with what a good answer looks like and the replies that should stop a purchase.
- AI Inventory and Check-Out Reports: A Letting Agent's Guide — How AI drafts inventory and check-out reports from photos or video, where its descriptions go wrong, and how to keep the report defensible in a deposit dispute.
- Should a Small Auto Repair Shop Use AI? A Decision Checklist — A checklist for small repair shops deciding on AI: which problems justify it, what your shop system already does, costs, trust and a scoring rule.
- AI Video Survey Tools for Movers Compared: Cost and Accuracy (2026) — Six video survey tools for removal firms compared on published price, pricing model and what their accuracy claims actually measure.
- Red Flags When Buying AI Marketing Software or Services — Fifteen warning signs in AI marketing quotes and contracts, each with what it looks like, how to check it and a marked-up deli quote.
- AI Tools and AI Development: The Complete 2026 Guide — the AI hub, including every tutorial in the AI-for-business series.
Sources: vendor help and policy pages on retired products and changed defaults (Clockwise, HubSpot Breeze agents, Kajabi, SimplePractice, Microsoft Lens, Intuit QuickBooks Online); ISO/IEC 42001 publication details; published guidance on SOC 2 Type I and Type II reports; EU AI Act Article 50 transparency duties.