To run a two-week AI tool trial, pick one job, write pass marks before you start, and collect 20 real examples to test with. Use the tool on real work every day in week one with a two-minute log. In week two, push it with awkward cases and a data-export test. Score it on day 13 and cancel or convert before billing starts.
Most trials fail for one reason: nobody decided in advance what "good enough" means. The tool gets a few impressive demos in the first two days, then nobody uses it because the day job takes over, and on day 15 the card is charged. Two weeks is plenty of time for a clear answer, but only if the trial is planned like a small project. The plan below takes about two hours to set up and 10-15 minutes a day to run.
Before day one: two hours of preparation
Choose one job, not the whole tool
AI tools do many things. Your trial tests one: answering phone calls out of hours, drafting customer updates, summarising supplier emails, writing product descriptions. Pick the job that prompted the purchase. If the tool can't do that one well, its other features don't matter.
Record how the job is done now
Time the job for three or four instances this week, and note its quality: how many need correcting, how long customers wait. That's your baseline. Our tutorial on piloting AI in shadow mode explains how to run a customer-facing tool alongside your current process without risk.
Write the pass marks
Three to five measurable conditions the tool must meet. For an illustrative car repair garage trialling an AI phone receptionist for calls outside opening hours:
The trial passes only if ALL of these are true:
1. At least 17 of 20 test calls are handled correctly
(right information, right next step, booking logged).
2. It never quotes a repair price or diagnoses a fault.
3. Every call it can't handle ends with a clear message
and a callback request in our inbox within 5 minutes.
4. Two of our three staff want to keep it.
5. The monthly cost at our real call volume is confirmed
in writing.
Pass marks like these make the final decision a check rather than a debate. For more on setting them, see how to set AI pilot success criteria that hold up.
Collect 20 real test cases
Real ones, from the last month: actual emails, call notes, customer questions, documents. Include easy ones, typical ones and at least five awkward ones. Remove or change personal details if the tool hasn't passed your data checks yet.
Do the data checks now, not after
- Is your data used to train the vendor's models, and can you switch that off?
- Is there a data processing agreement you can sign?
- How do you export what you create, and how do you delete everything if you leave?
- What does the tool need access to? Give it the minimum.
Questions to ask an AI vendor before you sign has the full list.
Diary the billing date
Many trials convert to paid automatically. Microsoft's one-month Copilot Business trial becomes a paid subscription at the end; OpenAI's help centre says ChatGPT trials renew unless cancelled, and recommends cancelling at least 24 hours before. HubSpot's Customer Agent can be tried for 14 days with a Professional or Enterprise seat. If there's no free trial, buy one month on monthly billing with the fewest seats allowed. Put the decision on day 13 and the cancellation deadline on day 14 in the diary now.
Put the plan on one page
Before day one, write the whole trial on a single page and share it with everyone involved. An illustrative version for the garage:
Trial: out-of-hours AI phone receptionist
Dates: Mon 6 Oct to Sun 19 Oct. Decision: Sat 18 Oct.
Cancel by: Sun 19 Oct, 5pm (renews Mon 20 Oct).
Owner: service adviser (15 min a day allowed).
Users: owner, service adviser, workshop manager.
Job tested: calls after 5.30pm and at weekends only.
Baseline: about 12 voicemails a week; 3 never called back.
Pass marks: see list (5 conditions).
Test cases: 20, in the shared folder "Trial calls".
Data: recordings deleted after 30 days; training off;
export tested on day 11.
It takes ten minutes and prevents most of the "I thought you were doing that" problems that sink trials in the second week.
Days 1-2: set up and run the test cases
Time: two to three hours in total.
Set the tool up the way you'd use it for real: your documents, your tone, your rules. Connect only what the chosen job needs. Then run all 20 test cases and record the result for each: right, nearly right (light edit needed), wrong.
An illustrative first round for the garage's phone receptionist, using staff playing callers from real call notes:
| Result | Count | Examples |
|---|---|---|
| Right | 13 | Opening hours, booking a service, "is my car ready?" (took a message correctly) |
| Nearly right | 4 | Booked a slot on a bank of time the garage keeps free for breakdowns |
| Wrong | 3 | Estimated a clutch replacement "usually around $800"; didn't recognise a caller wanting to cancel; kept a caller with a roadside breakdown on the line asking for their registration |
Thirteen of 20 is short of the 17 pass mark, but that's normal on day one. The price estimate is the serious one: it breaks pass mark 2. Fix the setup (add "never give prices; say a technician will call back with a quote", block the reserved slots, add a roadside-breakdown instruction to give the recovery number immediately) and rerun the three failures. Don't move on until the wrong answers are fixed or you know they can't be.
Days 3-7: use it on real work every day
Time: the job itself, plus two minutes a day for the log.
The people in the trial use the tool for every instance of the chosen job. For a customer-facing tool, keep a person in the loop: the AI answers, but staff check every call summary or draft before anything is promised. Each person keeps a short log.
Daily trial log - phone receptionist (illustrative)
Date | Calls | Handled well | Needed fixing | Time saved | Notes
Mon | 6 | 5 | 1 | 25 min | Took a
| | | | | booking for
| | | | | a Sunday
Tue | 4 | 4 | 0 | 20 min |
Wed | 9 | 7 | 2 | 30 min | Two callers
| | | | | asked for
| | | | | "the owner"
Thu | 5 | 5 | 0 | 25 min |
Fri | 7 | 6 | 1 | 25 min | Misheard a
| | | | | phone number
Two minutes a day is enough. The point is to capture problems while they're fresh, because by day 13 nobody remembers Wednesday.
If real volume turns out too low to learn from, say a quiet week brings only three out-of-hours calls, top it up. Have a colleague replay five of the test cases as live calls on two evenings, logged the same way. A trial that sees too little real work ends with a shrug, and a shrug usually means paying for another month to find out.
Day 8: a 15-minute midpoint check
Look at the logs together and ask three questions:
- Is it being used every day? If not, why not? Fix the reason (it's slow to open, nobody knows the login, it doesn't fit the workflow) or accept the trial has failed.
- Are the same problems repeating? Change the setup now, while there's time to see if the fix works.
- Is anything a deal-breaker already? If the tool has broken a must-never rule twice, you can stop early and save the rest of the fortnight.
For the garage, the repeated issue was misheard phone numbers. The fix was a setting to read the number back to the caller. That's the kind of change the midpoint exists for.
Days 9-12: push it on purpose
Time: about an hour in total, spread over the four days.
Week two is for the cases that real use doesn't throw up often enough. Test deliberately:
- Awkward inputs: for the garage, a caller with heavy background noise, one who swears, one who asks "am I talking to a robot?", one who wants a specific technician, one calling about a car already in the workshop.
- Volume: what happens at your busiest hour, or when you hit a usage limit or credit cap?
- Failure: switch off something it depends on (a calendar connection, a document) and see whether it fails safely, with a clear message, or confidently.
- Support: send the vendor one real support question. How long does the answer take, and is it useful?
- Exit: export your settings, documents and history. Can you get them out in a usable form?
An illustrative independent bookshop trialling an AI writing assistant for its newsletter ran a different set of stress tests: a 60-page publisher catalogue to summarise, an event listing with a missing date, a request to write in the voice of a named author (it should decline or keep it generic), and a newsletter with a price change halfway through the month. Different tool, same principle: find where it breaks before a customer does.
Day 13: score it
Time: 30 minutes.
Go back to the pass marks first. Then fill in a scorecard so you can compare with any other tool you trial. An illustrative scorecard for the garage's phone receptionist at the end of the fortnight:
| Criterion | Weight | Score (1-5) | Weighted | Evidence |
|---|---|---|---|---|
| Accuracy on the 20 test cases (after fixes) | 30% | 4 | 1.2 | 18 of 20 right on the final run |
| Never breaks must-never rules | 20% | 5 | 1.0 | No prices or diagnoses after day 2 fix |
| Time saved on real work | 20% | 4 | 0.8 | About 2 hours a week of callbacks avoided |
| Staff want to keep it | 15% | 4 | 0.6 | 2 of 3 yes; 1 unsure |
| Data, export and support | 15% | 3 | 0.45 | Export works; support took 2 days |
| Total | 100% | 4.05 out of 5 |
All five pass marks were met, and the monthly cost at the garage's volume was confirmed by email. A weighted score above 3.5 with every must-never rule met is a reasonable bar for converting. Below 3, cancel. In between, decide whether a specific, fixable problem is holding it back.
Day 14: decide, and act on the decision
There are three outcomes, and each has a short list of actions.
- Convert. Stay on monthly billing for another month or two while use settles, then consider annual. Write down the setup changes you made during the trial, so the next person can maintain them.
- Extend. Only with a named reason and a date: "the booking integration arrives next week; extend one week to test it". An extension without a reason is a slow way of saying no.
- Cancel. Before the deadline. Export anything you want, delete your data through the settings or by written request, disconnect it from email, files and calendars, and revoke any access keys. Keep the scorecard; it's useful when the next vendor calls.
What a trial that fails early looks like
Stopping a trial early is a good result if it saves you a year of the wrong tool. Take an illustrative funeral director trialling an AI note-taker for arrangement meetings with families, hoping to cut the hour of typing up notes after each one.
Days one and two went well on the test cases, which were role-played meetings between staff. On day four, with a real family, two problems appeared. The arranger felt she had to explain the recording at the start of a conversation with a grieving family, and one family declined, which is their right. And on the meetings that were recorded, the summaries got names of relatives and hymn titles wrong often enough that every note needed checking line by line, which took nearly as long as typing it.
At the day-eight midpoint, the log showed four real meetings, one declined, and an average of 40 minutes spent correcting each summary against a target of 15. The owner also checked the vendor's retention settings and found recordings were kept for 90 days by default. The trial stopped there. The better answer for this business was a structured arrangement form on a tablet, with an AI assistant tidying the typed notes afterwards, which kept recordings out of the room entirely.
The lesson isn't that note-takers are bad. It's that a trial on real work, with a midpoint, surfaces the problems that demos and role-play never do.
What the trial itself costs
A proper trial isn't free, even when the software is. Here's the internal time for the garage's fortnight, at an illustrative $30 an hour:
| Activity | Time | Cost |
|---|---|---|
| Preparation, pass marks and test cases | 2 hours | $60 |
| Setup and first test round | 2.5 hours | $75 |
| Daily log and checks, 10 working days | 2.5 hours | $75 |
| Midpoint, stress tests and scoring | 2 hours | $60 |
| Total | 9 hours | $270, plus one month's subscription if there's no free trial |
Compare that with the cost of committing to the wrong tool: an illustrative $60 a month on an annual plan is $720 for a year, and a tool on three seats at $50 each is $1,800. A $270 trial that prevents either is money well spent, and one that confirms the right choice means you start with a setup that already works.
Four ways trials go wrong
- Trialling on demo data. A tool that's excellent on the vendor's sample documents can struggle with your scanned delivery notes or your customers' questions. Use your own material from the first day. Running an AI software demo so you see the real product covers the same point for the sales stage.
- Too many people, too many jobs. An illustrative florist gives seven staff access to test "whatever they like". By day 14 there are seven opinions and no data. Two or three people on one job gives a clearer answer.
- Nobody owns it. Name one person who runs the log, the midpoint and the scoring, and give them the 15 minutes a day to do it.
- Judging on the best day. The first impressive answer sticks in the memory. The log and the scorecard are there to stop one good afternoon deciding a year's spend.
If you're comparing two tools, run them one after the other on the same 20 test cases, or side by side on the same real work, and use the same scorecard. For customer-facing chat tools specifically, testing a customer chatbot before it goes live adds checks this plan doesn't cover, and evaluating an AI software vendor scores the company behind the tool as well as the tool itself.
Trial questions worth settling early
What if the AI tool has no free trial?
Buy the smallest plan on monthly billing, with the minimum number of seats, and treat the first month as the trial. Put the renewal date in the diary the day you sign up. Avoid annual billing and lifetime deals until the tool has passed your scorecard, however large the discount.
Is two weeks long enough to judge an AI tool?
For a tool that does one clear job, yes, if you use it on real work every day and test awkward cases deliberately in the second week. For something that changes a whole process, such as a CRM with AI features, two weeks tells you whether to continue into a longer pilot, not whether to commit for a year.
Should customers be involved in the trial?
Not directly at first. Run customer-facing tools in shadow mode, where the AI drafts or suggests and a person decides what goes out. Only let it reply to customers directly once it has passed your test cases, and tell customers they're dealing with an AI where that applies.
What should happen to our data after a trial?
Before you start, find out how to export anything you create and how to delete your data if you don't continue. At the end of an unsuccessful trial, export what you want to keep, delete the rest through the tool's settings or a written request, and remove any connections to your email, files or other systems.
Further reads
- Annual or Monthly Billing for AI Tools: Which Saves You More? — When the trial passes: whether to switch to annual.
- AI Vendor Lock-In: How to Keep Your Data and Prompts Portable — Keep your prompts and data portable from day one.
- What If Your AI Vendor Shuts Down? Checks Before You Commit — Checks on the supplier, not just the software.
- What to Check in an AI Vendor's Data Processing Agreement — What to check in the vendor's data terms.
- Why Your AI Pilot Stalled, and How to Get It Live — If the trial passes but nobody keeps using it.
- How to Check an AI Software Vendor's Customer References — Talk to real customers before a big commitment.
- AI Tool Approval Process: How Staff Request a New AI Tool — The request form staff fill in, the checklist the reviewer works through, and the three replies, so new AI tools get approved in days, not months.
- What to Do Before You Buy Any AI Tool: A 10-Point Checklist — Ten checks to run before paying for any AI tool, from pricing the job it replaces to reading the exit terms, with a kitchen-fitter worked example.
- Is an AI Scribe Worth It for a Small Private Clinic? — When an AI scribe pays off in a small clinic, three clinic types with different verdicts, a two-week test scorecard and what to check in every draft note.
- AI for Veterinary Practice Owners: Costs, Tools, and First Steps — What AI costs a vet practice in 2026, which tools do which jobs, budgets for one-, three- and five-vet practices, and a six-week plan to start safely.
- Best AI Estimating Software for Small Electrical Firms (2026) — AI takeoff and estimating tools for electrical contractors compared on what the AI does, price, and fit, with a test method using an old job.
- AI Video Survey Tools for Movers Compared: Cost and Accuracy (2026) — Six video survey tools for removal firms compared on published price, pricing model and what their accuracy claims actually measure.
- AI Tools for Physiotherapy Clinics: What Each One Actually Does — What each AI tool a physio clinic might buy actually does, what it doesn't, and which two to start with.
- Best AI Business Intelligence Tools for Small Businesses (2026) — Ten reporting and analytics tools compared on what their AI really does, list prices, the tier each AI feature needs and which small businesses each suits.
- Is Claude Team Worth It for a Small Business? — When Claude Team earns its $20-$25 a seat over separate Pro accounts, three businesses that get different answers, and how to make a shared project pay.
- Which AI Is Best for a Small Business? — Choose an AI assistant through a practical task trial, with realistic prompts, cost checks and examples from service businesses.
- AI Tools and AI Development: The Complete 2026 Guide — the AI hub, including every tutorial in the AI-for-business series.
Sources: Microsoft 365 Copilot for business page (one-month trial that converts to paid); HubSpot knowledge base (Customer Agent 14-day trial with a Professional or Enterprise seat); OpenAI help centre on free trials and cancelling before renewal; checked September 2026. Business examples are illustrative.