Send the vendor your own scenarios a few days ahead and ask to see the live product loaded with your content, not slides or a recording. In the demo, type questions yourself, include three or four they haven't seen, ask to see the admin screens and a mistake being corrected, then get pricing terms in writing before any trial.
A standard sales demo is built to show the product at its best: clean sample data, questions the presenter has run a hundred times, and a smooth path through the features that sell. None of that is dishonest, but none of it tells you how the product copes with your customers, your data and your awkward cases. The whole job is to swap their script for yours without making the call adversarial.
What a scripted demo hides, and how to get past it
AI products are especially easy to demo well and especially hard to judge from a demo. A chatbot that answers ten rehearsed questions perfectly can still invent a policy on the eleventh. A document reader shown a crisp sample PDF may struggle with your scanned, photographed paperwork. The usual gaps between the demo and the product you'd get:
- Sample data instead of yours. Their demo account holds tidy, complete records. Your booking system has gaps, duplicates and free-text notes.
- Prepared prompts. The presenter may be pasting questions from a document that are known to work.
- Features that aren't finished. "That's rolling out next quarter" can cover the one feature you need.
- Integrations shown as slides. A diagram of the connection to your booking system is not a working connection.
- Only the customer's view. The chat window looks good; the screens your staff would use every day stay hidden.
Each has a simple counter: your data, your hands on the keyboard, a live connection, and a request to see the admin side. Before any of that, it helps to know whether the product is a thin layer over a general model; checking whether an AI tool is a ChatGPT wrapper covers the questions for that.
The brief to send three days before the demo
The brief sets the agenda. Vendors generally welcome it, because a prepared prospect is a serious one. Here is one from an illustrative 120-pitch campsite with glamping pods, looking at AI tools to answer guest messages:
Subject: Demo on Thursday: what we'd like to see
Thanks for setting this up. To make the hour useful, please:
1. Load the attached policies (arrival times, dogs, EV charging,
pitch sizes, cancellations) into a trial account before the call.
2. Show the live product, not slides or a recording. We'd like to
type some questions ourselves.
3. Use the 12 anonymised guest questions attached. We'll also bring
a few new ones on the day.
4. Show the connection to our booking system live, or tell us in
advance if that's not possible.
5. Show the admin screens: updating a policy, reviewing
conversations, and how a hand-off to staff works.
6. Bring the person who configures the product for customers.
7. Leave 10 minutes for pricing: what counts as a billable
conversation or resolution, minimum terms, and cancellation.
Attending from us: owner and reception manager.
The before-and-after is stark. The campsite's first request to another vendor had been "Could you show us your AI chatbot?", which produced a 45-minute presentation of features for hotel chains, a sample conversation about spa bookings and no time for questions. The brief above produced an hour spent entirely on the campsite's own policies.
An hour-long demo agenda that puts your scenarios first
Let the vendor lead for the first few minutes; you'll learn how they see their product. Then move to your agenda. A shape that works:
- Minutes 0 to 10: their overview, kept short. Note which features they lead with.
- Minutes 10 to 30: your 12 prepared questions, typed by you or read out and typed by them while sharing their screen, in the product as customers would see it.
- Minutes 30 to 40: the unseen questions and edge cases.
- Minutes 40 to 50: the admin side: fix a wrong answer, change a policy, review yesterday's conversations, see a hand-off arrive with a member of staff.
- Minutes 50 to 60: pricing, contract terms, data handling and next steps.
If the vendor can't fit this into an hour, ask for a second session rather than cutting the edge cases or the admin view. Those two blocks tell you most of what you need.
Questions to ask during a software demo, grouped by what they expose
Keep this list open during the call and tick off what you've seen with your own eyes, not what you've been told.
Is this the real product?
- Is this the version we'd get on the plan you're quoting, or does it include features from a higher tier?
- Which of the features shown today are in beta or not yet released?
- Can I type the next question myself?
Will it work with our data and systems?
- Can you show the connection to our booking system working now, with a real availability check?
- What happens when a record is incomplete, for example a booking with no arrival time?
- How long after we update a policy does the AI use the new version?
What happens when it's wrong?
- Show me how a member of staff corrects a wrong answer, and how long that takes.
- When does it hand over to a person, and what does the person see?
- Can we see a log of every answer it gave and what it based the answer on?
What will it cost, and what are we committing to?
- Exactly what counts as a billable conversation, resolution or credit?
- What's the minimum term, and what does cancelling involve?
- What did a customer with our volume pay last month, all in?
What happens to our data, and if we leave?
- Is our content or our guests' messages used to train models?
- How long are conversations kept, and can we delete them?
- Can we export our knowledge base and conversation history in a usable format?
These are the demo-day questions. The deeper checks on security, contracts and references come afterwards; evaluating an AI software vendor with a scorecard covers that stage.
Testing with questions the vendor hasn't seen
The unseen questions are the most revealing part of the hour. Pick ones your staff really get, including awkward ones. The campsite brought four:
- "Can we arrive at 11pm? Our ferry gets in late."
- "Is it OK to bring a trailer tent on a standard pitch?"
- "My daughter booked for us, can you tell me which pitch she's on?"
- "The shower block was filthy last night and I want some money back."
One vendor's product, loaded with the policies, answered the first with "Yes, late arrivals are welcome! Just let us know your expected time." The policies said the gate closes at 10pm and late arrivals must use the overflow car park until morning. That's the kind of answer that looks fine in a demo and causes an argument at the barrier at 11pm. What to note: not just that it was wrong, but what happened next. Did the presenter fix it live in the admin screen, and did the fix work on a reworded version of the same question? It did, in about two minutes, which counted in the vendor's favour.
The third question tests privacy: the right answer is a polite refusal and an offer to contact the person who booked. The fourth tests escalation: anything involving money or a complaint should go to a person, quickly and with the conversation attached. If you'd like a larger set of edge cases, you can ask a chat assistant to draft them from your policies:
Here are our campsite policies. Write 15 guest questions that are
likely to trip up an AI assistant: questions the policies only
partly answer, questions involving money or complaints, requests
for another guest's details, and late or unusual arrivals. Use the
casual wording real guests use in messages. Don't include questions
the policies answer in one obvious sentence.
[paste policies]
An illustrative result gave a useful set, but a few needed fixing: two asked about a swimming pool the site doesn't have (the model had assumed one), and several were written in tidy, formal English that no guest uses at 10pm from a car. Replace those with real examples from your own inbox, stripped of names, and keep the ones that probe gaps in the policies.
The admin screens matter more than the chat window
Your staff will spend far more time in the admin view than any guest spends in the chat window, so ask to drive it. Four things to see happen live:
- A correction. Staff fix a wrong answer, and the fix holds when the question is asked differently.
- A policy change. Update the dog charge and ask about dogs again. How long did it take to apply?
- A hand-off. A complaint arrives with a member of staff. Can they see the whole conversation, reply from the same place, and mark it done?
- The review screen. Yesterday's conversations, filterable by "handed off", "customer unhappy" or "no answer found".
If the answer to any of these is "our onboarding team does that for you", ask what it costs each time and how quickly it happens. A product your own staff can't correct becomes a product you pay someone else to babysit.
Pricing questions: what counts as a resolution or a credit
Many AI support tools charge per conversation or per "resolution", and the definition decides the bill. Real examples from vendors' own pages: Intercom charges $0.99 per resolved outcome for its Fin agent and counts an "assumed resolution" when a customer goes quiet for 24 hours after Fin's last answer; Zendesk closes a messaging conversation after two hours of inactivity by default (72 hours on email and web forms) and, since 18 May 2026, bills only resolutions its AI check verifies; Help Scout's AI Answers cost $0.75 per resolution. A guest who gives up and phones instead may count as "resolved". Silence isn't satisfaction, so ask how the vendor you're seeing defines it, and whether you can dispute charges for conversations that were handed off or abandoned.
Then do the sum at your own volume. In the campsite's peak season, about 600 guest messages a month turn into roughly 450 conversations. At an illustrative $0.90 per resolution, with 70% counted as resolved, that's 315 resolutions, or about $284 a month; if the definition counts abandoned chats and the rate rises to 85%, it's about $344. Against a flat-fee rival at, say, $199 a month, the answer depends entirely on that definition. The tutorial on outcome-based AI pricing goes further into these models.
Demos of AI that isn't a chatbot
The same principle, your inputs rather than theirs, applies to every kind of AI product. For a document tool, bring five of your own documents, including the worst one: the scanned form with a coffee ring, the supplier invoice in an odd layout. Ask them to process it live and show you the fields it extracted, including the ones it wasn't sure about.
For a forecasting or pricing tool, bring history you already know the answer to. A boutique hotel looking at an AI pricing tool, for instance, might send last season's occupancy and rates and ask: what would the tool have recommended for three specific dates, a quiet Tuesday, a festival weekend and the night of a big local wedding? Comparing its suggestions with what actually happened is far more telling than any chart of projected revenue uplift. If the vendor can't run your history before you sign, that tells you how much of the demo was built for the demo.
Signals worth writing down during the call
Keep a short list of observations as the call goes, because they fade quickly once the sales follow-up starts. Any of these deserves a note and a follow-up question:
- You're not allowed to type, or the presenter switches to slides when you ask to see something live.
- "We'll load your data after you sign" in reply to a brief sent days earlier.
- Key features described as "coming soon" or "in beta" without a date.
- The pricing unit is vague: "standard definition", "fair use", "depends on usage".
- A discount that expires within days of the demo.
- The word "AI" covering something that turns out to be fixed buttons and menus; spotting AI washing in software has the tests for that.
Scoring two demos for a 120-pitch campsite
The campsite's owner and reception manager scored both demos straight after each call, before discussing them, then compared notes:
| What was checked | Vendor A | Vendor B |
|---|---|---|
| 12 prepared questions answered correctly | 11 of 12 (said dogs were free; they're charged per night) | 12 of 12 |
| 4 unseen questions handled well | 3 of 4 (late arrival wrong, fixed live in 2 minutes) | 1 of 4 (invented a partial refund on the complaint) |
| Booking system connection shown live | Yes, real availability check | No: "our integration team sets that up after contract" |
| Staff could correct an answer themselves | Yes | Only by emailing support |
| Resolution definition given in writing | Yes, abandoned chats excluded | "Standard industry definition" |
| Minimum term | Monthly | 12 months |
Vendor B's perfect score on the prepared questions is the classic sign of a well-rehearsed demo. Its performance on the unseen questions, the missing integration and the vague pricing definition told the real story. Vendor A made more visible mistakes, but it showed them being fixed, which is what everyday use looks like.
From demo to a trial on your own traffic
A good demo earns a trial, not a contract. The campsite asked Vendor A for two weeks in which the AI drafted replies to real guest messages for reception to approve, so no guest saw an unchecked answer. Over the first week, of 200 real questions, 176 drafts were correct as written (88%), 16 were rightly handed off and 8 were wrong, mostly about the new EV charging prices that hadn't been added to the policies. After updating the policies, the second week reached 94% correct. Reception timed their handling at about 4 minutes a message before and 1.5 minutes with approved drafts: in a 600-message month, that's roughly 25 hours saved. The method is set out in piloting AI in shadow mode before customers see it, and the campsite-specific uses are in how campsites and glamping sites use AI to answer guests.
The trial also catches the mistake a demo can't. Picture a letting agency that bought a document-reading tool after a flawless demo on the vendor's sample inventory reports. Its own reports were scanned, with handwritten notes and photos, and in the first week about a third of the AI summaries missed rooms entirely. The problem showed up only when a tenant disputed a deposit deduction and the summary had no record of the room in question. A one-week trial on ten real reports would have shown it before signing. Before any contract, it's also worth speaking to a customer of a similar size who has used the product for at least a season.
AI software demos: common follow-up questions
Should I send my questions to the vendor before the demo?
Send most of them, and keep a few back. Sending scenarios in advance lets the vendor load your content, so you see the product working on your business rather than theirs. Holding back three or four questions shows you how it behaves on something nobody prepared for, which is closer to what customers will do.
Who from my business should attend an AI software demo?
The person who will run it day to day, not just the owner. They'll ask about the admin screens, the corrections and the hand-offs that decide whether it fits the job. Keep the group small; two or three people who have read the brief ask better questions than six who haven't.
Is a recorded demo video enough to judge an AI product?
Not on its own. A recording shows the product on a good day with prepared inputs. Use videos to decide who deserves a live demo, then insist on the live session with your scenarios and a few unseen questions before you consider a trial.
How long should a trial be after a good demo?
Long enough to see normal traffic and a few awkward cases: often two weeks, and longer if your enquiries are seasonal. Agree the success measures before it starts, and run it in a way customers don't see, such as drafts for staff approval, until the numbers are good enough.
Further reads
- Questions to Ask an AI Vendor Before You Sign Anything — The contract and data questions to settle after the demo.
- What to Ask an AI Chatbot Vendor Before You Sign Up — Extra questions specific to customer-facing chatbots.
- How to Check an AI Software Vendor's Customer References — Make reference calls with similar customers count.
- What Uptime and Support Should an AI Vendor Promise You? — What uptime and support terms to ask for next.
- How to Measure Whether Your AI Chatbot Is Actually Working — The measures to track once a trial is running.
- How to Read AI Software Pricing: Seats, Credits, and Usage Fees — Decode seats, credits and usage fees in the quote.
- Choosing a Mortgage CRM With AI Built In: What Actually Matters — A demo-ready checklist for mortgage CRMs with AI: document extraction, chasing, client-bank alerts, compliance logs, integrations and exit terms.
- Choosing Nursery Software With AI: What to Compare — What to compare when nursery software offers AI: children's data handling, role controls, a demo test with one observation, and a scored choice.
- Best AI Estimating Software for Small Electrical Firms (2026) — AI takeoff and estimating tools for electrical contractors compared on what the AI does, price, and fit, with a test method using an old job.
- How to Run a Two-Week AI Tool Trial Before You Commit — A day-by-day plan for trialling an AI tool on real work, with a test-case list, a daily log, a filled-in scorecard and the cancellation steps people forget.
- AI Tools and AI Development: The Complete 2026 Guide — the AI hub, including every tutorial in the AI-for-business series.
Sources: Intercom pricing and help pages on Fin resolutions; Zendesk help pages on automated resolutions; Help Scout pricing (checked September 2026).