Book 60 to 90 minutes with the four to eight people who will build, use and be affected by the AI tool. Tell them the launch happened six months ago and it failed. Everyone writes down why, silently, for ten minutes. Share the reasons one at a time, rank the top five, and give each a fix and a warning sign.
The technique comes from the psychologist Gary Klein, who described it in Harvard Business Review in 2007. He cites 1989 research finding that imagining an outcome has already happened, known as prospective hindsight, increased people's ability to correctly identify reasons for future outcomes by 30%. What follows is the script, the agenda and the AI-specific prompts to run one properly.
Why AI projects need a pre-mortem more than most
Most software fails loudly: it crashes, or it doesn't do the thing. AI tools fail quietly. They produce answers that look right and aren't. Staff stop trusting them and redo the work by hand without telling anyone. A usage-based bill creeps up. The vendor swaps the underlying model and the tool starts behaving slightly differently. None of these shows up on launch day.
A pre-mortem is different from filling in a risk register. A register records the risks you have already thought of. A pre-mortem gives people permission to say the ones they have been holding back: "the head receptionist will switch it off the first busy Monday", or "we're launching this because the owner saw a demo, not because anyone asked for it". Framing it as a story about a failure that already happened makes those things easier to say, because nobody is predicting doom; they are explaining history.
Who to invite, and the one-page brief they read first
Four to eight people. More than that and the quiet ones stop talking.
A two-person business can still run one, with two adjustments. Take a florist run by a couple, planning to let an AI assistant take wedding enquiries by web chat. With only two voices, the silent-writing step matters even more, because each partner knows how the other will react. Borrow the missing perspectives: before the session, spend ten minutes on the phone asking a trusted regular customer "how would this annoy you?", and ask the freelancer who helps at weekends the same. In this case the regular said, "I'd want to know it won't quote me a price for flowers that aren't in season," which neither partner had written down, and it became the assistant's first rule.
- The person who decided to do the project (usually the owner).
- Whoever will configure or build it, including an outside supplier's lead if you have one, though they shouldn't run the session.
- Two or three people who will use it every day.
- The person who deals with complaints.
- At least one sceptic. If nobody on the team is doubtful, you haven't asked around enough.
Send this brief two days before. Everyone should read the same description, or the session turns into an argument about what the project is.
PRE-MORTEM BRIEF: [project name]
What goes live: [one sentence]
Launch date: [date]
Who uses it: [roles]
What it WILL do: [3 bullets]
What it will NOT do: [3 bullets]
Success at 90 days: [one or two numbers, e.g. "missed calls under 5%"]
Budget: [set-up + monthly]
Tools and data it touches: [systems, and what kind of data]
Filled in for an illustrative three-person online plant shop about to switch on the AI agent in Shopify Inbox, it might read:
PRE-MORTEM BRIEF: AI replies in Shopify Inbox
What goes live: Shopify Inbox's AI agent answers customer
chat on its own, 24 hours a day
Launch date: 1 Nov (before the winter gift rush)
Who uses it: owner, packer, part-time customer service
What it WILL do: - answer order-status and delivery questions
- explain our returns policy
- answer basic care questions from our guides
What it will NOT do: - issue refunds or replacements
- diagnose a dying plant from a photo
- promise delivery dates in the gift rush
Success at 90 days: half of chats closed without a person;
no rise in "wrong answer" complaints
Budget: no extra software cost; ~2 hours a week
reading transcripts
Tools and data it touches: order history, customer names and
emails, our policy and care-guide pages
The "will NOT do" lines earned their place in the session. Shopify's agent can use web search as a secondary source, so one early failure story was "it gave a care tip from some other website that contradicted our own guide, and the customer's plant died". That became a pre-launch test: twenty care questions, each answer checked against the shop's guides.
The 75-minute agenda
| Minutes | Step | What the facilitator does |
|---|---|---|
| 0-5 | Set the scene | Reads the failure statement below. Confirms nothing said will be held against anyone. |
| 5-15 | Silent writing | Everyone writes at least five reasons, one per card or line. No talking. |
| 15-40 | Round-robin | One reason per person per turn, going round until the lists run out. Records every one. Allows questions of clarification only, no debate. |
| 40-50 | Cluster and rank | Groups duplicates. Each person gets three votes. The group rates the top clusters 1-3 for likelihood and 1-3 for impact. |
| 50-70 | Fixes and tripwires | For the top five: one change before launch, one warning sign to watch, one named owner. |
| 70-75 | Decide | Go, go with changes, or delay. Written down with the date. |
The failure statement, which you read out word for word:
"It is [launch date + 6 months]. The [project name] has been switched off.
People are relieved it's gone. It cost us money and some goodwill.
Take ten minutes and write down every reason you can think of for how
we got here. Include the awkward ones: people, money, customers, the
supplier, the owner's decisions. Nothing you write will be held against you."
What comes out of the silent writing is rough, and that's fine. A handful of the plant shop's cards, as written:
"It told someone their order would arrive by Christmas Eve"
"Customers didn't realise it was a bot and got cross"
"Nobody read the transcripts after week two"
"It promised a refund for a plant that arrived fine"
"The packer was the only one who knew how to switch it off"
"It said 'water daily' for a succulent"
In clustering, the first and fourth became one cluster ("promises it can't keep"), the third and fifth another ("nobody owns it after launch"), and the last joined the web-search card from the brief. Six cards became four clusters. The raw wording matters: "nobody read the transcripts after week two" is far more useful than "insufficient monitoring", because it tells you exactly what the tripwire should measure.
The ranking is a quick sum. With three people and three votes each, the plant shop's nine votes fell: "promises it can't keep" 4, "nobody owns it after launch" 3, "advice that contradicts our guides" 2, "customers didn't realise it was a bot" 0. Likelihood times impact then separated the top two: promises scored 3 × 3 = 9, because a missed Christmas delivery costs a customer for good; ownership scored 3 × 2 = 6. The zero-vote cluster wasn't dropped, since the agent's opening line already says it's an AI; it was simply checked and ticked off.
Two rules make or break the round-robin. First, the most senior person speaks last on every round, so they don't set the tone. Second, no one defends the project during the session. If the owner starts explaining why a failure couldn't happen, the room goes quiet and you lose the next ten ideas.
Failure prompts for when the room goes quiet
Small teams often dry up after fifteen reasons. Keep these cards ready and read one out when the round-robin stalls. They are the failure types that come up again and again with AI tools.
- Wrong answers that sounded right. It quoted an old price, invented a policy, or gave a customer an answer that was plausible and false.
- It couldn't see what it needed. The diary, stock or customer record wasn't connected, or was out of date, so it made bookings or promises against stale data.
- Staff routed around it. People didn't trust it, checked everything twice, and the project saved nothing. Or the one person who understood it left.
- Customers hated it. Particular groups, such as older customers or people with an urgent problem, couldn't reach a human and complained or left.
- The bill. Usage-based charges, per minute or per message or per credit, ran well past the estimate.
- The supplier changed. Prices rose at renewal, a feature moved to a higher tier, the model changed, or the product closed. That last one is real: the Clockwise calendar assistant closed on 27 March 2026 at short notice after its team joined Salesforce, and users' data was deleted rather than transferred.
- Rules and data. Customer data went somewhere it shouldn't, calls were recorded without telling people, or nobody could explain a decision the tool made.
For the supplier cards in particular, the tutorial on keeping your data and prompts portable gives you fixes to write straight into the plan.
A dental practice before switching on AI call answering
Say a three-surgery dental practice with two receptionists takes around 140 calls a day and misses roughly a third of them at peak times. It plans to switch on an AI receptionist to answer overflow and out-of-hours calls, book check-ups and hygiene appointments, and take messages for everything else. Seven people attend the pre-mortem. They produce 34 reasons, which cluster into nine. The practice and its numbers are hypothetical; the top five after voting look like this.
| Failure story | L | I | Change before launch | Tripwire | Owner |
|---|---|---|---|---|---|
| It booked a routine check-up for a patient with swelling who needed to be seen that day | 2 | 3 | Any mention of pain, swelling, bleeding or injury ends booking and goes to an urgent message or transfer; tested with 20 scripted calls | Any urgent call mishandled in the weekly transcript review: urgent path switches to transfer-only | Practice manager and lead dentist |
| Double bookings, because it wrote into the wrong diary template | 2 | 2 | AI can only book into slots marked as bookable by it | More than two clashes in a week: booking off, message-taking only | Senior receptionist |
| Receptionists didn't trust it and rang every patient back | 3 | 2 | Receptionists write the answers to the 25 most common questions and test it first | Callback rate above 30% in week two: review the script with them | Senior receptionist |
| Patients who dislike automated calls complained and some moved practice | 2 | 2 | Saying "person" or pressing 0 reaches the message line at any point; announced on the website first | Three or more phone complaints in a month | Practice manager |
| The bill came in at three times the quote because calls ran long | 1 | 2 | Monthly spending cap and usage alert set in the supplier's portal | Any week above 120% of forecast usage | Owner |
L and I are likelihood and impact on a 1-3 scale. The decision was "go with changes": launch two weeks later than planned, after the 20-call emergency test. That two-week delay is what a good pre-mortem usually produces. Very few end in "cancel", and very few end in "no changes".
Notice where the top risk came from. The lead dentist raised the swelling story, and the receptionists hadn't considered it because they triage by instinct every day. It's the kind of risk covered in dental practice AI mistakes, and the fix is written into the tool's call script and escalation rules before launch, not after the first complaint.
Turning failure stories into tripwires
A fix you apply before launch deals with the risks you can prevent. A tripwire deals with the ones you can't: a measurable warning sign, chosen in advance, with an action already agreed. The point is to stop you debating while something is going wrong. When the number is hit, the owner acts; there is no meeting.
Write each one in this form:
If [measure] goes above [threshold] in [period],
[named owner] will [specific action] within [time].
Example: If the callback rate goes above 30% in any week,
the senior receptionist will pause booking and review the
script with the front-desk team within two working days.
Weak tripwires sound like "keep an eye on complaints" or "monitor usage". They have no number and no owner, so nothing ever triggers them. If you can't name the number, you probably can't measure it yet, and setting up that measurement becomes one of the fixes.
Here's the same conversion done on a cost risk. Weak version, from an illustrative gym's pre-mortem for HubSpot's Customer Agent: "keep an eye on the credits". HubSpot charges the agent at 50 credits per resolved conversation, and credits cost $10 per 1,000, so each resolved chat is about $0.50. The strong version:
If the Customer Agent resolves more than 300 conversations in a
calendar month (about $150 of credits at list price), the owner
will check the transcripts for repeat questions within two days
and add those answers to the website FAQ.
The action isn't "switch it off". High volume there may mean the agent is working; the tripwire exists to make someone look before the bill decides for them.
Using ChatGPT or Claude as an extra voice in the room
AI can add failure ideas your team missed, but only after the human round. If you ask it first, its list anchors everyone and the local, awkward risks never get said. Paste in the brief, with no customer or staff names, and ask:
You are helping run a project pre-mortem for a small business.
Here is the project brief: [paste brief].
Here are the failure reasons our team already listed: [paste list].
Imagine it is six months after launch and the project was abandoned.
List 10 further plausible reasons it failed that are NOT on our list.
For each, give: the reason in one sentence, an early warning sign we
could measure in the first 30 days, and one change we could make before
launch. Prioritise reasons specific to this kind of tool and business
over generic project-management risks.
For the dental practice, an illustrative extract of what comes back, and what to do with each line:
| AI's suggested reason | Keep or bin |
|---|---|
| "Insufficient staff training on the new system" | Bin: generic, and already covered by the receptionists' test |
| "Adult children calling on behalf of elderly parents couldn't get past the identity questions, so booked nothing" | Keep: specific, measurable (count abandoned calls that mention a relative) |
| "Holiday opening hours weren't updated in the AI's script, so it booked patients into days the practice was closed" | Keep: add "update closures in the script" to the practice manager's calendar |
| "Lack of stakeholder buy-in" | Bin |
| "Poor mobile signal garbled dates of birth, so records were matched to the wrong patient" | Keep: test with ten calls from a car or a weak-signal spot |
Expect about half the answers to be generic ("insufficient training", "unclear goals"). Keep the two or three that are specific and bin the rest. It doesn't know your staff, your customers or your supplier's contract, and those are where the biggest risks usually sit.
Signs the pre-mortem was theatre
- Nobody said anything uncomfortable. If every reason is technical, people didn't feel safe.
- Every fix is "train the staff". Training is sometimes right, but it's also what people write when they don't want to name a real cause.
- No tripwire has a number in it.
- Delay was never an option, because the launch date had been announced to customers.
- The supplier ran the session. They have every reason to steer it away from price rises, lock-in and their own product's weak spots.
- Nobody looks at it again. Book a 30-day check now, and move the top risks into your AI risk register so they stay visible after launch.
The 30-day check is short if the tripwires were written properly: go down the list and read off the numbers. At the dental practice it took twenty minutes. The urgent-call tripwire hadn't fired; the weekly reviews found two calls mentioning pain, and both had gone to the urgent message line as intended. Double bookings: one clash in week one, none since. The callback rate had fired, at 38% in week two, and the review found the cause wasn't distrust: the AI's booking confirmations didn't mention which surgery, so receptionists were ringing patients to tell them. One added line to the confirmation brought it down to 12% by week four. Usage sat at 90% of forecast. One tripwire firing in a month, with a clear cause and a quick fix, is a healthy result.
If the project stalls despite all this, the pattern is usually one of the ones in why AI pilots stall, and your pre-mortem notes will tell you which failure story came true. That's the second job a pre-mortem does: it gives you a record to check reality against.
Further reads
- What Are the Risks of Using AI in My Small Business? — A wider list of risks to compare against yours.
- AI Use Case Template: Score Every Idea on One Page — Score the idea before you plan its launch.
- How to Review an AI Tool After 90 Days: Keep, Fix or Cancel — Check the tripwires against real results later.
- AI Incident Response Plan for Small Businesses (With Template) — What to do if a predicted failure happens anyway.
- How to Measure Customer Reaction After Introducing AI — Turn customer-reaction tripwires into proper measures.
- AI Change Management for Small Teams: A Practical Plan — Handle the people risks the session will surface.
- How to Run Your First AI Pilot Project in a Small Business — Six stages for a first AI pilot, from a one-page charter to the keep, fix or stop meeting, followed through an optician's email pilot with real-looking numbers.
- How to Write a Business Plan With AI, and What to Check Yourself — Draft each section of a business plan with AI from your own facts, build the financials yourself, and check the claims lenders and partners will test.
- How to Use AI for Scenario Planning: Best, Worst and Likely Cases — Build best, worst and likely cases from your own numbers, use AI to challenge the assumptions, and turn each case into triggers and pre-agreed actions.
- Why AI Projects Fail in Small Businesses (It's Rarely the Tech) — The seven organisational reasons small-business AI projects fail, what each looks like by week three, and an eight-question check to run before you start.
- How to Test a New Service Idea With AI Before You Launch It — A six-week way to test a new service with AI doing the drafting and real customers supplying the evidence, with pass marks set before the results arrive.
- How to Map a Business Process Before You Automate It — A practical process-mapping method for automation: six columns per step, swimlanes drawn with AI help, an exception count and a label for every step.
- How to Rescue a Stalled AI Project or Exit It Cleanly — Freeze spending, take stock of what was built, then re-scope to one live piece in 30 days or shut it down without losing your data or logins.
- How to Choose Your First AI Project: 7 Tests Before You Commit — Put each candidate for your first AI project through seven pass-or-fail tests, then commit to the one that passes all seven, with a one-page note.
- AI Tools and AI Development: The Complete 2026 Guide — the AI hub, including every tutorial in the AI-for-business series.
Sources: Gary Klein, 'Performing a Project Premortem', Harvard Business Review, September 2007; Clockwise's shutdown notice after its team joined Salesforce (March 2026); Shopify help documentation on the Shopify Inbox AI agent; HubSpot credits pricing (checked September 2026).