How to Review an AI Tool After 90 Days: Keep, Fix or Cancel

Coding Liquids tutorial cover featuring Sagnik Bhattacharya for How to Review an AI Tool After 90 Days: Keep, Fix or Cancel.
Coding Liquids tutorial cover featuring Sagnik Bhattacharya for How to Review an AI Tool After 90 Days: Keep, Fix or Cancel.

Pull four pieces of evidence: who actually uses it (the admin usage report), what changed against your pre-AI baseline, how often its output needs fixing, and the full monthly cost including staff time. Score them, then decide: keep if it beats the baseline for its cost, fix if results are good where it's used, cancel if neither.

The verdict rarely has to be all or nothing. A common and sensible outcome is a split: keep the tool for the two people who get real value, drop the seats nobody opens, and give one struggling use a 30-day repair with a named change.

Follow me on Instagram@sagnikteaches

Put the review in the diary before the renewal date decides for you

The biggest cost in a 90-day review is doing it too late. Before anything else, find three dates and write them down: when the subscription renews, whether it's a monthly or annual term, and any notice period for cancelling or reducing seats.

Connect on LinkedInSagnik Bhattacharya

Terms vary more than people expect. Microsoft, for example, only gives a prorated refund or credit if you cancel a business subscription within seven days of it starting or renewing; after that, you turn off recurring billing so it ends at the end of the term (see Microsoft's page on cancelling a business subscription). Credit-based tools have their own trap: HubSpot's included credits, for instance, don't roll over, so an unused allowance is simply lost each month. If you're not sure what your contract says, the tutorial on auto-renewals, price rises and notice periods shows what to look for.

Subscribe on YouTube@codingliquids

Schedule the review for around day 80 to 85. That leaves time to gather evidence, run a short fix if you need one, and still act before the next billing term locks you in.

The evidence pack: what to pull, and where from

Most of this comes from systems you already have. Budget about three hours in total.

EvidenceWhere to get itTime
Who uses it, and how oftenMicrosoft 365 admin centre: Reports, then Usage, then the Microsoft 365 Copilot report (active users, prompts and last activity per person over a chosen period; names are concealed by default). Google Workspace Admin console: Generative AI, then Gemini reports. Other tools: the admin or billing page. If there's no report, ask each person.15 minutes
Results against your baselineRe-run the same measure you took before launch (time per task, turnaround, volume handled) for two normal weeks1 to 2 hours
Quality: the rework ratePick 20 recent outputs at random and count how many needed more than light editing45 minutes
Full costInvoices, plus add-ons, credits or automation tasks the tool triggers, plus staff time spent on setup, admin and checking30 minutes
What staff thinkA five-question pulse (below), answered privately10 minutes each
Anything that went wrongYour incident or error log, and a quick question at the team meeting10 minutes

Look at the last four weeks of usage, not the 90-day average. Enthusiasm in the first fortnight inflates the average; weeks nine to twelve show the habit that's actually formed.

"More than light editing" needs defining for tools that don't write drafts. For an AI meeting-notes tool, a useful rule is that a summary needs heavy rework if any action item is missing, assigned to the wrong person or given the wrong date, because those are the lines people act on. Check 20 summaries against the recordings or your own notes on that basis only. A summary with clumsy wording but correct actions counts as light; a polished one that drops "send the revised quote by Friday" counts as heavy.

Credit-based tools need the cost line worked out rather than read off the invoice. Take a business on HubSpot's Service Hub Professional using Customer Agent, which uses 50 credits per resolved conversation. A monthly pool of around 3,000 included credits covers about 60 resolved conversations. If the agent resolved 45 last month, 750 credits went unused and didn't roll over, so the real cost per resolution was higher than the headline $0.50. If it resolved 90, the extra 1,500 credits cost about $15 at $10 per 1,000. Either way, put the number of resolved conversations and the credits used in the pack, not just the plan price, and check how the tool defines "resolved" before crediting it with every closed chat.

If you never took a baseline, you can still rebuild a rough one: time five tasks done the old way this week, or pull the same figures from records dated before the tool arrived. It's weaker evidence, so say so in your notes.

Five-question staff pulse (answer honestly; names won't be shared)

1. In a normal week, how many times do you use [tool]?
   Never / 1-2 / 3-10 / More than 10
2. Which task do you use it for most?
3. For that task, does it save you time once you've checked
   the output? Saves a lot / Saves a little / No difference /
   Costs me time
4. What stops you using it more?
5. If we cancelled it tomorrow, what would you miss, if anything?

Question five is the most revealing. "Nothing" from someone with a paid seat is a clear answer.

Scoring sheet: five questions, ten points

A score doesn't make the decision for you, but it stops the review turning into a debate between the person who loves the tool and the person who pays for it.

Score each 0, 1 or 2.

1. ADOPTION: share of licensed people using it in at least
   3 of the last 4 weeks
   0 = under 30%   1 = 30-70%   2 = over 70%

2. RESULT AGAINST BASELINE
   0 = no measurable change
   1 = better, but short of the target you set
   2 = met or beat the target

3. QUALITY: share of sampled outputs needing heavy rework
   0 = more than 1 in 4   1 = 1 in 10 to 1 in 4   2 = under 1 in 10

4. VALUE AGAINST FULL COST (staff time saved, valued at
   loaded hourly cost, versus subscription + extras + admin)
   0 = worth less than it costs   1 = roughly break-even
   2 = clearly worth more

5. RISK AND FIT
   0 = a data incident or an unresolved privacy concern
   1 = minor issues, caught in time
   2 = no issues; fits how the team works

8-10  KEEP. Renew; consider removing unused seats.
5-7   FIX. One named change, 30 days, then re-score.
0-4   CANCEL. Follow the exit checklist.
Override: a 0 on question 5 means fix that first, or cancel,
whatever the total.

The override line matters more than it looks. Picture a small accountancy practice whose AI tool for summarising client correspondence scores 2, 2, 2 and 2 on the first four questions: well used, clearly faster, accurate, worth its cost. Then the evidence pack turns up an incident. During a busy week, a trainee whose seat hadn't been set up yet used the same tool's free personal tier and uploaded a client's full bank statements, outside the firm's account and its data terms. Risk and fit scores 0, so the total of 8 doesn't matter: the verdict is fix first. The fix names the change (every user on the firm's account before they touch client files, personal tiers blocked on work devices), the practice decides with its data-protection adviser whether the client needs to be told, and the tool is re-scored in 30 days.

Score each seat or use separately if the picture differs a lot between people. That's how you arrive at a split verdict rather than a blunt yes or no.

Worked example: four Copilot seats in a picture-framing business

Take an illustrative picture framer with a shop counter, a workshop and six staff. Ninety days ago the owner added Microsoft 365 Copilot Business for four people on monthly billing, at $25.20 a seat: $100.80 a month at list price in USD. The goals were faster quote emails, product descriptions for the online shop, and quicker checks of supplier price lists.

What the evidence showed:

  • Usage (last 28 days): two people active on 15 or more days; the counter assistant on 3 days; the workshop manager not since week two.
  • Result: the baseline, timed on 20 quotes before launch, was 12 minutes per quote email. The two regular users now average 7 minutes on the same kind of quote. They handle about 40 of the shop's 45 weekly quote emails, so the saving is roughly 40 × 5 minutes, about 3.3 hours a week or 14 hours a month.
  • Quality: 3 of 20 sampled quote drafts had a wrong moulding or glass price. The cause: an out-of-date supplier price list still sitting in a folder Copilot could read. All three were caught before sending, but that's a 15% heavy-rework rate.
  • Cost: $100.80 a month, plus about two hours a month of the owner's time on admin and spot checks.
  • Staff pulse: the workshop manager wrote "I'm cutting mouldings all day, I don't know what I'd use it for." The counter assistant didn't know Copilot could read the price list at all.

Scores: adoption 1 (2 of 4 people is 50%); result 2; quality 1; value 2 (14 hours at an illustrative loaded cost of $20 an hour is about $280 a month against roughly $140 including the owner's time); risk and fit 1. Total: 7, which means fix.

The decision was split. Keep the two regular users. Remove the workshop manager's seat at the next monthly renewal, taking the bill to $75.60. Fix the quality problem by deleting old price lists from the shared folder, keeping one current price file, and adding a line to the team's quote prompt: "List every price you used and the file it came from." Give the counter assistant a 30-minute session on two tasks. Re-score in 30 days, with targets of under 1 in 10 heavy rework and the counter assistant active on at least 10 days.

Not every review ends in a split. For contrast, take an illustrative events-catering company that bought a standalone AI writing tool for three people, at an illustrative $30 a seat, to speed up menu proposals. At day 85 the evidence pack read: one of three people active in the last four weeks (adoption 0); proposals still taking about 50 minutes each, the same as the baseline, because the tool couldn't see the company's costings or past menus (result 0); drafts mostly usable once rewritten (quality 1); $90 a month for no measured saving (value 0); no data issues (risk 2). Total: 3, cancel. The staff pulse settled any doubt. The one active user wrote that she mainly used it for social captions, which the Copilot Chat already included in the company's Microsoft 365 plan could do at no extra cost.

What "fix" should mean: one named change and a 30-day deadline

"Let's give it a bit longer" is how unused subscriptions survive for years. A fix is only a fix if it names the change, the person doing it and the number that will tell you it worked. The usual causes and their matching fixes:

  • People don't know what to use it for. Attach it to one recurring task per role and show them that task done well. The tutorial on getting staff to actually use AI tools covers this in depth.
  • Output is wrong because the sources are messy. AI that reads your files repeats whatever is in them, including last year's prices. Tidy what it can see; organising shared files so AI tools can use them walks through it.
  • Quality varies from person to person. Agree shared prompts for the main tasks so everyone starts from a version that works.
  • The tool is fine but the seats are wrong. Move licences to the people with the right tasks rather than buying more.
  • It's the wrong tool for the job. If the task needs something the tool can't do, no amount of training fixes that. Cancel and look again.

Written down, the picture framer's fix fitted on a card pinned by the counter:

FIX PLAN: Copilot, quote emails            Re-score: day 115
Change 1:  Delete old supplier price lists; keep one current
           file in "Reference - prices"      Who: owner, by Fri
Change 2:  Add "list every price you used and its file" to the
           shared quote prompt               Who: owner, same day
Change 3:  30-minute session for the counter assistant on quote
           emails and stock questions        Who: regular user
Measure:   Heavy rework under 1 in 10 on 20 sampled quotes;
           counter assistant active on 10+ days in 28

At day 115, the quality target was met comfortably: 1 heavy rework in 20 sampled quotes, and the price-and-file line made the remaining error easy to trace. The counter assistant managed 6 active days, short of 10. Rather than give it "a bit longer", the owner moved that seat to the online-shop assistant, who had been asking for one, and set the same test for the new user.

Hold one rule firmly: if the same tool fails the same test twice, it goes.

Cancelling without leaving loose ends

A cancelled AI tool can keep reading your data, running automations or charging you if it isn't unplugged properly. Work through this list:

  1. Export what you need: chat histories worth keeping, custom instructions, saved prompts and any generated files. Check how long the vendor keeps data after cancellation.
  2. Stop the billing at the right moment: reduce seats or turn off recurring billing before the renewal, following the terms you noted at the start.
  3. Disconnect it: remove its access to your email, calendar and file storage in your Google or Microsoft account settings, delete any API keys, and switch off automations in Zapier or Make that call it.
  4. Ask for deletion if it held customer or staff data, and keep the vendor's confirmation.
  5. Update your process notes and templates that mention the tool, so new staff don't go looking for it.
  6. Tell customers if they used it directly, for example if a chat window disappears from your website, and give them the replacement route.
  7. Record the decision: what you tried, the score and why you stopped. Next year, when someone suggests the same tool, you'll know what happened.

Step 3 is the one people skip, and the result tends to surface a week or two later. A realistic version: a small agency cancels an AI meeting note-taker and forgets its calendar connection. The bot keeps joining client calls under the old account for another fortnight, until a client asks who the silent attendee is. Or the tool is cancelled while an automation still sends each new enquiry to it for summarising; the step fails, the automation stops at that point, and enquiries stop reaching the inbox without anyone being told. Search your automation platform for the tool's name before you cancel, not after.

If you've built prompts, assistants or workflows around the tool, keep them in a form you can move elsewhere. The tutorial on keeping your data and prompts portable explains how to avoid starting from scratch.

Five biases that skew a 90-day verdict

  • Sunk cost. "We spent three months setting it up" isn't a reason to keep paying. Only the next year's costs and benefits count.
  • The novelty curve. Use usually peaks in the first weeks and settles. Judge the settled level.
  • The loudest user. One enthusiast's results aren't the team's. Score seats separately.
  • Logins mistaken for value. Someone who opens the tool daily to ask trivia saves nothing. Adoption only counts alongside a result.
  • A moving baseline. If the 90 days covered your busiest season, or your quietest, adjust before comparing. Note it in the review so next quarter's comparison is fair.

The moving baseline catches seasonal businesses most. Say an illustrative landscaping firm launched AI-drafted quotes in June and reviewed in September. Hours saved per month looked healthy. Had it reviewed in December, when quote requests drop to a fraction of the summer level, the same tool would have looked barely worth its seat. The per-quote figure (minutes saved on each quote) stays steady across the year, so review on that, and judge the annual value from a whole year's volume rather than from whichever quarter the review happens to fall in.

Once each tool has a verdict, put the numbers into a simple return calculation so next year's budget conversation starts from evidence. The worked method in how to calculate AI ROI uses the same inputs you've just gathered, and the next review, at day 180, should take half the time.

Further reads

Sources: Microsoft Learn pages on Microsoft 365 Copilot usage reports and on cancelling Microsoft business subscriptions; Google Workspace Admin Help on Gemini usage reports; Microsoft 365 Copilot Business list pricing (checked September 2026).

Coming up to a renewal on an AI tool?

On a 1:1 call we'll go through your usage and results, decide which seats or features earn their cost, and plan either the fix or a clean exit before the renewal date.

Book a 1:1 call with me