A small law firm's first 90 days with AI should include, in order: a named partner owner and a time baseline in week one, a policy with approved tools and client wording by day 15, two low-risk pilots with checking routines from days 16 to 60, measurement, and a go/no-go decision with a budget by day 90.
The order matters more than the tools. Firms that buy a subscription first find staff using it for legal research before anyone has written down how research must be checked, and the first invented citation on a live file ends the experiment. Policy first, narrow pilots second, and measurement before any wider rollout keeps the risk small and gives the partners real evidence to decide on.
The plan on one page
The illustration throughout is a six-person private client and residential property firm: three lawyers (a partner in wills and probate, a partner in residential property, an associate doing both), a paralegal, a legal secretary and an office manager. The partner in wills and probate owns the programme.
| Days | Goal | Lead | Firm time | Done when |
|---|---|---|---|---|
| Before day 1 | Owner named, baseline logged | Owning partner | About 4 hours plus 2 weeks of light logging | Baseline sheet complete |
| 1–15 | Policy, approved tool, client wording | Owning partner, office manager | About 10 hours | Policy signed and briefed to all staff |
| 16–30 | Two pilots chosen and designed | Owning partner | About 6 hours | Checking routine written for each |
| 31–45 | Pilots in shadow mode | Pilot users | About 1 extra hour a week each | Error log shows the failure patterns |
| 46–60 | Pilots live with review | Pilot users | Review time only | Two weeks with no uncaught errors |
| 61–75 | Measurement | Office manager | About 5 hours | Scorecard compared with baseline |
| 76–90 | Go/no-go, budget, next pilot | Partners | About 4 hours | Decision memo signed |
Who does what in a firm of six
Small firms don't have an innovation team, so roles have to fit around fee-earning. In the illustration they were split like this:
- Owning partner: writes the policy, chooses the pilots, chairs a 20-minute check-in every Friday, pilots both workflows and signs the decision memo. Two hours a week on top of the pilot work.
- Second partner: pilots both workflows and acts as the sceptical reviewer, reading a sample of five items from colleagues' files each week without being told which were AI-assisted.
- Associate: pilots both workflows, and produces more attendance notes and letters than anyone else, so the associate's figures carry most weight at day 75.
- Office manager: sets up accounts, keeps the baseline and error log, and runs the day-75 measurement.
- Paralegal and secretary: not piloting in the first quarter, but briefed on the policy so nobody uses a personal AI account on firm work in the meantime.
The blind sample turned out to be the most useful habit in the programme. In week seven the second partner flagged an update letter as "a bit stiff", then learnt it was one the associate had written without AI. That settled a running argument about whether clients could tell the difference.
Before day 1: an owner and a baseline
Pick one partner to own the programme for the quarter, with two protected hours a week. An associate or office manager can do much of the work, but a partner has to own it: the policy needs partner authority, and the go/no-go is a partner decision.
Then log two ordinary weeks. Everyone notes, roughly, how long they spend on six recurring tasks. Don't aim for precision; aim for a number you'll be able to compare against later. Here's the firm's completed baseline (illustrative):
| Task | Times per month | Average minutes each | Hours per month |
|---|---|---|---|
| Attendance notes (calls and meetings) | 120 | 11 | 22 |
| Client update letters and emails | 50 | 18 | 15 |
| Reports on title, first drafts | 14 | 75 | 17.5 |
| Will drafts from instruction forms | 18 | 60 | 18 |
| Replying to routine enquiries | 90 | 6 | 9 |
| Chasing third parties (banks, agents, other side) | 70 | 7 | 8 |
The instruction to staff was one sentence: "For the next two weeks, whenever you finish one of these six tasks, add a line to the sheet with the task, the date and a rough number of minutes." A shared spreadsheet with four columns (name, task, date, minutes) is enough. Averages from two weeks are rough, but they're your own, and they're far more persuasive at day 90 than a vendor's claim.
About 90 hours a month across six people. Some of it will never be AI work, but the table gives the pilots something to be measured against.
Days 1–15: policy, tools and client wording
Choose one approved tool
For a firm this size, one general business AI plan is enough for the first quarter. Claude Team and ChatGPT Business both cost $25 per seat a month on monthly billing or $20 billed annually, with a two-seat minimum, and neither uses business content for training by default. If your practice-management system includes AI features, test those first; if they cover the pilots, you may not need anything else yet. Custom GPTs are being retired, so build any shared instructions in projects instead. Legal research tools can wait: neither pilot below needs one.
Write a one-page policy
Keep it short enough that people read it. The firm's version (illustrative) said:
1. Client information goes only into the approved tool, signed in with a firm account. No personal or free AI accounts for firm work, on any device.
2. AI may draft attendance notes, letters and summaries from material we provide. It may not be used to find or state the law during this pilot.
3. Every AI draft is reviewed by the fee earner responsible before it is saved to the file or sent. The reviewer, not the tool, is responsible for its content.
4. Errors found at review are recorded in the error log, whether or not they were caught.
5. Questions go to [owning partner]. This policy will be reviewed on day 90.
A fuller template, with sections for supervision and incidents, is in an AI acceptable use policy for a small professional firm.
Add a line to client terms
Something like: "We use AI tools within our secure systems to help prepare drafts. A qualified lawyer reviews all work, and your information is not used to train AI models." Check it against your regulator's guidance and your insurer's requirements before using it.
Brief everyone for an hour
Walk through the policy, show the approved tool on a dummy matter, and explain the error log. The firm's agenda was: ten minutes on why (the baseline numbers), fifteen on the policy line by line, twenty on a live demonstration drafting an attendance note from a made-up dictation, and fifteen on questions. Staff who understand why the rules exist follow them; staff who are handed a PDF don't.
Days 16–30: choosing two pilots
Score each candidate task from 1 to 5 on four things: volume (how often it happens), risk (reversed, so 5 means low risk), checkability (can a reviewer spot errors quickly?) and independence (does it work without integrations?). The firm's scoring (illustrative):
| Task | Volume | Low risk | Checkability | Independence | Total |
|---|---|---|---|---|---|
| Attendance notes from dictation | 5 | 4 | 5 | 5 | 19 |
| Client update letters | 4 | 4 | 5 | 5 | 18 |
| Routine enquiry replies | 5 | 3 | 4 | 3 | 15 |
| Report on title first drafts | 3 | 2 | 3 | 4 | 12 |
| Will drafts from instruction forms | 3 | 1 | 2 | 4 | 10 |
Will drafting scored lowest, which surprised the owning partner, who had assumed it was the obvious win. The problem is checkability: an error in a will can sit unnoticed for years, surface only after death, and be impossible to fix. Attendance notes and update letters, by contrast, are checked by the person who was on the call, minutes after they're drafted. Start where errors are cheap and visible. What to automate first in a small law firm goes through more candidates.
For each chosen pilot, write down: who takes part, the exact steps, the prompt, the checking routine, and what would stop the pilot. The update-letter card, filled in, was short enough to pin above a desk (illustrative):
PILOT: Client update letters (routine property and probate matters)
WHO: Both partners, associate. Secretary formats and sends.
STEPS: 1) Fee earner writes 3-6 bullet notes on the file.
2) Paste notes + matter type into the "Update letters" project.
3) Review the draft against the notes and the file.
4) Secretary sends from the case management system as usual.
CHECKS: Every date, figure and name against the file.
Nothing in the letter that isn't in the notes.
Mortgage/cash, sole/joint, executor names: matter-specific.
STOP IF: A letter reaches a client with a wrong date, figure or name.
For attendance notes, the stop rule was: "Any AI note saved to a file with an error that changes the meaning of advice or instructions pauses the pilot until the partners review it."
Days 31–45: pilots in shadow mode
Shadow mode means the AI drafts in parallel while people carry on as normal. Nothing AI-drafted goes on the file or to a client. The pilot users compare the AI's version with their own and log every difference that matters. It costs an extra hour or so a week per person, and it's the cheapest way to learn how the tool fails on your matters. The general method is set out in piloting AI in shadow mode.
For attendance notes, fee earners dictated a quick summary into their phone after each call, then pasted the transcript into a project set up with this instruction:
Turn the dictation below into an attendance note in our format:
Matter / Date / Duration / Type / Attendees / Summary / Advice given /
Client instructions / Actions (with owner and date).
Rules: record only what the dictation says. If something was being
considered, write "considering", not "instructed". If advice was
deferred ("we'll check and come back"), record it as an action, not as
advice. Flag any name, figure or date you are unsure of with [CHECK].
A dictation from the wills and probate partner:
attendance note, Mrs [client], estate of her late husband, phone call
twenty minutes. she's found another account, the building society one,
about eighteen thousand she thinks. asks whether she can sell the car
before probate, told her we'd check and come back. thinking about
whether the daughter should be an executor too, not decided. action
send her the building society letter template, diary two weeks
And the first AI draft (illustrative):
Summary: Client has located a further building society account,
balance approximately 18,000.
Advice given: Client should not sell the deceased's car until the
grant of probate is issued.
Client instructions: Client instructs us to add her daughter as an
executor.
Actions: Send building society letter template (fee earner, today).
Diary review in two weeks.
Two errors, both of the same kind. The partner had deferred the car question, and the draft turned it into advice she never gave. The daughter was a possibility, and the draft made it an instruction. It also missed an action: check the position on the car and reply. This pattern, AI hardening uncertainty into decisions, showed up in seven of the first 40 notes, even after the rules were added to the prompt. That's exactly what shadow mode is for. The routine that came out of it: every reviewer reads the "Advice given" and "Client instructions" sections word by word against their memory of the call, and treats every [CHECK] flag as a question to resolve.
If you'd rather record meetings than dictate afterwards, read whether lawyers can use AI note takers in client meetings first; consent and retention change the risk.
Days 46–60: pilots go live with review
Once shadow mode shows the error patterns and the checking routine catches them, the AI draft becomes the first draft. The reviewer edits it, saves it and records any errors in the log. The log from the first live week looked like this (illustrative):
| Date | Pilot | Error | Caught at review? | Cause |
|---|---|---|---|---|
| Day 47 | Update letter | Completion date given as 14th, not 15th | Yes | Two dates in the notes |
| Day 48 | Attendance note | "Instructed" for "considering" | Yes | Known pattern |
| Day 50 | Attendance note | Client's surname misspelt | Yes | Dictation mishearing |
| Day 52 | Update letter | Letter referred to "your mortgage offer", but the client is a cash buyer | Yes | Precedent assumed a mortgage |
All caught, and each one pointed to a fix: a second precedent for cash purchases, and a line at the top of every dictation giving the client's name spelt out. An error log that records causes improves the process; one that only records mistakes just makes people nervous.
Days 61–75: measuring against the baseline
The office manager repeated the two-week log for the two pilot tasks and filled in a scorecard (illustrative):
| Measure | Attendance notes | Update letters |
|---|---|---|
| Baseline minutes per item | 11 | 18 |
| Pilot minutes per item, including review | 5 | 8 |
| Items in the live phase | 86 | 40 |
| Errors caught at review | 9 | 4 |
| Errors that reached a file or client | 0 | 0 |
| Staff who want to keep it (of those piloting) | 3 of 3 | 2 of 3 |
At normal monthly volumes (120 notes, 50 letters), that's 6 minutes saved on each note and 10 on each letter: about 12 plus 8 hours, so roughly 20 hours a month. The one dissenting vote on letters came from the associate, who found editing AI drafts slower than writing from scratch for complex matters. The fair response was to keep AI letters optional for complex matters and default for routine ones. For a more careful approach to the numbers, see measuring time saved after rolling out AI in a small firm.
Days 76–90: the go/no-go and the next quarter
Decide against criteria written on day 16, not criteria invented after the results come in. The firm's criteria were: at least 25% time saved per item, no uncaught errors in the live phase, and a majority of pilot users wanting to continue. Both pilots passed. The decision memo, filled in, was one page:
DECISION: Go, both pilots, from day 91.
EVIDENCE: 20 hours/month saved at current volumes; 13 errors caught,
0 uncaught; 5 of 6 pilot votes to continue.
CONDITIONS: Review step stays mandatory. Error log continues.
Letters: AI draft default for routine matters, optional for complex.
BUDGET: 6 seats at $20/month (annual) = $120/month. No other spend.
NEXT PILOT (days 91-180): routine enquiry replies, shadow first.
NOT YET: will drafting; any use of AI to research or state the law.
OWNER: [partner] continues; office manager keeps the log.
REVIEW: Day 180.
The "not yet" line is as important as the "go". Writing down what the firm has decided not to do stops the programme drifting into the risky tasks by default.
What the 90 days cost this firm
Money was the small part. The firm paid for six seats on monthly billing during the pilot, so it could stop without an annual commitment, and switched to annual billing only after the go decision.
| Item | Cost over 90 days |
|---|---|
| 6 seats at $25 a month, monthly billing | $450 |
| Baseline logging, all staff | About 6 hours |
| Owner's planning, policy and pilot design | About 20 hours |
| Staff briefing | 6 hours (one each) |
| Shadow-mode extra time | About 6 hours |
| Measurement and decision | About 9 hours |
| Total firm time | About 47 hours |
Valued at an illustrative blended $150 an hour, that's about $7,050 of time plus $450 of software. Against roughly 20 hours a month saved, worth about $3,000 a month at the same rate, the quarter pays for itself within about three months of going live. The more important point is that the partners could see that number for themselves, from their own logs, before committing to anything annual.
What derails the 90 days, and the early signs
- The owning partner stops turning up. Early sign: the weekly check-in is moved twice in a row. Protect the two hours or hand ownership to someone who can.
- Scope creep during the pilot. Early sign: someone mentions they've "also been using it for" research or will clauses. Remind everyone of the policy and add the idea to the next-quarter list instead.
- An error log nobody fills in. Early sign: no entries in a week of live use. That almost never means no errors. Ask each reviewer for one example, face to face.
- Buying before the pilots finish. Early sign: a vendor demo booked for week five. Keep it, learn from it, but don't sign until day 90.
- Measuring only time. Early sign: the scorecard has no error column. Time saved with errors reaching files is a loss, not a gain.
If the programme does stall, why AI pilots stall, and how to get them live covers the usual causes and the ways back.
Questions partners raise about the plan
Can we compress the plan into 30 days?
You can compress the policy stage, but not the pilots. Two weeks of shadow running is the minimum to see how often the AI gets things wrong on your own matters, and a fortnight of live use to see whether the review step holds under normal workload. A firm that skips shadow testing tends to find its first serious error on a live file, which usually ends the experiment.
What if a partner refuses to use AI at all?
Don't make it compulsory in the first 90 days. Run the pilots with willing fee earners, share the error log and the time figures openly, and invite the sceptical partner to review a sample of AI-assisted work blind. Evidence from the firm's own matters is more persuasive than anything a vendor says, and a sceptic reviewing output often improves the checking routine.
Should we tell clients about the pilot?
Tell them in general terms through your engagement letter, and specifically whenever AI will process something unusual, such as a recording of their meeting. There's no need to announce each AI-assisted letter, because a lawyer reviews and signs every one. Check your regulator's guidance and your professional indemnity insurer's expectations before settling the wording.
Further reads
- ChatGPT or a Legal AI Tool: Which Should a Small Firm Use? — Choosing between a general AI plan and a legal tool.
- Per-Seat or Pay-As-You-Go? Legal AI Pricing for Small Firms — What the day-76 budget should allow for, and why.
- What an AI Consultant Does for a Law Firm, and What It Costs — If you want outside help running the 90 days.
- AI Time Capture for Solicitors: Recover Hours You Never Billed — A strong candidate for your next quarter's pilot.
- How to Stop AI Inventing Case Law: A Small-Firm Checking Routine — The checking routine to add before any research use.
- Can Solicitors Use ChatGPT Without Breaching Confidentiality? — The confidentiality questions behind the day-15 policy.
- How Small Law Firms Use AI to Draft Letters and Routine Documents — A precedent-first way to draft routine legal letters with AI: risk bands, a cleaned precedent pack, a fill-and-flag prompt and a fee-earner checklist.
- AI Client Intake for Law Firms: Qualify Enquiries Out of Hours — Set up overnight AI intake that gathers facts, screens fit and urgency, books consultations and hands a clean summary to a person each morning.
- How Conveyancers Use AI to Cut Admin on Each Transaction — Stage-by-stage admin savings for a conveyancing file, a title-summary prompt with sample output, chaser wording and the tasks AI must never touch.
- Which Legal Tasks Should a Small Firm Never Hand to AI? — A four-question test for legal AI risks, the seven jobs that stay with a named lawyer, and wording to put the line in your firm's AI policy.
- AI Tools and AI Development: The Complete 2026 Guide — the AI hub, including every tutorial in the AI-for-business series.
Sources: Anthropic Claude Team and OpenAI ChatGPT Business pricing pages; OpenAI help pages on the retirement of custom GPTs.