Keep one shared log where anyone records an AI mistake in under two minutes: what went wrong, where it was caught, how bad it was, and the likely cause. Review it monthly, group entries by cause, and close each one only when a prevention change exists: a prompt edit, a glossary entry, a checklist line or a tool setting.
Most logs fail by recording the mistake and stopping there. "AI got the date wrong, fixed it" tells you nothing next month. The column that makes a log work is "what we changed so it can't happen the same way again", and an entry without it stays open.
The error log's columns, and what each one is for
An error log template needs enough columns to find patterns and few enough that people fill it in. Twelve works. The first six are filled in by whoever spots the mistake; the last six by the log's owner at review.
| Column | Filled by | Why it's there |
|---|---|---|
| ID and date found | Finder | So entries can refer to each other ("repeat of E-034") |
| Job or client reference | Finder | Lets you check the original work later |
| Tool and task | Finder | Which AI, doing what: "ChatGPT, quote email" |
| What went wrong, quoted | Finder | The actual wrong text, not a summary of it |
| Where it was caught | Finder | Before or after it reached a client or customer, and by whom |
| Severity guess | Finder | 1, 2 or 3; the owner can change it |
| Cause category | Owner | One of eight categories, so patterns can be counted |
| Fix to this item | Owner | What was done about this particular error |
| Prevention change | Owner | What changed so it can't recur the same way |
| Prevention owner and date | Owner | Who makes the change, by when |
| Status | Owner | Open, prevented, or accepted (with a reason) |
| Repeat of | Owner | Links to an earlier entry; the most telling column of all |
The difference between a useless entry and a useful one is mostly the quote. Before: "Chatbot said something wrong about sworn translations. Fixed." After: "Website chatbot, 15 Sep, told a visitor 'we offer sworn translations in all languages'; we offer them in four. Visitor asked for Greek. Caught by the visitor's follow-up email." The second takes thirty seconds longer and tells the owner exactly what to look for in the knowledge base.
"Where it was caught" deserves a word. An error the reviser caught is a check working. An error a client caught is a check that failed. Both belong in the log, but they mean different things, and a log that only records client-found errors will always look worse than the work really is while telling you less about where to strengthen.
Eight entries from a translation agency's log
Here is a September extract from the log of an illustrative ten-person translation agency that uses AI for first-draft translations, email drafting, summaries of client briefs and a website chatbot:
ID Found By Tool / task What went wrong (quoted) Caught Sev Cause
E-031 3 Sep reviser MT + AI post-edit, "1,250.50" kept with a decimal point; before 2 instruction gap
financial report client's style needs "1.250,50"
E-032 4 Sep client MT, product manual client's product name translated as a after 2 terminology
common noun throughout
E-033 5 Sep reviser MT, safety leaflet "Do not immerse the unit in water" before 3 model omission
(Dutch) came out without "niet": the warning
said the opposite
E-034 8 Sep PM ChatGPT, quote email "delivery within 24 hours" for a before 2 missing source
9,000-word job (standard: 5 days)
E-035 10 Sep account Claude, brief summary summary omitted the deadline, which was after 2 model omission
manager in the email's last line
E-036 12 Sep reviser MT, customer letter formal and informal "you" mixed in one before 1 instruction gap
letter
E-037 15 Sep visitor website chatbot "we offer sworn translations in all after 2 missing source
email languages" (we offer 4)
E-038 18 Sep PM automation + AI step enquiry in Spanish acknowledged in after 1 tool setting
English
And the owner's columns for three of them, after the monthly review:
E-033 Fix: corrected by reviser before delivery.
Prevention: new check step for all safety content: a second AI
compares every negation, warning and number in source and target;
human revision stays mandatory. Owner: head of production, 20 Sep.
Status: prevented.
E-034 Fix: email corrected before sending.
Prevention: turnaround table added to the email project's files;
instruction added: "Never state a delivery time that isn't in the
turnaround table; write [CHECK TURNAROUND] instead."
Owner: PM lead, 12 Sep. Status: prevented.
E-037 Fix: visitor emailed with the correct list of languages.
Prevention: service list with language coverage added to the
chatbot's knowledge base; test question added to the monthly run.
Owner: marketing, 17 Sep. Status: prevented. Repeat of: E-019.
E-037's "repeat of E-019" is the line that matters. The chatbot had made a similar overclaim in August, and the fix then had been to correct that one answer. Only when the prevention went into the knowledge base did it stop.
Severity: three levels anyone can apply
Severity scales with more than three levels get argued over. Three levels, each with examples from your own work, get applied consistently:
- 1, minor. Wrong but harmless, or caught before anyone outside saw it and easy to fix: a register slip, a formatting inconsistency, an awkward phrase. E-036 and E-038.
- 2, significant. Would have misled a client or customer, cost money or time, or needed a correction after delivery: a wrong turnaround, a translated product name, an overclaim by the chatbot. E-031, E-032, E-034, E-035, E-037.
- 3, serious. A safety, legal, financial or data-protection consequence, or a client relationship at risk, whether or not it was caught in time. E-033 was caught, but a safety warning that says the opposite of the original is a level 3 near miss.
Level 3 does two things at once: it goes in the log, and it triggers your incident process if the error reached anyone. The log isn't the place to manage a live incident; an AI incident response plan covers containment and notification, and what to do when AI gets something wrong with a customer covers putting it right with the person affected. The log's job is to make sure the same thing doesn't happen again.
Cause categories that point straight to a fix
The cause column is what turns a list of mistakes into a list of fixes. Keep the categories few and practical, and define each by the kind of change that prevents it:
| Cause | Example | Prevention that usually works |
|---|---|---|
| Missing source | Quote email without the turnaround table | Add the source to the project or knowledge base; tell the AI to flag gaps |
| Stale source | An old rate card still in the project files | Remove superseded files; add "valid until" dates |
| Instruction gap | No rule on number formats or formal address | Add the rule to the prompt template or brief |
| Terminology | Product name translated or misspelt | Glossary or do-not-translate entry |
| Model invention or omission | Dropped negation; invented figure | A targeted check step; keep a human check for that content type |
| Tool or automation setting | Language detection not configured | Change the setting, add a filter step, test it |
| Process skipped | Revision bypassed on a rush job | Make the step impossible to skip in the workflow |
| Late human edit | Figure changed after revision, unchecked | Any change after sign-off goes back for a re-check |
Notice that only one category, model invention or omission, is really about the AI. Most entries in most logs turn out to be about what the AI was given or how the work flowed around it. That's good news, because those are the causes you control. For terminology entries, the fix is a proper glossary with a never-write column; building a product glossary AI must use shows how. For translation-specific entries such as E-031 and E-036, most prevention belongs in the mechanical checks described in how to check AI translations before customers see them.
Making logging take under two minutes
If logging takes longer than fixing the error, nobody logs. Three things keep it quick.
A form, not a spreadsheet. A short Google Forms or Microsoft Forms questionnaire feeding a shared sheet takes about 90 seconds to fill in and works from a phone. Six required fields: job reference, tool and task, what went wrong (paste the text), where it was caught, severity guess, your name. Everything else is the owner's job.
A blameless rule, written at the top of the form. "This log is about fixing processes, not about who made the mistake." If entries are used to criticise people, the log empties within a month, and the errors don't stop; they just stop being written down.
Let AI tidy the messy ones. People often report errors in a rushed chat message. The owner can turn those into clean entries with a prompt:
Turn this report into an error log entry with these fields:
tool/task, what went wrong (quote the wrong text if given), where
caught (before/after reaching a client), severity (1 minor,
2 significant, 3 serious) with a one-line reason, and the most likely
cause from this list: missing source, stale source, instruction gap,
terminology, model invention or omission, tool setting, process
skipped, late human edit.
If something needed isn't in the report, write "ask reporter".
Report: [paste message]
Given the chat message "chatbot told someone yesterday we do sworn translations in every language, they asked for Greek, we don't do Greek sworn", an illustrative result:
Tool/task: website chatbot, answering a service question
What went wrong: told a visitor "we offer sworn translations in all
languages"; the visitor requested a sworn translation into Greek,
which the agency doesn't offer
Caught: after reaching a customer (visitor enquiry)
Severity: 2, misled a prospective customer about a service
Likely cause: missing source (service coverage not in knowledge base)
Ask reporter: date of the conversation; was the visitor corrected?
The "ask reporter" line is the useful part: the AI identifies what's missing rather than filling it in. The owner still confirms the cause, because "missing source" is a guess until someone checks what the chatbot's knowledge base actually said.
Letting automations report their own failures
Not every AI error has a person nearby to notice it. Automations that call AI steps fail quietly, and the platforms' own alerts are easy to miss. A few behaviours are worth knowing. Zapier pauses a Zap automatically when 95% of its runs error over seven days, and sends no error emails when an error handler runs, so a handled failure can go unseen. Make's error handlers (now named Skip, Retry, Resume, Commit and Rollback) can quietly skip a bad record. A Power Automate flow that fails for 14 days straight gets switched off automatically.
Rather than relying on alerts, give each automation an error path that writes a row straight into the log: the automation's name, the step that failed, the input it was working on, and the error message. The agency added this to its enquiry-acknowledgement automation after E-038, so the next time language detection failed, the log recorded it the same minute instead of a client noticing a week later.
Chatbots can feed the log too. Most chatbot tools let you export or search conversations; a weekly search for phrases such as "that's not right", "you said" and "speak to a person" surfaces the conversations where something went wrong, and each confirmed error becomes an entry.
Closing the loop: from entry to prevention
An entry is closed when its prevention change exists, not when the individual error is fixed. Three prevention changes from the agency's September log, before and after:
- E-034, the quote email. Before: the drafting prompt said "Write a friendly quote email for this job." After: "Write a friendly quote email for this job. Take turnaround and price only from the attached tables. If the job isn't covered, write [CHECK TURNAROUND] or [CHECK PRICE] and don't estimate."
- E-035, the brief summary. Before: "Summarise this client brief." After: "Summarise this client brief. Then list separately, quoting the original words: every date or deadline, every number, every named deliverable, and any condition the client attached."
- E-032, the translated product name. Before: nothing in the brief about product names. After: a do-not-translate field in the brief template, loaded into the translation tool's glossary for that client before any file is translated.
Each change is small and specific, and each is testable: run the same job again and see whether the error recurs. Where a prevention change can't be made, for example because a tool has no setting for it, mark the entry "accepted" with the reason and the manual check that covers it. Accepted is a legitimate status. Open for three months is not.
Reading the log for patterns once a month
Individual entries fix individual errors. Patterns fix categories of error. Once a month, the owner sorts the log three ways and looks for the biggest pile:
- By cause. In September, instruction gaps and missing sources made up 11 of the agency's 23 entries. That pointed to the prompt templates and project files, not the AI models.
- By tool and task. Email drafting produced more significant errors than translation did, largely because email prompts had never been written down; each person used their own.
- By where caught. Six of 23 reached a client. All six came from tasks with no second check: emails, summaries and the chatbot. Translation errors were almost all caught by revisers.
That third view is the most useful for deciding where to add checking. The review itself can be a standing half-hour; how to run a monthly AI quality review in 30 minutes sets out an agenda that the log's patterns slot straight into.
Three months of the agency's log, in numbers
The illustrative agency's first quarter with the log went like this:
- September: 23 entries. 15 minor, 7 significant, 1 serious (the safety leaflet near miss). 6 reached a client. 9 were repeats of the same few problems: number formats, turnaround times and product names.
- October: 17 entries. 4 reached a client. 4 repeats, all from causes whose prevention wasn't finished yet.
- November: 9 entries. 1 reached a client. 1 repeat.
The falling count isn't only fewer errors; some of it is errors now caught by the new checks before they'd be worth logging. That's fine, and it's also why the first month's number shouldn't alarm anyone. A new log usually fills fast, because people finally have somewhere to put what they've been quietly fixing for months.
The time cost in September: 23 entries at about 2 minutes each (46 minutes across the team), a 30-minute monthly review, and 12 prevention changes averaging 15 minutes (3 hours). About 4 hours 15 minutes in all, or roughly $235 at an internal $55 an hour. By November, with fewer entries and fewer changes needed, it was under two hours. Set that against a single level 3 error reaching a client, such as a safety warning that says the opposite of the original, and the log is the cheapest part of the agency's quality process.
Why error logs get abandoned, and how to keep yours going
Logs rarely fail because the template was wrong. They fail for human reasons, each with a straightforward remedy:
- Nobody owns it. Entries pile up without causes or prevention. Name one owner, with the monthly review in their calendar.
- It feels like a blame list. People stop logging their own mistakes, and then other people's. Keep names in the "found by" column only, and praise finds, especially of your own errors.
- Too many fields. A form with fifteen required fields gets filled in twice. Six required fields for the finder, the rest for the owner.
- Entries never close. A log of 200 open entries is a museum. Close each one as prevented or accepted within a month, or escalate it.
- Nothing visibly changes. Share one line each month with the team: "Your entries led to these three changes." People log more when they can see it matters.
Kept up for a few months, the log becomes the most practical record you have of where AI helps your business and where it needs a person beside it, which is worth knowing every time you decide what to automate next.
Further reads
- AI Content Approval Workflow: Draft, Check, Sign Off — The review stages where most log entries get caught.
- How to Set Up Human Review for AI Work Without Slowing Down — Set review levels so the log's catches come early.
- How to Catch Outdated Information in AI Answers — Prevent the stale-fact entries that fill many logs.
- Chatbot Guardrails: Stop AI Promising What You Don't Offer — Fix chatbot promise errors at the source.
- How to Spot-Check AI Support Replies: A Weekly Sampling Routine — Find errors in support replies before customers do.
- Common AI Mistakes Small Businesses Make and How to Avoid Them — The mistakes worth watching for from day one.
- What an AI Consultant Can't Do for You, and What You Must Own — Seven responsibilities that stay with the business when you hire AI help, an ownership card for every automation, and promises no consultant should make.
- A Simple AI Risk Register for Small Businesses (With Template) — A one-table AI risk register with a scoring scale, a template to copy, twelve filled-in rows from a florist and the triggers for updating it.
- How to Handle Staff Who Over-Rely on AI — Signs of AI over-reliance, a conversation script, a three-rule standard to put in writing, and when a pattern needs a formal process.
- How to Choose and Support an AI Champion in a Small Team — A weighted scoring sheet for candidates, a one-page remit template, a monthly check-in agenda and a music school example with time costs.
- The Limits of AI: What It Still Gets Wrong in a Small Business — Eight things AI still gets wrong in a small firm, where each one bites, the warning signs, and the specific check that catches it.
- AI Hallucinations Explained for Business Owners: Causes and Fixes — Why AI invents statistics, quotes and policies, where that hurts a business most, and two copy-ready prompts plus a checking routine that catch it.
- A Wedding Planner's AI Workflow From Enquiry to Final Timeline — One illustrative planner's year with AI, stage by stage: consultation notes, proposal, suppliers, guests, confirmations and the wedding-day timeline.
- How to Stop AI Inventing Case Law: A Small-Firm Checking Routine — A seven-step routine and a copyable citation log for catching AI-invented or misquoted authorities before anything reaches a court or client.
- How Payroll Bureaus Use AI to Cut Errors and Queries — Four points in the pay cycle where AI cuts bureau errors and employee queries, while the calculations stay inside your payroll software.
- How Tutors Should Check AI-Made Worksheets Before Using Them — The ten-minute, four-pass check that catches wrong answer keys, two-answer questions, off-level reading and garbled diagrams before a pupil sees them.
- AI Tools and AI Development: The Complete 2026 Guide — the AI hub, including every tutorial in the AI-for-business series.
Sources: Zapier, Make and Power Automate help pages as summarised in our verified fact sheet (Zapier auto-pause and error-handler emails, Make error handler names, Power Automate flows switching off after continuous failure). The log format, severity levels and cause categories are practical recommendations.