AI Error Log: Track Mistakes and Stop Them Happening Again

Coding Liquids tutorial cover featuring Sagnik Bhattacharya for AI Error Log: Track Mistakes and Stop Them Happening Again.
Coding Liquids tutorial cover featuring Sagnik Bhattacharya for AI Error Log: Track Mistakes and Stop Them Happening Again.

Keep one shared log where anyone records an AI mistake in under two minutes: what went wrong, where it was caught, how bad it was, and the likely cause. Review it monthly, group entries by cause, and close each one only when a prevention change exists: a prompt edit, a glossary entry, a checklist line or a tool setting.

Most logs fail by recording the mistake and stopping there. "AI got the date wrong, fixed it" tells you nothing next month. The column that makes a log work is "what we changed so it can't happen the same way again", and an entry without it stays open.

Follow me on Instagram@sagnikteaches

The error log's columns, and what each one is for

An error log template needs enough columns to find patterns and few enough that people fill it in. Twelve works. The first six are filled in by whoever spots the mistake; the last six by the log's owner at review.

Connect on LinkedInSagnik Bhattacharya
ColumnFilled byWhy it's there
ID and date foundFinderSo entries can refer to each other ("repeat of E-034")
Job or client referenceFinderLets you check the original work later
Tool and taskFinderWhich AI, doing what: "ChatGPT, quote email"
What went wrong, quotedFinderThe actual wrong text, not a summary of it
Where it was caughtFinderBefore or after it reached a client or customer, and by whom
Severity guessFinder1, 2 or 3; the owner can change it
Cause categoryOwnerOne of eight categories, so patterns can be counted
Fix to this itemOwnerWhat was done about this particular error
Prevention changeOwnerWhat changed so it can't recur the same way
Prevention owner and dateOwnerWho makes the change, by when
StatusOwnerOpen, prevented, or accepted (with a reason)
Repeat ofOwnerLinks to an earlier entry; the most telling column of all

The difference between a useless entry and a useful one is mostly the quote. Before: "Chatbot said something wrong about sworn translations. Fixed." After: "Website chatbot, 15 Sep, told a visitor 'we offer sworn translations in all languages'; we offer them in four. Visitor asked for Greek. Caught by the visitor's follow-up email." The second takes thirty seconds longer and tells the owner exactly what to look for in the knowledge base.

Subscribe on YouTube@codingliquids

"Where it was caught" deserves a word. An error the reviser caught is a check working. An error a client caught is a check that failed. Both belong in the log, but they mean different things, and a log that only records client-found errors will always look worse than the work really is while telling you less about where to strengthen.

Eight entries from a translation agency's log

Here is a September extract from the log of an illustrative ten-person translation agency that uses AI for first-draft translations, email drafting, summaries of client briefs and a website chatbot:

ID     Found  By        Tool / task             What went wrong (quoted)                 Caught   Sev  Cause
E-031  3 Sep  reviser   MT + AI post-edit,      "1,250.50" kept with a decimal point;     before   2    instruction gap
                        financial report        client's style needs "1.250,50"
E-032  4 Sep  client    MT, product manual      client's product name translated as a     after    2    terminology
                                                common noun throughout
E-033  5 Sep  reviser   MT, safety leaflet      "Do not immerse the unit in water"        before   3    model omission
                        (Dutch)                 came out without "niet": the warning
                                                said the opposite
E-034  8 Sep  PM        ChatGPT, quote email    "delivery within 24 hours" for a          before   2    missing source
                                                9,000-word job (standard: 5 days)
E-035  10 Sep account   Claude, brief summary   summary omitted the deadline, which was   after    2    model omission
              manager                           in the email's last line
E-036  12 Sep reviser   MT, customer letter     formal and informal "you" mixed in one    before   1    instruction gap
                                                letter
E-037  15 Sep visitor   website chatbot         "we offer sworn translations in all       after    2    missing source
              email                             languages" (we offer 4)
E-038  18 Sep PM        automation + AI step    enquiry in Spanish acknowledged in        after    1    tool setting
                                                English

And the owner's columns for three of them, after the monthly review:

E-033  Fix: corrected by reviser before delivery.
       Prevention: new check step for all safety content: a second AI
       compares every negation, warning and number in source and target;
       human revision stays mandatory. Owner: head of production, 20 Sep.
       Status: prevented.
E-034  Fix: email corrected before sending.
       Prevention: turnaround table added to the email project's files;
       instruction added: "Never state a delivery time that isn't in the
       turnaround table; write [CHECK TURNAROUND] instead."
       Owner: PM lead, 12 Sep. Status: prevented.
E-037  Fix: visitor emailed with the correct list of languages.
       Prevention: service list with language coverage added to the
       chatbot's knowledge base; test question added to the monthly run.
       Owner: marketing, 17 Sep. Status: prevented. Repeat of: E-019.

E-037's "repeat of E-019" is the line that matters. The chatbot had made a similar overclaim in August, and the fix then had been to correct that one answer. Only when the prevention went into the knowledge base did it stop.

Severity: three levels anyone can apply

Severity scales with more than three levels get argued over. Three levels, each with examples from your own work, get applied consistently:

  • 1, minor. Wrong but harmless, or caught before anyone outside saw it and easy to fix: a register slip, a formatting inconsistency, an awkward phrase. E-036 and E-038.
  • 2, significant. Would have misled a client or customer, cost money or time, or needed a correction after delivery: a wrong turnaround, a translated product name, an overclaim by the chatbot. E-031, E-032, E-034, E-035, E-037.
  • 3, serious. A safety, legal, financial or data-protection consequence, or a client relationship at risk, whether or not it was caught in time. E-033 was caught, but a safety warning that says the opposite of the original is a level 3 near miss.

Level 3 does two things at once: it goes in the log, and it triggers your incident process if the error reached anyone. The log isn't the place to manage a live incident; an AI incident response plan covers containment and notification, and what to do when AI gets something wrong with a customer covers putting it right with the person affected. The log's job is to make sure the same thing doesn't happen again.

Cause categories that point straight to a fix

The cause column is what turns a list of mistakes into a list of fixes. Keep the categories few and practical, and define each by the kind of change that prevents it:

CauseExamplePrevention that usually works
Missing sourceQuote email without the turnaround tableAdd the source to the project or knowledge base; tell the AI to flag gaps
Stale sourceAn old rate card still in the project filesRemove superseded files; add "valid until" dates
Instruction gapNo rule on number formats or formal addressAdd the rule to the prompt template or brief
TerminologyProduct name translated or misspeltGlossary or do-not-translate entry
Model invention or omissionDropped negation; invented figureA targeted check step; keep a human check for that content type
Tool or automation settingLanguage detection not configuredChange the setting, add a filter step, test it
Process skippedRevision bypassed on a rush jobMake the step impossible to skip in the workflow
Late human editFigure changed after revision, uncheckedAny change after sign-off goes back for a re-check

Notice that only one category, model invention or omission, is really about the AI. Most entries in most logs turn out to be about what the AI was given or how the work flowed around it. That's good news, because those are the causes you control. For terminology entries, the fix is a proper glossary with a never-write column; building a product glossary AI must use shows how. For translation-specific entries such as E-031 and E-036, most prevention belongs in the mechanical checks described in how to check AI translations before customers see them.

Making logging take under two minutes

If logging takes longer than fixing the error, nobody logs. Three things keep it quick.

A form, not a spreadsheet. A short Google Forms or Microsoft Forms questionnaire feeding a shared sheet takes about 90 seconds to fill in and works from a phone. Six required fields: job reference, tool and task, what went wrong (paste the text), where it was caught, severity guess, your name. Everything else is the owner's job.

A blameless rule, written at the top of the form. "This log is about fixing processes, not about who made the mistake." If entries are used to criticise people, the log empties within a month, and the errors don't stop; they just stop being written down.

Let AI tidy the messy ones. People often report errors in a rushed chat message. The owner can turn those into clean entries with a prompt:

Turn this report into an error log entry with these fields:
tool/task, what went wrong (quote the wrong text if given), where
caught (before/after reaching a client), severity (1 minor,
2 significant, 3 serious) with a one-line reason, and the most likely
cause from this list: missing source, stale source, instruction gap,
terminology, model invention or omission, tool setting, process
skipped, late human edit.
If something needed isn't in the report, write "ask reporter".
Report: [paste message]

Given the chat message "chatbot told someone yesterday we do sworn translations in every language, they asked for Greek, we don't do Greek sworn", an illustrative result:

Tool/task: website chatbot, answering a service question
What went wrong: told a visitor "we offer sworn translations in all
  languages"; the visitor requested a sworn translation into Greek,
  which the agency doesn't offer
Caught: after reaching a customer (visitor enquiry)
Severity: 2, misled a prospective customer about a service
Likely cause: missing source (service coverage not in knowledge base)
Ask reporter: date of the conversation; was the visitor corrected?

The "ask reporter" line is the useful part: the AI identifies what's missing rather than filling it in. The owner still confirms the cause, because "missing source" is a guess until someone checks what the chatbot's knowledge base actually said.

Letting automations report their own failures

Not every AI error has a person nearby to notice it. Automations that call AI steps fail quietly, and the platforms' own alerts are easy to miss. A few behaviours are worth knowing. Zapier pauses a Zap automatically when 95% of its runs error over seven days, and sends no error emails when an error handler runs, so a handled failure can go unseen. Make's error handlers (now named Skip, Retry, Resume, Commit and Rollback) can quietly skip a bad record. A Power Automate flow that fails for 14 days straight gets switched off automatically.

Rather than relying on alerts, give each automation an error path that writes a row straight into the log: the automation's name, the step that failed, the input it was working on, and the error message. The agency added this to its enquiry-acknowledgement automation after E-038, so the next time language detection failed, the log recorded it the same minute instead of a client noticing a week later.

Chatbots can feed the log too. Most chatbot tools let you export or search conversations; a weekly search for phrases such as "that's not right", "you said" and "speak to a person" surfaces the conversations where something went wrong, and each confirmed error becomes an entry.

Closing the loop: from entry to prevention

An entry is closed when its prevention change exists, not when the individual error is fixed. Three prevention changes from the agency's September log, before and after:

  • E-034, the quote email. Before: the drafting prompt said "Write a friendly quote email for this job." After: "Write a friendly quote email for this job. Take turnaround and price only from the attached tables. If the job isn't covered, write [CHECK TURNAROUND] or [CHECK PRICE] and don't estimate."
  • E-035, the brief summary. Before: "Summarise this client brief." After: "Summarise this client brief. Then list separately, quoting the original words: every date or deadline, every number, every named deliverable, and any condition the client attached."
  • E-032, the translated product name. Before: nothing in the brief about product names. After: a do-not-translate field in the brief template, loaded into the translation tool's glossary for that client before any file is translated.

Each change is small and specific, and each is testable: run the same job again and see whether the error recurs. Where a prevention change can't be made, for example because a tool has no setting for it, mark the entry "accepted" with the reason and the manual check that covers it. Accepted is a legitimate status. Open for three months is not.

Reading the log for patterns once a month

Individual entries fix individual errors. Patterns fix categories of error. Once a month, the owner sorts the log three ways and looks for the biggest pile:

  1. By cause. In September, instruction gaps and missing sources made up 11 of the agency's 23 entries. That pointed to the prompt templates and project files, not the AI models.
  2. By tool and task. Email drafting produced more significant errors than translation did, largely because email prompts had never been written down; each person used their own.
  3. By where caught. Six of 23 reached a client. All six came from tasks with no second check: emails, summaries and the chatbot. Translation errors were almost all caught by revisers.

That third view is the most useful for deciding where to add checking. The review itself can be a standing half-hour; how to run a monthly AI quality review in 30 minutes sets out an agenda that the log's patterns slot straight into.

Three months of the agency's log, in numbers

The illustrative agency's first quarter with the log went like this:

  • September: 23 entries. 15 minor, 7 significant, 1 serious (the safety leaflet near miss). 6 reached a client. 9 were repeats of the same few problems: number formats, turnaround times and product names.
  • October: 17 entries. 4 reached a client. 4 repeats, all from causes whose prevention wasn't finished yet.
  • November: 9 entries. 1 reached a client. 1 repeat.

The falling count isn't only fewer errors; some of it is errors now caught by the new checks before they'd be worth logging. That's fine, and it's also why the first month's number shouldn't alarm anyone. A new log usually fills fast, because people finally have somewhere to put what they've been quietly fixing for months.

The time cost in September: 23 entries at about 2 minutes each (46 minutes across the team), a 30-minute monthly review, and 12 prevention changes averaging 15 minutes (3 hours). About 4 hours 15 minutes in all, or roughly $235 at an internal $55 an hour. By November, with fewer entries and fewer changes needed, it was under two hours. Set that against a single level 3 error reaching a client, such as a safety warning that says the opposite of the original, and the log is the cheapest part of the agency's quality process.

Why error logs get abandoned, and how to keep yours going

Logs rarely fail because the template was wrong. They fail for human reasons, each with a straightforward remedy:

  • Nobody owns it. Entries pile up without causes or prevention. Name one owner, with the monthly review in their calendar.
  • It feels like a blame list. People stop logging their own mistakes, and then other people's. Keep names in the "found by" column only, and praise finds, especially of your own errors.
  • Too many fields. A form with fifteen required fields gets filled in twice. Six required fields for the finder, the rest for the owner.
  • Entries never close. A log of 200 open entries is a museum. Close each one as prevented or accepted within a month, or escalate it.
  • Nothing visibly changes. Share one line each month with the team: "Your entries led to these three changes." People log more when they can see it matters.

Kept up for a few months, the log becomes the most practical record you have of where AI helps your business and where it needs a person beside it, which is worth knowing every time you decide what to automate next.

Further reads

Sources: Zapier, Make and Power Automate help pages as summarised in our verified fact sheet (Zapier auto-pause and error-handler emails, Make error handler names, Power Automate flows switching off after continuous failure). The log format, severity levels and cause categories are practical recommendations.

Want an error log that actually changes things?

On a 1:1 call we'll look at where AI mistakes turn up in your work, set up a log your team will fill in, and agree how each kind of error gets fixed at the source.

Book a 1:1 call with me