Assume every Zap and scenario will eventually fail without an error. Switch on error notifications that reach two people, attach error handlers that alert a named person, add a daily heartbeat or volume check that fires when expected runs don't happen, and spend 15 minutes a week reading run history. Together those catch almost every quiet failure.
Error emails only cover errors. The failures that hurt are the quiet ones: a Zap that was switched off, a filter that now blocks everything because someone renamed a form field, an AI step that returns an empty value which the next step happily saves. None of those raises an alarm. You need checks that look for missing results, not only failed ones.
Six ways automations fail without telling anyone
| Failure | Illustrative example | Why no alert arrives | How to catch it |
|---|---|---|---|
| The trigger stops firing | The website form tool is updated and its connection to Zapier breaks | Nothing runs, so nothing errors | Volume or heartbeat check |
| A filter blocks everything | A form field is renamed from "Job type" to "Service"; the filter condition never matches | Filtered runs aren't errors | Daily count of runs that got through |
| The automation is switched off | Auto-paused after repeated errors while the owner was on holiday | The alert went to one inbox; after that nothing runs | Heartbeat check, shared alert address |
| Success with bad data | An AI step returns "I can't access that file", which is saved as a CRM note | Every step technically succeeded | Validate outputs before saving |
| An error handler swallows the error | A Skip handler drops failed bundles and marks the run a success | Handled errors don't send the usual emails | Handlers that notify a person |
| The allowance runs out | Monthly tasks or credits used up on the 24th | Runs stop or are held, often with one email to the account owner | Usage alert, weekly usage check |
Most small businesses discover these the expensive way. An illustrative sign maker found two in one quarter: quote requests from its website stopped reaching the job board for 11 days after a form field was renamed, and a Make scenario that booked installation slots was switched off after a Google connection expired while the owner was away. The rest of this tutorial is what it put in place afterwards.
Layer 1: make sure the right people hear about errors
In Zapier, go to Settings, then Notifications, then My notification settings, and edit the default. The options are Immediately (an email per error, the default), Immediately, then hourly summary, Hourly summary, and Never. Choose one of the first two for anything important. On Pro plans and above you can set custom rules for individual Zaps or folders, for example immediate alerts for the quote Zap and an hourly summary for the rest. On Team and Enterprise plans, owners and super admins can configure alerts for any Zap, and you can add labels that appear in the email subject line to make triage quicker.
In Make, the platform emails you when an error isn't handled by any error handler, and again when a scenario is disabled because of repeated errors. Check your profile's notification settings to confirm those emails are on and going to an address someone reads.
Then fix the address problem, which causes more silent failures than any technical issue. At the sign maker, error emails went to the former office manager's mailbox, which was being forwarded to nobody. Create a shared address or channel for automation alerts, make sure two people see it, and use it for every tool's notifications. If you're not sure which automations exist or who owns them, start with an automation audit to find the Zaps and scenarios nobody owns.
Layer 2: error handlers that tell a person
Notifications tell you an error happened. Error handlers decide what the automation does next, and the important design rule is that every handler must alert someone.
Zapier's error handler is available from Professional. Click the three dots on any action step (not the trigger or a Paths step itself, though you can add one to an action inside a path) and add an error handler. It runs as an alternative route when that step fails, and the error message is available as a field. A useful minimum is one step that posts to your alerts channel: "Quote Zap failed at Create job card: [error message]. Customer email: [email]." You can add Paths inside the handler to treat different errors differently.
One detail matters a great deal: Zapier doesn't send its usual error notification emails when an error handler runs. If your handler quietly logs the problem to a spreadsheet nobody reads, you've turned a noisy failure into a silent one. The handler must notify.
Zapier's Autoreplay setting, on Professional and above, retries failed steps automatically: up to five times, with the last attempt about 10 hours and 35 minutes after the first error. It's worth enabling for temporary problems, but note the delay: an error that autoreplay eventually can't fix reaches you half a day later.
Make's error handlers attach to individual modules, and there are five: Skip, Retry, Resume, Commit and Rollback. (Older guides call two of them Ignore and Break.) Activating a handler doesn't consume operations. How each behaves, and the silent-failure risk in each:
| Handler | What it does | Run status | Silent-failure risk |
|---|---|---|---|
| Skip | Drops the failed bundle, carries on with the rest | Success | High: data is lost and the run looks fine |
| Retry | Stores the failed bundle as an incomplete execution and retries per your settings | Warning | Low, if someone checks incomplete executions |
| Resume | Replaces the failed output with a substitute value you define | Success | High, unless the substitute is obviously a placeholder |
| Commit | Stops the run, keeps changes made so far in database apps | Warning | Medium |
| Rollback | Stops the run, reverts changes where the app supports it; the default when you set nothing and incomplete executions are off | Error | Low: it errors loudly, and repeated errors disable the scenario |
Two practical rules for Make. First, an error-handling route doesn't need a handler at the end: you can put a single Slack or email module on it to announce the problem. Second, if you use Resume, make the substitute impossible to miss, such as "CHECK-MISSING" instead of a plausible fake value. At the sign maker, a Resume handler that filled missing phone numbers with "000" had quietly created 40 contacts nobody could call.
Also switch on Store incomplete executions in each important scenario's settings. It's off by default. With it on, a failed run is kept so you can fix the cause and resume from where it stopped, instead of losing the data.
Layer 3: check for runs that didn't happen
This is the layer most businesses lack, and the one that would have caught both of the sign maker's failures. There are two simple forms.
A heartbeat check. A monitoring service gives you a unique web address. Your automation visits it at the end of every successful run, and the service alerts you if the visit doesn't happen within the expected window. Healthchecks.io, for example, has a free Hobbyist plan that monitors 20 jobs. In Make, add an HTTP module at the end of the scenario; in Zapier, use a webhook step (Webhooks by Zapier isn't available on the free plan). For a scenario that runs every hour, set the check to alert if there's been no ping for two hours.
A volume check. Heartbeats suit scheduled automations. For event-driven ones, like "new quote request", check the output instead. A daily scheduled Zap or scenario counts yesterday's new rows in the job board sheet and posts an alert if the number is below a threshold. The sign maker averages about 14 quote requests per weekday, so its rule is: if fewer than 4 have arrived by 15:00 on a weekday, post to the alerts channel. It fired twice in the next quarter; once was a genuinely quiet public-holiday week, and once was a broken form connection, found on the same afternoon rather than 11 days later.
Zapier Manager also has a New Zap Error trigger that fires when one of your Zaps errors, which you can route to Slack or SMS. It's useful, but it depends on an error happening, so it doesn't replace a volume check. If you've enabled Autoreplay, expect it to fire later, after the replays.
Layer 4: validate AI output before it's saved or sent
AI steps add a new kind of quiet failure: the step succeeds, but what it returned is useless. The model might return an empty field, a refusal ("I'm sorry, I can't open attachments"), a label that isn't on your list, or a summary of the wrong email. Add a check straight after every AI step:
- Not empty: a Filter (Zapier) or filter on the route (Make) that only continues if the key output field exists.
- Allowed values: for classifications, continue only if the category is one of your list; anything else goes to a Needs review route.
- Sensible length: a 10-word "summary" of a 2,000-word thread, or a 900-word "subject line", means something went wrong.
- No refusal phrases: route outputs containing "I can't", "I'm unable" or "as an AI" to review.
The sign maker found 37 CRM notes reading "I'm unable to access the attached PDF" before it added the last rule; the AI step had been asked to summarise attachments it couldn't read. Setting up the AI steps themselves, including output fields that make validation easy, is covered in adding AI steps to Zapier and, for Make, building your first AI automation in Make.
The automatic switch-off rules to know
Both platforms turn automations off by themselves under certain conditions. After that, nothing runs, so nothing else errors, which is the quietest failure of all.
Zapier pauses a Zap when 95% or more of its runs have errored in the last 7 days. Team accounts get a 24-hour grace period before the pause and Enterprise accounts 72 hours; Company-plan customers can override the behaviour per Zap. An owner email is sent, but if that owner has left, nobody sees it.
Make disables a scenario's scheduling after a set number of consecutive failed runs; the Number of consecutive errors setting, in advanced settings, defaults to 3. Some cases skip the count entirely: scenarios with instant triggers (webhooks) are disabled immediately when an error happens, and account validation, operations-limit-exceeded and data-size-limit errors disable scheduling straight away. If incomplete executions storage fills up and the Enable data loss setting is off, Make disables the scenario too. Warnings, including errors handled by Retry, don't count towards consecutive errors.
The practical defence is the heartbeat check from Layer 3, which notices a switched-off automation within one missed cycle, whatever the reason.
A 15-minute weekly review
Put a recurring 15 minutes in one person's calendar, ideally Monday morning, and work through the same list:
- Zap history: filter the past week for runs with problem statuses, including Errored, Safely halted, On hold, Handled error and Scheduled (runs waiting for Autoreplay). Open anything unfamiliar.
- Make scenario history and the Incomplete executions tab: resolve or delete every stored incomplete run.
- Last successful run for each critical automation. If a daily one last succeeded four days ago, investigate.
- Usage: tasks and credits used against last week and the monthly allowance.
- Heartbeat dashboard: any checks that went red and recovered.
- Alerts address: still reaching two people?
Keep a short log. The sign maker's looks like this (illustrative):
| Week | Found | Action | Minutes |
|---|---|---|---|
| 1 | 3 incomplete executions: supplier price file had a new column | Remapped, resumed all three | 20 |
| 2 | Nothing | None | 10 |
| 3 | Task usage 40% above normal: a loop in the invoice Zap | Fixed filter; usage back to normal | 25 |
| 4 | Heartbeat red for 2 hours on Tuesday: Google outage | None needed; runs caught up | 10 |
Over time the log becomes a record of what breaks and why, which is the same idea as an AI error log that stops mistakes recurring, applied to automations.
The sign maker's setup, start to finish
Illustratively, the sign maker had nine automations across Zapier and Make: quote requests, job cards, installation bookings, supplier invoices, review requests, a weekly summary, and three smaller ones. Here's what protecting them took:
- Ranking: 30 minutes listing all nine by the damage a silent week would cause. Quote requests, job cards and installation bookings came top.
- Notifications: 20 minutes creating a shared alerts address and pointing both tools at it, with immediate Zapier alerts for the top three.
- Handlers: about 90 minutes adding Zapier error handlers and Make Retry handlers with Slack alerts to the top three, and replacing the "000" Resume value.
- Heartbeat and volume checks: about an hour for three heartbeat checks on the free Healthchecks.io plan and one volume check on quote requests.
- AI output validation: 30 minutes adding not-empty and refusal-phrase filters after two AI steps.
About four and a half hours in total, plus 15 minutes a week. In the following quarter, three problems occurred, and all three were spotted within a day: the form connection (volume check), an expired Google connection (heartbeat) and the invoice loop (weekly usage check). The previous quarter's two failures had taken 11 days and 6 days to notice. Deciding who looks after all this, especially once there are more than a handful of automations, is part of the wider question in who should own AI in a small business. And if any of these automations talk to customers directly, the extra checks in monitoring AI that talks to customers apply on top.
Keeping automations honest
Should Zapier's Autoreplay always be switched on?
For most Zaps, yes, because many errors are temporary, such as an app briefly unavailable. Be careful with steps that create or send things. If a step timed out after the other app had already acted, a replay can create a duplicate invoice, email or record. For those, prefer an error handler that alerts a person, or add a search step that checks whether the item already exists.
Who should receive automation alerts in a small business?
At least two people, through a shared address or channel rather than one person's inbox. One should be the automation's owner, who can fix it, and the other someone who notices when a process stops, such as the office manager. Alerts sent only to the person who built the automation are the most common reason failures go unnoticed after staff changes or during holidays.
How many automations can one person realistically monitor?
With alerts, heartbeat checks and a weekly review in place, one person can comfortably look after 20 to 40 automations in a small business, spending 15 to 30 minutes a week. Without those, even five is hard, because every check means opening each Zap or scenario by hand. Rank them by the damage a silent failure would cause and protect the top five first.
Further reads
- How to Add Human Approval Steps to AI Automations — Add a person before the actions that can't be undone.
- Power Automate for Small Businesses: When It Beats Zapier — How Power Automate's own auto-off rules differ.
- Zapier vs Make vs n8n for AI Automation: Which Fits Your Business? — Compare monitoring features when choosing a platform.
- When Does Zapier Get Too Expensive? Finding the Tipping Point — Running out of tasks is a failure mode too.
- How to Map a Business Process Before You Automate It — Know what 'working' looks like before you monitor it.
- A Simple AI Risk Register for Small Businesses (With Template) — Record your critical automations and their checks.
- How to Tell If a Process Is Ready to Automate With AI — An eight-point readiness scorecard with two automatic vetoes, and a driving school's cancellation process scored 7, fixed, then rescored 13.
- AI Incident Response Plan for Small Businesses (With Template) — A two-page AI incident plan for small teams: what counts as an incident, who leads, severity levels, first-hour actions and a template to copy.
- Is It Worth Automating a Task You Only Do Once a Week? — The payback sum for weekly tasks, the factors beyond time that change the answer, and three worked cases: one yes, one partly, one no.
- Can AI Work With the Tools Your Business Already Uses? — The four ways AI connects to the software you run on, a 30-minute stack audit, an insurance broker's tools mapped, and workarounds for old systems.
- What Is a Webhook? Why Some Automations Run Instantly — Why some automations fire in seconds and others lag by minutes, how to tell which you have, and how to keep webhooks from failing silently.
- How to Chase Missing Client Records With AI Before Deadlines — Build a missing-items list, send AI-written chasers on a ladder counted back from each deadline, and let AI sort the replies for you.
- How Coaches Use AI to Qualify Discovery Call Bookings — Set fit criteria, ask four questions before the calendar, let AI sort and brief each booking, and send people who aren't ready somewhere useful.
- Replying to Trade Enquiries in 60 Seconds With AI — Why most automations can't reply in 60 seconds, the three builds that can, and what a first reply to a roofing or plumbing enquiry should actually say.
- Can AI Handle Vaccination and Check-Up Reminders for a Vet? — What AI can and can't do for a vet practice's vaccination and check-up reminders, with a message sequence, a reply-triage prompt and worked numbers.
- Do You Need an AI Consultant to Set Up Your Marketing Automation? — A scoring table and worked example showing which marketing automations you can build yourself and when outside help pays for itself.
- AI Tools and AI Development: The Complete 2026 Guide — the AI hub, including every tutorial in the AI-for-business series.
Sources: Zapier help articles on error notifications, custom error handling, Autoreplay, run statuses and Zapier Manager; Make help pages on error handling, error handlers, incomplete executions and scenario settings; Healthchecks.io pricing page. Checked September 2026.