What Uptime and Support Should an AI Vendor Promise You?

Coding Liquids tutorial cover featuring Sagnik Bhattacharya for What Uptime and Support Should an AI Vendor Promise You?
Coding Liquids tutorial cover featuring Sagnik Bhattacharya for What Uptime and Support Should an AI Vendor Promise You?

For a tool your business depends on, an AI vendor should promise at least 99.9% monthly uptime in writing (about 43 minutes of downtime a month), service credits when it misses, a support response within an hour for a full outage during your working hours, and at least 60 days' notice before retiring or changing the AI model underneath.

Most self-serve AI plans promise none of that, and the ones that do often cover less than you'd think. Google's Workspace SLA (service level agreement) commits to 99.9% for Gmail, Docs, Drive and the rest, but the Gemini app isn't on its list of covered services. Anthropic describes its standard API tier as best-effort availability. And an uptime promise measures whether the service is reachable, not whether it's fast or right.

Follow me on Instagram@sagnikteaches

What 99.9% really allows: the downtime maths

Uptime percentages sound alike and aren't. Here is what each level allows in a 30-day month and over a year:

Connect on LinkedInSagnik Bhattacharya
Uptime promisedDowntime allowed per 30-day monthDowntime allowed per year
99%7.2 hoursAbout 3.7 days
99.5%3.6 hoursAbout 43.8 hours
99.9%43 minutesAbout 8.8 hours
99.95%22 minutesAbout 4.4 hours
99.99%4 minutesAbout 53 minutes

Two details change what those figures mean in practice. The first is the measurement period: a monthly commitment is stricter than a yearly one, because a yearly 99.9% can absorb a single eight-hour outage and still be met. The second is the definition of downtime. Google's Workspace SLA counts a period as downtime when the web interface has more than a five percent user error rate, measured on its servers. A partial outage that breaks the service for 4% of your requests doesn't count, even if one of them is yours.

Subscribe on YouTube@codingliquids

For AI tools there's a third gap. The service can be up while the AI inside it is slow, rate-limited or answering worse than last week. No uptime percentage covers quality. That needs different promises, covered in the checklist below.

Who promises what today: three reference points

It helps to know what large vendors commit to before you judge a small one. These are from the vendors' own documents, as published in September 2026:

  • Google Workspace: 99.9% monthly uptime for the covered services. If it misses, the credit is 3 days of service added for uptime between 99.0% and 99.9%, 7 days between 95.0% and 99.0%, and 15 days below 95.0%, capped at 15 days a month. You must claim through a support case within thirty days or lose the credit, and the SLA is your sole and exclusive remedy for missed uptime.
  • Anthropic's API: the standard tier runs with best-effort availability. Priority Tier, which targeted 99.5% uptime, is no longer sold; existing commitments run to their end date, and anyone needing guaranteed capacity is pointed to sales.
  • OpenAI's API: its Scale Tier, where enterprise customers buy capacity in advance, comes with a 99.9% uptime SLA. Ordinary pay-as-you-go use doesn't carry that promise.

Model retirement is the AI-specific promise most buyers never ask about, and the two big model providers publish theirs. Anthropic gives at least 60 days' notice before retiring a publicly released model: it told developers on 5 June 2026 that Claude Opus 4.1 would retire, and it did on 5 August 2026. OpenAI promises at least six months for generally available models, at least three for specialised variants, and warns that preview models may go with much shorter notice, such as two weeks. Requests to a retired model simply fail, so any automation pinned to one stops on that date.

Service credits won't cover what an outage costs you

Service credits look like compensation until you do the sum. Take an illustrative packaging supplier with 25 people on Google Workspace Business Standard at about $14 a user a month, roughly $350 a month in total. Suppose one month's uptime falls to 99.5%, which is 3.6 hours of downtime:

Service credit:  3 days of service = $350 x 3/30        = about $35
Hours lost:      3.6 hours x 25 staff x $30 an hour     = $2,700
Credit as a share of the loss: $35 / $2,700             = about 1.3%

The credit covers about one percent of the cost, and only if someone remembers to open a support case within thirty days. So don't choose a vendor, or pay more, for a generous credit table. An SLA is worth having for three other reasons: it tells you the vendor measures its own reliability, it gives you a written standard to hold them to, and, if you negotiate for it, it gives you the right to leave. That last one is the clause to ask for: termination without penalty if the vendor misses its SLA in, say, three months out of any six.

The promises to ask for, and how to check each one

Use this as the checklist for any AI tool your business will rely on. The middle column is what I'd treat as a reasonable minimum for a small business; it isn't a market average, and your own risk decides whether to ask for more.

Uptime terms

PromiseReasonable minimumHow to check it
A written SLAIn the contract or a document it links toIf it's only on a marketing page, you don't have one
Uptime level and period99.9%, measured monthlyRead the definition of "Monthly Uptime Percentage" or similar
What counts as downtimeIncludes partial outages of the features you useLook for error-rate thresholds and what's excluded
The AI features are coveredNamed in the covered servicesCompare the covered list with the features you're buying for
Third-party exclusionsThe vendor's own model provider isn't excludedSearch the SLA for "third party" and "upstream"
Scheduled maintenanceAnnounced in advance, outside your working hoursCheck the notice period and the usual window
Credits and claimsA clear table and a claim window of at least 30 daysDiary the claim deadline after any incident
Chronic failure exitRight to terminate after repeated missesUsually negotiated; ask before signing

Support terms

PromiseReasonable minimumHow to check it
Hours and channelsYour working hours, in your time zone, with a way to reach a personOpen a test ticket during the trial and time the reply
Severity definitionsWritten, with "service down" as the top levelAsk for the support policy document
Response targetsService down: 1 hour. Major feature broken: 4 hours. Minor: next working dayTargets should be in the policy, not just the sales deck
Updates during incidentsStatus page updates at least hourly during an outageSubscribe to the status page during the trial
Post-incident reportsA written cause and fix for major incidentsAsk to see a past one
EscalationA named route beyond the first-line queueGet it in writing at onboarding

AI-specific terms

PromiseReasonable minimumHow to check it
Notice before model changesAt least 60 days, matching Anthropic's floor for its own modelsAsk how you'll be told, and whether you can test the new model first
Behaviour changesRelease notes when prompts, models or features change outputAsk where changes are announced
Capacity at busy timesA stated policy on rate limits and throttlingAsk what happens when their upstream model is overloaded
FallbackA plan if their model provider has an outageAsk whether they can switch providers, and how fast
Exit and data exportFull export in a usable format, with a period after terminationTest an export during the trial
Discontinuation noticeMonths, not weeks, before a product is withdrawnAsk what notice the contract gives

The last row isn't theoretical. OpenAI closed the Sora apps on 26 April 2026 and the Sora API on 24 September 2026, and the calendar tool Clockwise shut down on 27 March 2026 and deleted user data rather than transferring it. No uptime percentage protects you from a product being withdrawn. Checks before committing to an AI vendor that might shut down covers that risk in depth, and keeping your data and prompts portable covers the export side.

Matching the promise to how you use the tool

A single standard for every tool wastes effort. The question is what a two-hour outage would actually cost, and when it would happen. Four illustrative businesses:

Business and toolA two-hour outage meansUptime to ask forSupport hours needed
Care agency: AI answering carers' and families' calls out of hoursMissed calls about visits and welfare at 7am99.9% plus a manual fallbackIncluding evenings and weekends
Packaging supplier: AI quoting assistant used 9 to 5Quotes go out a morning late99.9%Business hours
Import-export business: overnight extraction of shipping documentsThe batch reruns an hour later99.5% is fineNext working day
Laboratory: AI transcription of internal meetingsSomeone takes notes by handNo SLA neededEmail is enough

The care agency is the only one where the vendor's support hours matter as much as its uptime. A 99.9% promise with support from 9 to 5 on weekdays still leaves a Saturday-morning outage unhandled until Monday.

A realistic SLA trap: the upstream exclusion

An illustrative home-care provider bought an AI scheduling assistant whose SLA promised 99.9% monthly uptime. One Tuesday the assistant stopped working for five hours because the model provider it relies on had an outage. The provider claimed its service credit, and the vendor declined: the SLA excluded "failures of third-party services, including upstream AI model providers". By the vendor's own definitions, the month's uptime had been met.

It showed up only when the credit claim was refused, months after signing. The fix is to search any SLA for "third party", "upstream" and "subprocessor" before signing, and to ask directly: if your model provider goes down, does that count as your downtime? A vendor built on another company's model can't promise more than that company gives it, so the honest answers are either "yes, it counts" or a clear description of how they fail over to a second provider. The same sub-processor chain matters for security, which reading an AI vendor's SOC 2 and ISO 27001 paperwork covers.

Measuring a vendor's uptime yourself during a trial

Don't take a vendor's historical uptime on trust. For a 30-day trial, subscribe to its status page, and keep a simple log of every incident that affected your team: start time, end time, what broke. A filled-in example from the packaging supplier's trial:

Date     Start  End    Minutes  What broke
03 Sep   10:12  10:41  29       Quote drafts failing (vendor status: "degraded")
11 Sep   14:05  14:23  18       Login errors for all users
24 Sep   09:30  10:18  48       AI replies timing out (vendor status: "operational")
Total                  95

Uptime for the month: (43,200 - 95) / 43,200 = 99.78%

Two things stand out. The month came in below 99.9%, which is useful when you negotiate. And on 24 September the vendor's status page said everything was operational while the AI replies were timing out, which is exactly the kind of degraded performance an uptime definition can miss. Raise both with the vendor before you sign, and ask how it would have classified that 48 minutes.

Test the support promise as well

Response targets are easy to write and harder to meet, so test them. During the same trial, the packaging supplier opened three genuine tickets at different times and noted how long each took to get a first reply from a person:

TicketOpenedSeverity chosenFirst human replyAgainst the stated target
How to export quote historyTuesday 10:05Minor3 hours 40 minutesMet (next working day)
AI drafts failing for every userFriday 16:50Service down2 hours 10 minutesMissed (1 hour)
Question about user permissionsSaturday 11:20MinorMonday 09:15Met (weekends not covered)

The one that matters is the Friday afternoon outage, and it missed. That doesn't have to end the trial, but it's the question to raise with the account manager: was that a one-off, or is late Friday outside the hours the target really applies to? The answer tells you more about the support you'll get than the policy document does.

Getting an assistant to pull the terms out of an SLA

SLAs are short compared with most contracts, but the terms you need are scattered across definitions, exclusions and credit tables. An AI assistant on a business plan can pull them into one place in a minute. A prompt the laboratory's office manager used on a vendor's published SLA:

Below is a vendor's service level agreement. Extract, quoting the
exact clause for each:
1. Uptime percentage and measurement period
2. How downtime is defined
3. Which services or features are covered, as a list
4. Everything excluded from downtime
5. The service credit table, any cap, and the claim deadline
6. Any right to terminate for repeated failures
If a point isn't in the text, write "not stated". Don't infer.
[SLA text pasted here]

Part of the illustrative answer:

1. 99.9%, calendar month (clause 2.1)
3. Covered: the platform, including all AI features (clause 1.4)
6. Not stated

Point 3 needed checking, and it was wrong. Clause 1.4 listed "the web application, the mobile application and the API"; the AI summaries ran on a separate add-on that the SLA never mentioned. The assistant filled the gap with an assumption, even though the prompt asked it not to. Point 6, "not stated", was accurate and useful: it told the lab to ask for an exit clause before signing. Use the extraction as a map of where to look, then read each quoted clause yourself, especially any answer that sounds reassuring.

A fallback plan for when the AI is down anyway

Even a 99.9% promise allows 43 minutes a month, and outages don't book appointments. For anything customer-facing, write a one-page fallback before go-live. A filled-in version for the care agency's out-of-hours line:

Trigger:      AI line not answering, or status page shows an incident,
              for more than 10 minutes.
Who decides:  On-call coordinator (rota on the office wall and in
              the shared drive).
Action 1:     Divert the out-of-hours number to the on-call mobile
              (steps saved in the phone provider's app).
Action 2:     Text carers due in the next 2 hours: "Our phone
              assistant is down. Call [on-call mobile] with any
              visit changes."
Action 3:     Log every call by hand on the paper sheet in the
              on-call bag.
Restore:      Switch the divert back, test with one call, and type
              up the paper log.
Afterwards:   Note the start, end and cause; claim any credit within
              the SLA's deadline; review at the next team meeting.

Test the divert once during a quiet hour. A fallback nobody has tried is a plan on paper only.

Asking for terms the standard contract doesn't include

Small customers can't always renegotiate, but asking costs one email and often gets a written answer you can rely on later. A filled-in request from the care agency:

Subject: Service levels for [product] before we sign
Hello [first name], before we commit to an annual plan, could you confirm in writing: your uptime commitment and how downtime is measured; whether the AI call-answering feature is covered; whether an outage at your AI model provider counts as downtime; your support hours and response targets for a full outage, including weekends; how much notice you give before changing the underlying model; and whether we can terminate without penalty if the SLA is missed in three months out of six. Thank you, [name], Registered Manager.

If the answers matter enough to change your decision, make sure they end up in the order form or contract, not just the email thread. The tutorial on AI software contracts, renewals and notice periods covers getting side letters and email promises into the paperwork, and what to ask an AI chatbot vendor has the wider question list if the tool talks to your customers.

SLA questions small buyers ask

Is 99.9% uptime good enough for a small business?

For most tools, yes. It allows about 43 minutes of downtime in a 30-day month, usually spread across a few short incidents. Ask for more only if a customer-facing process stops completely when the tool does, and even then a fallback plan is cheaper than a higher uptime tier. Check how downtime is defined, because a generous definition can make 99.9% mean less.

Do ChatGPT or Claude business plans come with an uptime guarantee?

Look for an SLA in the terms or order form for the plan you're buying; if there isn't one, you don't have a guarantee. On the API side, Anthropic describes its standard tier as best-effort availability, and OpenAI reserves its 99.9% SLA for enterprise capacity bought in advance. Larger contracts can negotiate terms that self-serve plans don't include.

What's the difference between response time and resolution time?

Response time is how quickly the vendor acknowledges your ticket and starts work. Resolution time is how quickly the problem is fixed. Most support promises are response times only, so a one-hour response target can still mean a day-long outage. Ask for both, and for regular updates during a serious incident.

Should I pay extra for premium support?

Only if the tool runs a process that can't wait until the next working day, or you have no one who can troubleshoot it. Price the premium against one realistic outage: the hours lost, the customers affected, the work redone. If that figure is smaller than a year of premium support, a fallback plan is the better buy.

Further reads

Sources: Google Workspace Service Level Agreement (last modified 31 August 2026); Anthropic API documentation on service tiers and model deprecations; OpenAI API deprecations page; OpenAI Scale Tier page (search listing). Checked September 2026.

Relying on an AI tool you can't afford to lose?

On a 1:1 call we'll look at how your business depends on the tool, what its current terms actually promise, and the fallback that keeps work moving when it's down.

Book a 1:1 call with me