/interfacer.
FeaturesLong read

What Your SLA Actually Covers When a Third-Party API Goes Down

Most SLA credits reimburse a tiny fraction of what downtime actually costs your business.

Senior Writer · · 8 min read
Cover illustration for “What Your SLA Actually Covers When a Third-Party API Goes Down”
Features · August 6, 2026 · 8 min read · 1,875 words

Start with the number everyone anchors to: 99.99% uptime. Four nines. The majority of organizations now require it as a minimum from their vendors. It sounds airtight.

It is not.

99.99% uptime means your vendor is contractually allowed to be down for 52.6 minutes per year. That is 4.38 minutes per month. One honest multi-hour outage eats through that entire annual budget in a single afternoon. Think of it like a gas tank labeled "full" that only holds enough fuel for a short trip — the label is accurate, but the capacity is not what you imagined.

Now layer in what major providers actually deliver. The big cloud infrastructure players commit to 99.99% per region in their SLAs. In 2025, actual delivered uptime across regions averaged somewhere between 99.95% and 99.97%. Below their own stated commitments. And each of them had at least one multi-hour regional outage that year, which means the gap between the number on paper and the number in practice is not hypothetical. It happened.

Enterprise SaaS vendors are usually less aggressive than that. Standard agreements often commit to 99.5%. Getting to 99.9% typically costs a 15 to 30 percent price premium.

Then there is the measurement-period trap, which is where things get quietly absurd.

Many providers measure SLA compliance monthly, not annually. Nine separate 55-minute outages spread across nine different months will each stay under the monthly threshold. Zero credits triggered. Take that exact same total downtime and concentrate it in a single month, and you hit the maximum credit tier. Same amount of downtime. Completely different outcome depending on the calendar. When you are reviewing or negotiating an SLA, push for annual measurement periods.

One more thing the numbers do not capture: the definition of "downtime" itself is usually narrow. Partial degradation does not count. Elevated error rates also fail to qualify. Slow responses do not count. None of those qualify, even when they are breaking your dependent integrations just as thoroughly as a full outage would. So your users are getting a broken experience, your engineers are in a war room, and the clock technically is not running.

The standard carve-outs that remove third-party failures from SLA coverage entirely

Nearly every enterprise SLA contains explicit language that removes provider liability when a third-party API or infrastructure failure is the root cause. This is not a loophole. It is not sneaky fine print buried on page 47. This is standard contract drafting, and it shows up in publicly available agreements from real vendors.

One commercial SLA excludes coverage for "circuit or connectivity outages, or third-party connection or system outages (such as payment processors or third-party loyalty hosts)" and "issues resulting from inadequate bandwidth or related to third-party software or services." Supabase's published SLA states they are "not responsible for outages or service interruptions caused by factors outside of its reasonable control," explicitly naming AWS, Cloudflare, GCP, Azure, and GitHub as excluded categories. Striim Cloud's SLA says outages from their cloud service providers "will not be included in the Downtime Minutes."

Read that last one again slowly.

Most APIs run on major cloud infrastructure. AWS or Azure goes down, the vendor goes down with it. And the same outage that took the vendor offline also triggers the contractual exclusion that protects the vendor from liability. The outage causes the exclusion. You cannot separate them. That is not a drafting accident. That is the contract working as intended. It is like being handed an umbrella with a hole in it — the hole is not a defect, it is a feature of the design.

Force majeure adds a second layer of protection for the vendor. SaaS providers routinely argue that a cloud provider outage is "outside their reasonable control." Striim's language covers "systemic electrical, telecommunications, or other utility failures," which is broad enough to absorb most infrastructure events. There is a reasonable counterargument that choosing a single cloud provider without redundancy is a foreseeable and preventable risk, not an act of God. Courts and arbitrators genuinely do not agree on where that line sits, which means you are litigating ambiguity in an already bad situation.

One more distinction: some providers offer SLOs rather than SLAs. Service level objectives rather than agreements. An SLO is a target. No contractual remedy attached. Some providers will tell you there is no meaningful difference. There is, and it matters when you are the one holding the bag after an outage.

If your product depends on third-party APIs, and those APIs run on shared cloud infrastructure, the outage scenario most likely to affect your users is exactly the one your SLA is least likely to cover.

What service credits are actually worth when downtime costs hundreds of thousands per hour

A major e-commerce retailer lost $2.1 million in revenue over six hours during a payment processing outage. Their SaaS provider's SLA offered $3,200 in service credits. That is roughly 0.15% of actual losses.

The contract performed exactly as written.

That is not an extreme outlier. That is just how the math works. A company paying $10,000 a month for a critical SaaS application typically receives somewhere between $500 and $1,500 in credits for a qualifying outage. If that outage causes $300,000 in business losses, the SLA covers less than one percent of the damage. The credit is calculated against what you pay the vendor, not against what the outage costs your business. Those two numbers are not even in the same conversation.

The AWS EC2 SLA is a useful reference because it is public and widely cited. Below 99.99% but above 99.0%, you get a 10% credit. Below 99.0% but above 95.0%, you get 30%. A 100% credit only kicks in below 95.0% uptime, a threshold that is almost never breached. The tier that would actually move the needle is the tier that almost never applies.

Then there are credit caps. Supabase caps total service credits at 20% of fees paid in the preceding 12 months. Anything above that ceiling is simply forfeited, regardless of what the outage actually cost you.

And many SLAs include what is called a "sole and exclusive remedy" clause. That language means service credits are your only recourse for SLA violations, full stop. It closes off litigation even when your actual losses are catastrophic and the credit ceiling is a rounding error by comparison. You agreed to it when you signed.

The Delta Airlines situation with CrowdStrike is the large-scale version of this. Roughly $500 million in business losses. A fraction of that recovered through SLA mechanisms. A gap of about seven times between what the outage cost and what the contract covered. Delta is not a small company with a weak negotiating position. That gap exists even at scale.

The broader picture: 93% of organizations report that a single hour of downtime costs more than $300,000. Average downtime cost reached $8,600 per minute in 2025, up from $5,600 per minute in 2022. The actual cost of an outage is rising faster than contract values. The gap is getting wider, not narrower.

Diagram: The Credit Gap: $2.1M Loss, $3,200 Covered. Visualizes: Show the brutal magnitude contrast between actual downtime costs and SLA service credits.

How the claims process is designed to reduce what providers actually pay out

Even for outages that clearly qualify under the SLA definition, credits are not automatic. You have to claim them. That burden sits entirely with you.

AWS requires credit claims within 30 business days of the end of the affected billing cycle. Azure requires submission within two months of the end of the billing month containing the incident. Google Cloud Maps requires the customer to report the suspected problem within 30 days, and a provider-issued outage notice does not automatically substitute for a formal claim. You still have to file.

You also need evidence. Logs. Request IDs. Timestamps. Error data showing your specific account was affected. A globally documented outage is sometimes accepted as sufficient proof, but not always, and the standard does not guarantee it.

I once watched a team spend three days reconstructing an incident timeline after the fact — digging through logs they were not specifically keeping, trying to meet a documentation standard they did not know existed until they needed it. Their engineers joked, grimly, that the only thing worse than the outage was filling out the paperwork to prove it happened. The claims window was already closing while they were still figuring out what they needed to submit.

Combine a narrow definition of downtime with a short claims window and a documentation burden, and you have a structure that will reduce the volume of credits actually paid out. Not through bad faith. Just through design.

The practical implication is simple. You need monitoring, logging, and a claims-ready incident record before an outage happens, not after. By the time you are reconstructing what happened from memory and incomplete logs, you are already behind.

What realistic resilience looks like once you accept what the SLA won't cover

Start with contract hygiene, because even if the SLA cannot make you whole, a better-negotiated one is still worth having.

A few things worth pushing on before you sign:

  • Annual measurement periods instead of monthly. The math is just better for you.
  • Explicit coverage or defined liability for third-party infrastructure failures. If the contract only covers direct provider failures, you have now documented in writing the problem you are most likely to face and excluded it from protection.
  • A broader definition of downtime that includes partial degradation and elevated error rates above a defined threshold, not just complete unavailability.
  • "Sole and exclusive remedy" language. Understand exactly what it forecloses before you agree to it, because by the time you care about it, it is too late to negotiate.
  • Force majeure scope. Get clarity on what qualifies, and whether single-cloud dependency without a failover plan is treated as a customer risk or a vendor risk.

That said, even a well-negotiated SLA is compensatory, not preventive. Credits arrive weeks after the incident. Your users experienced the outage regardless of what the contract says. The contract does not un-ring that bell.

The layer that actually prevents user impact is operational. Monitor third-party API health directly, not just your own services. Cascade failures need to be caught upstream before they propagate. Design for graceful degradation, so a third-party outage produces a limited, degraded experience rather than a hard failure. Track API versioning and deprecation actively, because a lot of outages are not actually unplanned. They are planned breaking changes that land without adequate notice. And look at where managed connectors or abstraction layers can shrink the surface area where a provider-side change silently breaks something in your product.

The 60% increase in API downtime in Q1 2025 is not an anomaly. It is a direction. Treating the SLA as your primary safety net means accepting that your users will absorb the cost every time an outage lands inside one of the standard exclusions.

The SLA documents the minimum acceptable service level and the limited, compensatory recourse available when that standard is missed. It is a paper record of what you are owed after something goes wrong. Operational resilience is what keeps something from going wrong in the first place. Those are two different jobs, and conflating them is how you end up surprised by a $3,200 credit on a $2.1 million problem.

Sources

  1. truto.one
  2. supabase.com
  3. jchanglaw.com

More in Features