When Every AI Provider Goes Down at the Same Time
Early this month three of the big model providers degraded within a few hours of each other, and a lot of AI features went dark with them. If your product calls one provider and waits, their outage is now your outage. This is how to build a resilience layer in front of your model calls: a circuit breaker that fails fast, health-aware failover across providers, request hedging for the slow tail, and a degraded mode that keeps the lights on when nothing is healthy.
On the third of this month, within the space of an afternoon, three of the major model providers degraded one after another. Not a clean total outage you can point at and wait out, but the worse kind: elevated error rates, timeouts, partial failures, the sort of thing where half your requests succeed and the other half hang for ninety seconds before dying. Teams that had quietly depended on a single provider spent that afternoon watching their AI features return spinners and 500s, and watching support tickets pile up for something they could not fix from their side.
If you run anything agentic or LLM-backed in production, the lesson was not subtle. Your model provider is a dependency like any other, and it will have bad days. The only question is whether a bad day over there turns into a bad day over here. This post is about the layer that decides the answer: a small resilience gateway that sits between your application and the providers, and makes a sensible choice when one of them starts to wobble.
The business problem: one provider is now single point of failure
It is easy to treat the model API as infrastructure that is simply always there, the way you treat DNS. It is not. The providers are younger, they are under enormous load, and this month proved they can degrade at the same time rather than politely taking turns. When you call one provider and block on the response, you have inherited its entire reliability curve. Its p99 is your p99. Its outage is your outage. Its rate limit, hit during a traffic spike, is your checkout flow timing out.
The cost of getting this wrong is not abstract. If your product generates support replies, summarizes documents, or runs an agent that customers are mid-task with, a thirty minute provider wobble is thirty minutes of failed work, abandoned sessions, and a spike in churn risk that outlasts the outage. The fix is not heroics during the incident. It is a boring layer you build once, that turns a provider failure into a slightly slower or slightly cheaper response instead of an error.
That layer has four jobs: stop sending traffic to a provider that is clearly failing, move that traffic somewhere healthy, shave the slow tail when a provider is up but sluggish, and have something honest to return when genuinely nothing is healthy. Let us build each piece.
Retries are not failover, and mixing them up hurts
The first instinct when a call fails is to retry it. For a transient blip, a single dropped socket, a brief 429 with a short reset, that is exactly right. Retry with a little backoff and you recover invisibly.
The trap is retrying your way through a real outage. If the provider is down hard for twenty minutes, every incoming request dutifully burns its full retry budget, each attempt waiting for a timeout before the next. Your worker pool fills with requests that are all sleeping between doomed retries. New requests queue behind them. Latency climbs for everyone, including the requests that could have been served by a healthy fallback. You have converted one provider's outage into your own resource exhaustion, which is strictly worse because now you are down even for the users whose work did not need that provider at all.
So the rule is: retries for blips, failover for outages, and a circuit breaker to tell the two apart. The circuit breaker is the component that remembers recent failures and fails fast once a provider has clearly gone bad, so requests stop piling up behind a dead dependency.
A circuit breaker that fails fast
A circuit breaker is a tiny state machine with three states. Closed means traffic flows normally. Open means the provider has failed enough recently that we skip it entirely and fail instantly. Half-open is the careful probe: after a cooldown, we let a single request through to see if the provider has recovered.
import time
from enum import Enum
class State(Enum):
CLOSED = "closed" # normal, let traffic through
OPEN = "open" # provider is bad, fail fast
HALF_OPEN = "half_open" # cooldown elapsed, allow one probe
class CircuitBreaker:
def __init__(self, fail_threshold=5, cooldown=30.0):
self.fail_threshold = fail_threshold # consecutive fails before opening
self.cooldown = cooldown # seconds to wait before probing
self.failures = 0
self.state = State.CLOSED
self.opened_at = 0.0
def allow(self) -> bool:
"""Call before using the provider. False means skip it."""
if self.state is State.OPEN:
# Enough time passed? Move to half-open and allow one probe.
if time.monotonic() - self.opened_at >= self.cooldown:
self.state = State.HALF_OPEN
return True
return False
return True # CLOSED or HALF_OPEN both allow the attempt
def record_success(self):
# A good probe (or any success) resets the breaker fully.
self.failures = 0
self.state = State.CLOSED
def record_failure(self):
self.failures += 1
# A failed probe in half-open sends us straight back to open.
if self.state is State.HALF_OPEN or self.failures >= self.fail_threshold:
self.state = State.OPEN
self.opened_at = time.monotonic()
The whole point is the allow() check at the top. When a provider is in the open state, that call returns False in microseconds, and your request skips it without paying the timeout. During the outage this month, that single check is the difference between a request that fails over in a few milliseconds and one that hangs for a minute and a half first.
Keep one breaker per provider, and ideally per model, because a provider can be fine on one model and rate-limiting another. Tune fail_threshold and cooldown to taste: a low threshold trips fast but is twitchy on noisy networks, a longer cooldown keeps you off a flapping provider but slows recovery once it is genuinely back.
A failover gateway with health-aware routing
Now wrap the breakers in a gateway that tries providers in order and skips the ones whose circuit is open. This is the component your application actually calls instead of the raw SDK.
from dataclasses import dataclass
from typing import Callable
@dataclass
class Provider:
name: str
call: Callable[[str], str] # your adapter: prompt in, text out, raises on failure
breaker: CircuitBreaker
class LLMGateway:
def __init__(self, providers: list[Provider], max_attempts: int = 3):
# Order matters: cheapest/best first, fallbacks after.
self.providers = providers
self.max_attempts = max_attempts # hard cap across ALL providers
def complete(self, prompt: str) -> str:
attempts = 0
last_error: Exception | None = None
for provider in self.providers:
if attempts >= self.max_attempts:
break
if not provider.breaker.allow():
continue # circuit open: skip without spending a timeout
attempts += 1
try:
result = provider.call(prompt)
provider.breaker.record_success()
return result
except Exception as e: # network error, 5xx, 429, timeout
provider.breaker.record_failure()
last_error = e
continue # fall through to the next provider
# Everyone we were allowed to try has failed.
raise RuntimeError("all providers exhausted") from last_error
Two details carry most of the weight here. The max_attempts cap is a cost and latency guard: without it, an outage that trips every breaker could still have a request walk your entire provider list, and the fallbacks are usually pricier per token. An outage plus uncapped failover is how a bad afternoon becomes a five-figure invoice, so the cap is not optional. The breaker.allow() skip is what keeps the gateway fast under load, because open circuits cost nothing to pass over.
Order your provider list deliberately. The primary is your best price-for-quality choice, and the ones after it are there to keep you serving, not to be perfect. This pairs naturally with cost-aware model routing: routing picks the right model for the job on a normal day, and this gateway sits underneath it to keep requests flowing when a provider that routing wants is down.
One more thing worth doing: because fallback providers return different output shapes and refuse different things, validate every response against the same schema no matter who served it. A fallback you only hit twice a year is exactly the path most likely to return something malformed, and you do not want to discover that during the outage you built it for.
Hedging and degraded modes for the slow tail
Failover handles a provider that is clearly down. The harder case is a provider that is up but slow: most requests are fine, but your p99 just blew out to twelve seconds because a fraction of calls are crawling. Retries do not help because the request has not failed, it is just late.
Hedging helps here. Fire the request at your primary, and if it has not answered within a short deadline, fire a second copy at a fallback without cancelling the first. Take whichever returns first. You pay for one extra call on the slow tail in exchange for cutting the worst-case latency hard.
import asyncio
async def hedged_complete(gateway_primary, gateway_fallback, prompt: str,
hedge_after: float = 2.0) -> str:
"""Start primary; if it is slow, race a fallback and take the first done."""
primary = asyncio.create_task(gateway_primary(prompt))
done, _ = await asyncio.wait({primary}, timeout=hedge_after)
if primary in done:
return primary.result() # primary beat the hedge deadline
# Primary is slow. Race it against a fallback; first finisher wins.
fallback = asyncio.create_task(gateway_fallback(prompt))
done, pending = await asyncio.wait(
{primary, fallback}, return_when=asyncio.FIRST_COMPLETED
)
for task in pending:
task.cancel() # stop paying for the loser
return next(iter(done)).result()
Hedge only a small slice of traffic, because hedging everything doubles your spend and doubles the load you put on providers during the exact moment they are already struggling. A deadline near your normal p95 is a good starting point: most requests finish before it and never hedge at all. Keep an eye on the interaction with backpressure and rate limiting, since a hedge storm is a quiet way to trip your own concurrency limits.
Finally, the honest part. Sometimes every provider is down at once, which is precisely what happened this month, and no amount of failover conjures a healthy backend. Decide in advance what you return then. For some features that is a cached answer, for a search box it is keyword results instead of a generated summary, for an agent it is a clear message that the model is unavailable and the user's work is saved. A degraded mode that tells the truth beats a spinner that lies, and it is the single thing users remember kindly after an outage.
The traps that turn failover into a bigger outage
Correlated failure is the one that bites hardest. If your primary and your fallback are the same underlying model family hosted two ways, or two providers that both sit on the same cloud region, they go down together and your failover is theater. Pick fallbacks that fail for different reasons than your primary, and actually check where they run.
Untested fallbacks rot. A failover path you never exercise is a path that has silently broken since you wrote it: an expired key, a renamed model, a schema that drifted. Run a tiny scheduled job that sends a real request down every provider in your chain, so you learn the fallback is broken on a Tuesday morning rather than during the incident.
Breakers that are too sensitive cause their own outage. Set the failure threshold too low and a couple of unlucky timeouts trip the circuit, dumping all your traffic onto a fallback that was never provisioned for full load, which then falls over too. Watch how often circuits open in normal operation and tune the threshold so that number is near zero on a good day.
Cost blowout hides in the recovery. The scenario to model is not the clean outage, it is the flapping provider: up, down, up, down, every flap pushing traffic to a pricier fallback and back. Track failover volume and fallback spend as first-class metrics, not something you reconstruct from the bill at month end.
The takeaway
The outage this month was a reminder that model providers are a dependency, and dependencies fail, sometimes all together. You cannot stop that from happening, but you can decide what it does to you. A circuit breaker that fails fast, a gateway that fails over to a healthy provider within a hard attempt cap, hedging to tame the slow tail, and an honest degraded mode for the worst case: that is the whole kit, and it is a few hundred lines that you build once and are grateful for exactly when everything else is on fire. Build it before the next bad afternoon, not during it.
If your AI features currently ride on a single provider and you would rather not find out the hard way what that costs, book a consultation call and we can map your failure modes and put a resilience layer in front of them.
Viral Ruparel
Generative AI consultant helping teams ship reliable LLM and agent systems in production.
Contact Viral about your AI project →Frequently Asked Questions
What is the difference between a retry and a failover for LLM calls?+
A retry sends the same request to the same provider again, which is the right move for a transient blip like a single dropped connection or a 429 with a short reset. A failover gives up on that provider and sends the request to a different one. The mistake teams make is retrying their way through a real outage: the primary is down for twenty minutes, every request burns its full retry budget waiting, and the queue backs up until the whole service falls over. Use retries for blips and failover for outages, and let a circuit breaker decide which situation you are in.
Will multi-provider failover make my output quality inconsistent?+
It can, and you should plan for it. A fallback model may format differently, refuse different things, or be weaker at your task. Pin a specific model per provider, validate the output against the same schema regardless of which provider served it, and keep a short eval that runs against every provider in your chain so a fallback you rarely exercise does not silently return garbage the day you finally need it. Treat the fallback as a real path, not a hope.
How do I stop failover from causing a surprise bill?+
Fallback routes are often pricier per token, and an outage plus aggressive failover plus retries is how a bad afternoon becomes a five-figure invoice. Cap the total attempts per request across all providers, put a hard per-request token ceiling on fallbacks, and track failover volume as its own metric so a provider flapping in and out does not quietly double your spend. Pair this with your existing cost controls rather than bolting on a second uncapped path.
Do I need a circuit breaker if I already have retries with backoff?+
Yes, because they solve different problems. Backoff spaces out your retries so you do not hammer a struggling provider, but every request still pays the latency of discovering the provider is down before it gives up. A circuit breaker remembers that the last several calls failed and fails the next one instantly, so requests stop piling up behind a dead dependency. Backoff is polite to the provider, the circuit breaker is protective of your own service.
Related Articles
Your Agent Got the Right Answer the Wrong Way
Your eval checks the final answer and goes green. Meanwhile the agent called six tools to do the work of two, hit a write endpoint it never needed, and landed on the right output by luck. Output-only evals cannot see any of that, and the newer reasoning models take longer, more autonomous paths where it matters more. Trajectory evaluation grades the steps: which tools ran, in what order, with what arguments. Here is how to build it and gate it in CI.
Your Agent Would Rather Guess Than Admit It Doesn't Know
Most agents never say "I don't know." They were trained on benchmarks that reward a confident guess over an honest abstention, so in production they invent a policy, a number, or a citation and say it with a straight face. An abstention gate fixes that: a cheap confidence signal, a threshold you calibrate against a real error budget, and a decision to answer, defer, or escalate. This is how you get an agent that knows its own limits.
Nobody Can Tell You How Many Agents You Are Running
A year of shipping agents leaves most teams with a fleet nobody can inventory: no owner, no declared budget, no record of which tools each one can reach. That is agent sprawl, and it shows up as a surprise invoice and an incident with no name on it. The fix is boring and it works: an agent registry. A typed manifest every agent must declare, a CI gate that blocks anything unregistered or over-budget, and a reconciliation job that flags an agent doing more than it said it would.