Circuit Breakers for LLM API Calls
When an LLM provider degrades, retries make it worse. A practical guide to adding a circuit breaker in Python: the three states, tuning, and failure modes.
In an earlier post on rate limits, retries, and backoff, I mentioned almost in passing that one reason to write your own retry loop, instead of leaning on the SDK, is that you might want a circuit breaker. This post is the part I skipped. Retries handle a request that failed for no lasting reason: a dropped connection, a single 503, a blip. They are the wrong tool when the provider is actually down.
Here is the difference. A retry says “that call failed, try it again in a moment.” When the upstream is healthy, that usually works, because the failure was transient. During a real outage, though, every retry is another request thrown at a service that cannot answer, and the caller waits out its backoff while holding a worker. Multiply that by every in-flight request, and one slow dependency has turned into a fully stalled service. The Google SRE book has a whole chapter on how this cascades.
A circuit breaker is the piece that says “stop trying for a bit.” This post is for engineers running a FastAPI or similar backend that calls an LLM API, who want the service to stay up, and stay honest with its own users, when the model endpoint does not. I build one, tune it, and spend most of the space on the ways it bites you in production. The running example is the kind of backend I worked on for Archi, the retrieval copilot for CMS operations at CERN, where the answer path depends on an external model API that is not under our control.
The pattern: three states
The circuit breaker comes from Michael Nygard’s Release It! and was popularized by a much-cited write-up by Martin Fowler. The name is an analogy. An electrical breaker trips to protect the wiring, and you flip it back once the fault is cleared. In software, the “wiring” is your own service, and the fault is a dependency that has stopped behaving.
A breaker has three states, and the whole design is the movement between them.
Closed is normal operation. Calls go through to the provider, and the breaker counts failures. If the failure count crosses a threshold inside some window, the breaker trips.
Open is the interesting state. The breaker rejects calls immediately, without touching the provider. This is the point of the whole thing: you fail fast instead of piling load onto something that is already struggling, and you give the provider room to recover. A cooldown timer runs while the breaker is open.
Half-open is how the breaker tests the water. When the cooldown expires, it lets a single trial call through. If that call succeeds, the breaker closes and traffic resumes. If it fails, the breaker snaps back to open and the cooldown restarts. That single probe matters: without it, you either reopen the floodgates all at once or stay open forever.
A minimal breaker in Python
You do not need a library to understand this. Here is an async breaker that wraps a call, small enough to read in one sitting. Watch how call moves from open to half-open once the cooldown has passed, how any success resets it to closed, and how _on_failure reopens it. It is deliberately not thread-safe yet; I come back to that later.
import time
from enum import Enum
class State(str, Enum):
CLOSED = "closed"
OPEN = "open"
HALF_OPEN = "half_open"
class CircuitOpenError(Exception):
"""Raised instead of calling the provider while the breaker is open."""
class CircuitBreaker:
def __init__(self, fail_max: int = 5, cooldown: float = 30.0):
self.fail_max = fail_max # failures before we trip
self.cooldown = cooldown # seconds to stay open
self.state = State.CLOSED
self.failures = 0
self.opened_at = 0.0
async def call(self, fn, *args, **kwargs):
if self.state == State.OPEN:
if time.monotonic() - self.opened_at >= self.cooldown:
self.state = State.HALF_OPEN # time to probe
else:
raise CircuitOpenError("circuit is open")
try:
result = await fn(*args, **kwargs)
except Exception:
self._on_failure()
raise
else:
self._on_success()
return result
def _on_success(self):
self.failures = 0
self.state = State.CLOSED
def _on_failure(self):
self.failures += 1
if self.state == State.HALF_OPEN or self.failures >= self.fail_max:
self.state = State.OPEN
self.opened_at = time.monotonic()Wrapping an LLM call is then just a matter of handing the request coroutine to call:
import httpx
breaker = CircuitBreaker(fail_max=5, cooldown=30.0)
async def ask_model(client: httpx.AsyncClient, prompt: str) -> str:
async def _request():
resp = await client.post(
"https://api.provider.example/v1/chat",
json={"model": "some-model", "messages": [{"role": "user", "content": prompt}]},
timeout=httpx.Timeout(30.0, connect=5.0),
)
resp.raise_for_status()
return resp.json()["choices"][0]["message"]["content"]
return await breaker.call(_request)In a FastAPI handler, catching CircuitOpenError is where you decide what the user sees when the model is unreachable. It could be a cached answer, a plain “the assistant is temporarily unavailable,” or a 503 with a Retry-After. The decision is yours, and any of them is better than a request that hangs for 30 seconds and then returns a 500.
For anything real, reach for a maintained library rather than the toy above. pybreaker is the well-worn choice and supports async. purgatory is a newer async-first option with pluggable state storage. Both handle the concurrency and bookkeeping I glossed over, and both give you listener hooks for metrics.
Deciding what counts as a failure
This is the tuning decision people get wrong, and the toy code above gets it wrong on purpose so I can point at it: it treats every exception as a failure. raise_for_status() throws on any 4xx or 5xx, so a 400 Bad Request caused by a malformed prompt in your own code will happily trip the breaker. Now one buggy request path takes down the model for every other user for thirty seconds. That is not the provider’s fault, and opening the circuit does nothing but hide your bug.
The rule: a circuit breaker should trip on failures of the dependency, not on your own mistakes. Sort the responses by whose problem they are:
- 5xx, timeouts, connection errors are the provider’s problem. These should count.
- 429 Too Many Requests is arguable, and it depends. Throttling means the provider is telling you to slow down, which a breaker does, so counting it can be right. But 429 is really the job of backoff and a rate limiter, and leaning on the breaker for it is coarse.
- 4xx other than 429 are your problem. A
400or422means the request was wrong. Retrying or tripping on these is pointless, because the same request will fail the same way. Do not count them.
So the breaker needs a way to classify errors instead of catching everything. A small predicate does it. It returns True for timeouts, connection errors, and 5xx responses, and False for everything else:
def is_dependency_failure(exc: Exception) -> bool:
if isinstance(exc, (httpx.TimeoutException, httpx.ConnectError)):
return True
if isinstance(exc, httpx.HTTPStatusError):
return exc.response.status_code >= 500
return FalseThen call only counts failures that pass the predicate, and re-raises the rest without touching its counters. Every serious library gives you this hook under a name like exclude or a custom exception filter. Use it, because the default of “everything is a failure” is rarely what you want.
Tuning the thresholds
Two numbers matter, and neither has a universal right answer.
fail_max, the number of failures before tripping, trades sensitivity against twitchiness. Set it to 1 and a single unlucky timeout opens the circuit for everyone. Set it to 50 and you send a lot of doomed traffic before you react. A count in the low single digits to low tens is the usual range.
A rolling window or a failure rate behaves better under mixed traffic than a raw consecutive count. For example, you might trip if more than half of the last 20 calls failed. The reason is that one success in the middle should not fully reset your read on a provider that is flapping, and with a consecutive count, it does.
cooldown, how long to stay open, trades recovery speed against pressure on the upstream. Too short, and half-open keeps poking a provider that has not recovered, which slows its recovery. Too long, and you stay degraded well after the outage is over. Thirty seconds to a couple of minutes is a reasonable starting point for an external API. The honest way to set it is to look at how long your provider’s incidents actually last, pick something on that order, and adjust from what you see.
The one thing I would not do is treat these as set-once constants. Emit the state transitions as metrics (the library hooks make this a few lines), and watch how often the breaker opens and how long it stays open. If it never trips, it is not protecting you, and your thresholds are too loose. If it is open more than it is closed, either your provider is genuinely bad or your thresholds are too tight. Both are worth knowing.
Where it bites you in production
The breaker is per-process. The CircuitBreaker object lives in one Python process. Run four replicas of your FastAPI service behind a load balancer, as you would on Kubernetes, and you have four independent breakers that learn about the outage separately. That is usually acceptable: each instance still protects itself, and finding out four times is not much worse than once. If you want a shared view, you need shared state (Redis, typically), and now the breaker check is a network call with its own failure modes. purgatory supports a Redis backend for exactly this. Most teams do not need it, so know which camp you are in before you add the dependency.
Half-open can turn into a thundering herd. The pattern allows a single probe. Under concurrent load, though, a naive implementation lets every request through the instant the cooldown expires, because they all read the state as half-open at once. You then slam the recovering provider with the full backlog. A correct breaker admits one trial and holds the rest. This is a real reason to use a maintained library, because the sketch above has this exact bug.
A breaker plus retries can multiply. These two patterns compose, but think about the order. Retry inside the breaker, and a single logical call can become several provider hits before it is finally counted as one failure. The cleaner arrangement is retries for the transient case, wrapped by the breaker for the sustained case, with the retries kept small: a couple of quick retries with backoff, and if the whole thing still fails, that is one failure the breaker records. Do not stack aggressive retries under a lenient breaker and expect either to help.
An open breaker can hide the outage from your dashboards. Failing fast is good for users and bad for whoever is on call, because the provider’s errors stop showing up as your errors. Make “breaker opened” a first-class, alertable event. Otherwise the first sign of trouble is a user asking why the assistant keeps saying it is unavailable.
What I would do differently
The first time I added a breaker to a service, I put it in the wrong place: around a function that did retrieval and the model call. When the vector store hiccuped, the breaker tripped and cut off the model too, even though the model was fine. A breaker should wrap exactly one dependency, so that tripping tells you something specific and the fallback can be specific too. If a call fans out to several dependencies, that means several breakers, not one around the group.
The other thing I would build in from the start is a real fallback, not just an error. An open breaker is an opportunity: it is exactly the moment to serve a cached response, a smaller or self-hosted model, or a plainly worded degraded answer. Decide that in advance, per endpoint, and an outage turns from a wall of 503s into something the product mostly rides through.
Closing
A circuit breaker does not make your LLM provider more reliable. It makes your service reliable in the face of a provider that is not, which is the only kind of reliability you actually control when the model lives behind someone else’s API. It pairs with the retry and backoff logic from the earlier post and with keeping the event loop unblocked so a slow dependency cannot stall everything: retries for the blip, the breaker for the outage, async so neither one holds a worker hostage. That combination is most of what keeps the answer path of a copilot like Archi responsive on a bad day for the upstream.