{"componentChunkName":"component---src-templates-blog-post-js","path":"/blog/2026-08-29-circuit-breaker-llm-api-calls/","result":{"data":{"site":{"siteMetadata":{"title":"M.Hassan Ahmed","author":"Hassan11196"}},"markdownRemark":{"id":"e799f4e1-c88c-58fc-90d3-aa3a9d528733","excerpt":"In an earlier post on rate limits, retries, and backoff, I mentioned almost in passing that one reason to write your own retry loop, instead of leaning on the…","html":"<p>In an earlier post on <a href=\"/blog/2026-07-11-llm-api-rate-limits-retries-backoff/\">rate limits, retries, and backoff</a>, I mentioned almost in passing that one reason to write your own retry loop, instead of leaning on the SDK, is that you might want a circuit breaker. This post is the part I skipped. Retries handle a request that failed for no lasting reason: a dropped connection, a single 503, a blip. They are the wrong tool when the provider is actually down.</p>\n<p>Here is the difference. A retry says “that call failed, try it again in a moment.” When the upstream is healthy, that usually works, because the failure was transient. During a real outage, though, every retry is another request thrown at a service that cannot answer, and the caller waits out its backoff while holding a worker. Multiply that by every in-flight request, and one slow dependency has turned into a fully stalled service. The <a href=\"https://sre.google/sre-book/addressing-cascading-failures/\">Google SRE book</a> has a whole chapter on how this cascades.</p>\n<p>A circuit breaker is the piece that says “stop trying for a bit.” This post is for engineers running a FastAPI or similar backend that calls an LLM API, who want the service to stay up, and stay honest with its own users, when the model endpoint does not. I build one, tune it, and spend most of the space on the ways it bites you in production. The running example is the kind of backend I worked on for <a href=\"/project/archi/\">Archi</a>, the retrieval copilot for CMS operations at CERN, where the answer path depends on an external model API that is not under our control.</p>\n<h2>The pattern: three states</h2>\n<p>The circuit breaker comes from Michael Nygard’s <em>Release It!</em> and was popularized by a much-cited <a href=\"https://martinfowler.com/bliki/CircuitBreaker.html\">write-up by Martin Fowler</a>. The name is an analogy. An electrical breaker trips to protect the wiring, and you flip it back once the fault is cleared. In software, the “wiring” is your own service, and the fault is a dependency that has stopped behaving.</p>\n<p>A breaker has three states, and the whole design is the movement between them.</p>\n<p><img src=\"/66967e989473ecf5e8c65fc1d200cd34/circuit-breaker-states.svg\" alt=\"A state diagram. CLOSED, where calls pass through and failures are counted, moves to OPEN when failures cross a threshold. OPEN fails calls fast, leaves the provider alone, and runs a cooldown timer, then moves to HALF-OPEN when the cooldown expires. HALF-OPEN lets one trial call through: if it succeeds the breaker returns to CLOSED, if it fails it goes back to OPEN and the cooldown restarts.\"></p>\n<p><strong>Closed</strong> is normal operation. Calls go through to the provider, and the breaker counts failures. If the failure count crosses a threshold inside some window, the breaker trips.</p>\n<p><strong>Open</strong> is the interesting state. The breaker rejects calls immediately, without touching the provider. This is the point of the whole thing: you fail fast instead of piling load onto something that is already struggling, and you give the provider room to recover. A cooldown timer runs while the breaker is open.</p>\n<p><strong>Half-open</strong> is how the breaker tests the water. When the cooldown expires, it lets a single trial call through. If that call succeeds, the breaker closes and traffic resumes. If it fails, the breaker snaps back to open and the cooldown restarts. That single probe matters: without it, you either reopen the floodgates all at once or stay open forever.</p>\n<h2>A minimal breaker in Python</h2>\n<p>You do not need a library to understand this. Here is an async breaker that wraps a call, small enough to read in one sitting. Watch how <code class=\"language-text\">call</code> moves from open to half-open once the cooldown has passed, how any success resets it to closed, and how <code class=\"language-text\">_on_failure</code> reopens it. It is deliberately not thread-safe yet; I come back to that later.</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\"><span class=\"token keyword\">import</span> time\n<span class=\"token keyword\">from</span> enum <span class=\"token keyword\">import</span> Enum\n\n\n<span class=\"token keyword\">class</span> <span class=\"token class-name\">State</span><span class=\"token punctuation\">(</span><span class=\"token builtin\">str</span><span class=\"token punctuation\">,</span> Enum<span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n    CLOSED <span class=\"token operator\">=</span> <span class=\"token string\">\"closed\"</span>\n    OPEN <span class=\"token operator\">=</span> <span class=\"token string\">\"open\"</span>\n    HALF_OPEN <span class=\"token operator\">=</span> <span class=\"token string\">\"half_open\"</span>\n\n\n<span class=\"token keyword\">class</span> <span class=\"token class-name\">CircuitOpenError</span><span class=\"token punctuation\">(</span>Exception<span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n    <span class=\"token triple-quoted-string string\">\"\"\"Raised instead of calling the provider while the breaker is open.\"\"\"</span>\n\n\n<span class=\"token keyword\">class</span> <span class=\"token class-name\">CircuitBreaker</span><span class=\"token punctuation\">:</span>\n    <span class=\"token keyword\">def</span> <span class=\"token function\">__init__</span><span class=\"token punctuation\">(</span>self<span class=\"token punctuation\">,</span> fail_max<span class=\"token punctuation\">:</span> <span class=\"token builtin\">int</span> <span class=\"token operator\">=</span> <span class=\"token number\">5</span><span class=\"token punctuation\">,</span> cooldown<span class=\"token punctuation\">:</span> <span class=\"token builtin\">float</span> <span class=\"token operator\">=</span> <span class=\"token number\">30.0</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n        self<span class=\"token punctuation\">.</span>fail_max <span class=\"token operator\">=</span> fail_max        <span class=\"token comment\"># failures before we trip</span>\n        self<span class=\"token punctuation\">.</span>cooldown <span class=\"token operator\">=</span> cooldown        <span class=\"token comment\"># seconds to stay open</span>\n        self<span class=\"token punctuation\">.</span>state <span class=\"token operator\">=</span> State<span class=\"token punctuation\">.</span>CLOSED\n        self<span class=\"token punctuation\">.</span>failures <span class=\"token operator\">=</span> <span class=\"token number\">0</span>\n        self<span class=\"token punctuation\">.</span>opened_at <span class=\"token operator\">=</span> <span class=\"token number\">0.0</span>\n\n    <span class=\"token keyword\">async</span> <span class=\"token keyword\">def</span> <span class=\"token function\">call</span><span class=\"token punctuation\">(</span>self<span class=\"token punctuation\">,</span> fn<span class=\"token punctuation\">,</span> <span class=\"token operator\">*</span>args<span class=\"token punctuation\">,</span> <span class=\"token operator\">**</span>kwargs<span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n        <span class=\"token keyword\">if</span> self<span class=\"token punctuation\">.</span>state <span class=\"token operator\">==</span> State<span class=\"token punctuation\">.</span>OPEN<span class=\"token punctuation\">:</span>\n            <span class=\"token keyword\">if</span> time<span class=\"token punctuation\">.</span>monotonic<span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span> <span class=\"token operator\">-</span> self<span class=\"token punctuation\">.</span>opened_at <span class=\"token operator\">>=</span> self<span class=\"token punctuation\">.</span>cooldown<span class=\"token punctuation\">:</span>\n                self<span class=\"token punctuation\">.</span>state <span class=\"token operator\">=</span> State<span class=\"token punctuation\">.</span>HALF_OPEN   <span class=\"token comment\"># time to probe</span>\n            <span class=\"token keyword\">else</span><span class=\"token punctuation\">:</span>\n                <span class=\"token keyword\">raise</span> CircuitOpenError<span class=\"token punctuation\">(</span><span class=\"token string\">\"circuit is open\"</span><span class=\"token punctuation\">)</span>\n\n        <span class=\"token keyword\">try</span><span class=\"token punctuation\">:</span>\n            result <span class=\"token operator\">=</span> <span class=\"token keyword\">await</span> fn<span class=\"token punctuation\">(</span><span class=\"token operator\">*</span>args<span class=\"token punctuation\">,</span> <span class=\"token operator\">**</span>kwargs<span class=\"token punctuation\">)</span>\n        <span class=\"token keyword\">except</span> Exception<span class=\"token punctuation\">:</span>\n            self<span class=\"token punctuation\">.</span>_on_failure<span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span>\n            <span class=\"token keyword\">raise</span>\n        <span class=\"token keyword\">else</span><span class=\"token punctuation\">:</span>\n            self<span class=\"token punctuation\">.</span>_on_success<span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span>\n            <span class=\"token keyword\">return</span> result\n\n    <span class=\"token keyword\">def</span> <span class=\"token function\">_on_success</span><span class=\"token punctuation\">(</span>self<span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n        self<span class=\"token punctuation\">.</span>failures <span class=\"token operator\">=</span> <span class=\"token number\">0</span>\n        self<span class=\"token punctuation\">.</span>state <span class=\"token operator\">=</span> State<span class=\"token punctuation\">.</span>CLOSED\n\n    <span class=\"token keyword\">def</span> <span class=\"token function\">_on_failure</span><span class=\"token punctuation\">(</span>self<span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n        self<span class=\"token punctuation\">.</span>failures <span class=\"token operator\">+=</span> <span class=\"token number\">1</span>\n        <span class=\"token keyword\">if</span> self<span class=\"token punctuation\">.</span>state <span class=\"token operator\">==</span> State<span class=\"token punctuation\">.</span>HALF_OPEN <span class=\"token keyword\">or</span> self<span class=\"token punctuation\">.</span>failures <span class=\"token operator\">>=</span> self<span class=\"token punctuation\">.</span>fail_max<span class=\"token punctuation\">:</span>\n            self<span class=\"token punctuation\">.</span>state <span class=\"token operator\">=</span> State<span class=\"token punctuation\">.</span>OPEN\n            self<span class=\"token punctuation\">.</span>opened_at <span class=\"token operator\">=</span> time<span class=\"token punctuation\">.</span>monotonic<span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span></code></pre></div>\n<p>Wrapping an LLM call is then just a matter of handing the request coroutine to <code class=\"language-text\">call</code>:</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\"><span class=\"token keyword\">import</span> httpx\n\nbreaker <span class=\"token operator\">=</span> CircuitBreaker<span class=\"token punctuation\">(</span>fail_max<span class=\"token operator\">=</span><span class=\"token number\">5</span><span class=\"token punctuation\">,</span> cooldown<span class=\"token operator\">=</span><span class=\"token number\">30.0</span><span class=\"token punctuation\">)</span>\n\n<span class=\"token keyword\">async</span> <span class=\"token keyword\">def</span> <span class=\"token function\">ask_model</span><span class=\"token punctuation\">(</span>client<span class=\"token punctuation\">:</span> httpx<span class=\"token punctuation\">.</span>AsyncClient<span class=\"token punctuation\">,</span> prompt<span class=\"token punctuation\">:</span> <span class=\"token builtin\">str</span><span class=\"token punctuation\">)</span> <span class=\"token operator\">-</span><span class=\"token operator\">></span> <span class=\"token builtin\">str</span><span class=\"token punctuation\">:</span>\n    <span class=\"token keyword\">async</span> <span class=\"token keyword\">def</span> <span class=\"token function\">_request</span><span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n        resp <span class=\"token operator\">=</span> <span class=\"token keyword\">await</span> client<span class=\"token punctuation\">.</span>post<span class=\"token punctuation\">(</span>\n            <span class=\"token string\">\"https://api.provider.example/v1/chat\"</span><span class=\"token punctuation\">,</span>\n            json<span class=\"token operator\">=</span><span class=\"token punctuation\">{</span><span class=\"token string\">\"model\"</span><span class=\"token punctuation\">:</span> <span class=\"token string\">\"some-model\"</span><span class=\"token punctuation\">,</span> <span class=\"token string\">\"messages\"</span><span class=\"token punctuation\">:</span> <span class=\"token punctuation\">[</span><span class=\"token punctuation\">{</span><span class=\"token string\">\"role\"</span><span class=\"token punctuation\">:</span> <span class=\"token string\">\"user\"</span><span class=\"token punctuation\">,</span> <span class=\"token string\">\"content\"</span><span class=\"token punctuation\">:</span> prompt<span class=\"token punctuation\">}</span><span class=\"token punctuation\">]</span><span class=\"token punctuation\">}</span><span class=\"token punctuation\">,</span>\n            timeout<span class=\"token operator\">=</span>httpx<span class=\"token punctuation\">.</span>Timeout<span class=\"token punctuation\">(</span><span class=\"token number\">30.0</span><span class=\"token punctuation\">,</span> connect<span class=\"token operator\">=</span><span class=\"token number\">5.0</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">,</span>\n        <span class=\"token punctuation\">)</span>\n        resp<span class=\"token punctuation\">.</span>raise_for_status<span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span>\n        <span class=\"token keyword\">return</span> resp<span class=\"token punctuation\">.</span>json<span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">[</span><span class=\"token string\">\"choices\"</span><span class=\"token punctuation\">]</span><span class=\"token punctuation\">[</span><span class=\"token number\">0</span><span class=\"token punctuation\">]</span><span class=\"token punctuation\">[</span><span class=\"token string\">\"message\"</span><span class=\"token punctuation\">]</span><span class=\"token punctuation\">[</span><span class=\"token string\">\"content\"</span><span class=\"token punctuation\">]</span>\n\n    <span class=\"token keyword\">return</span> <span class=\"token keyword\">await</span> breaker<span class=\"token punctuation\">.</span>call<span class=\"token punctuation\">(</span>_request<span class=\"token punctuation\">)</span></code></pre></div>\n<p>In a FastAPI handler, catching <code class=\"language-text\">CircuitOpenError</code> is where you decide what the user sees when the model is unreachable. It could be a cached answer, a plain “the assistant is temporarily unavailable,” or a 503 with a <code class=\"language-text\">Retry-After</code>. The decision is yours, and any of them is better than a request that hangs for 30 seconds and then returns a 500.</p>\n<p>For anything real, reach for a maintained library rather than the toy above. <a href=\"https://github.com/danielfm/pybreaker\"><code class=\"language-text\">pybreaker</code></a> is the well-worn choice and supports async. <a href=\"https://github.com/mardiros/purgatory\"><code class=\"language-text\">purgatory</code></a> is a newer async-first option with pluggable state storage. Both handle the concurrency and bookkeeping I glossed over, and both give you listener hooks for metrics.</p>\n<h2>Deciding what counts as a failure</h2>\n<p>This is the tuning decision people get wrong, and the toy code above gets it wrong on purpose so I can point at it: it treats <em>every</em> exception as a failure. <code class=\"language-text\">raise_for_status()</code> throws on any 4xx or 5xx, so a <code class=\"language-text\">400 Bad Request</code> caused by a malformed prompt in your own code will happily trip the breaker. Now one buggy request path takes down the model for every other user for thirty seconds. That is not the provider’s fault, and opening the circuit does nothing but hide your bug.</p>\n<p>The rule: a circuit breaker should trip on failures of the <em>dependency</em>, not on your own mistakes. Sort the responses by whose problem they are:</p>\n<ul>\n<li><strong>5xx, timeouts, connection errors</strong> are the provider’s problem. These should count.</li>\n<li><strong>429 Too Many Requests</strong> is arguable, and it depends. Throttling means the provider is telling you to slow down, which a breaker does, so counting it can be right. But 429 is really the job of <a href=\"/blog/2026-07-11-llm-api-rate-limits-retries-backoff/\">backoff and a rate limiter</a>, and leaning on the breaker for it is coarse.</li>\n<li><strong>4xx other than 429</strong> are your problem. A <code class=\"language-text\">400</code> or <code class=\"language-text\">422</code> means the request was wrong. Retrying or tripping on these is pointless, because the same request will fail the same way. Do not count them.</li>\n</ul>\n<p>So the breaker needs a way to classify errors instead of catching everything. A small predicate does it. It returns <code class=\"language-text\">True</code> for timeouts, connection errors, and 5xx responses, and <code class=\"language-text\">False</code> for everything else:</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\"><span class=\"token keyword\">def</span> <span class=\"token function\">is_dependency_failure</span><span class=\"token punctuation\">(</span>exc<span class=\"token punctuation\">:</span> Exception<span class=\"token punctuation\">)</span> <span class=\"token operator\">-</span><span class=\"token operator\">></span> <span class=\"token builtin\">bool</span><span class=\"token punctuation\">:</span>\n    <span class=\"token keyword\">if</span> <span class=\"token builtin\">isinstance</span><span class=\"token punctuation\">(</span>exc<span class=\"token punctuation\">,</span> <span class=\"token punctuation\">(</span>httpx<span class=\"token punctuation\">.</span>TimeoutException<span class=\"token punctuation\">,</span> httpx<span class=\"token punctuation\">.</span>ConnectError<span class=\"token punctuation\">)</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n        <span class=\"token keyword\">return</span> <span class=\"token boolean\">True</span>\n    <span class=\"token keyword\">if</span> <span class=\"token builtin\">isinstance</span><span class=\"token punctuation\">(</span>exc<span class=\"token punctuation\">,</span> httpx<span class=\"token punctuation\">.</span>HTTPStatusError<span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n        <span class=\"token keyword\">return</span> exc<span class=\"token punctuation\">.</span>response<span class=\"token punctuation\">.</span>status_code <span class=\"token operator\">>=</span> <span class=\"token number\">500</span>\n    <span class=\"token keyword\">return</span> <span class=\"token boolean\">False</span></code></pre></div>\n<p>Then <code class=\"language-text\">call</code> only counts failures that pass the predicate, and re-raises the rest without touching its counters. Every serious library gives you this hook under a name like <code class=\"language-text\">exclude</code> or a custom exception filter. Use it, because the default of “everything is a failure” is rarely what you want.</p>\n<h2>Tuning the thresholds</h2>\n<p>Two numbers matter, and neither has a universal right answer.</p>\n<p><strong><code class=\"language-text\">fail_max</code></strong>, the number of failures before tripping, trades sensitivity against twitchiness. Set it to 1 and a single unlucky timeout opens the circuit for everyone. Set it to 50 and you send a lot of doomed traffic before you react. A count in the low single digits to low tens is the usual range.</p>\n<p>A rolling window or a failure <em>rate</em> behaves better under mixed traffic than a raw consecutive count. For example, you might trip if more than half of the last 20 calls failed. The reason is that one success in the middle should not fully reset your read on a provider that is flapping, and with a consecutive count, it does.</p>\n<p><strong><code class=\"language-text\">cooldown</code></strong>, how long to stay open, trades recovery speed against pressure on the upstream. Too short, and half-open keeps poking a provider that has not recovered, which slows its recovery. Too long, and you stay degraded well after the outage is over. Thirty seconds to a couple of minutes is a reasonable starting point for an external API. The honest way to set it is to look at how long your provider’s incidents actually last, pick something on that order, and adjust from what you see.</p>\n<p>The one thing I would not do is treat these as set-once constants. Emit the state transitions as metrics (the library hooks make this a few lines), and watch how often the breaker opens and how long it stays open. If it never trips, it is not protecting you, and your thresholds are too loose. If it is open more than it is closed, either your provider is genuinely bad or your thresholds are too tight. Both are worth knowing.</p>\n<h2>Where it bites you in production</h2>\n<p><strong>The breaker is per-process.</strong> The <code class=\"language-text\">CircuitBreaker</code> object lives in one Python process. Run four replicas of your FastAPI service behind a load balancer, as you would on Kubernetes, and you have four independent breakers that learn about the outage separately. That is usually acceptable: each instance still protects itself, and finding out four times is not much worse than once. If you want a shared view, you need shared state (Redis, typically), and now the breaker check is a network call with its own failure modes. <code class=\"language-text\">purgatory</code> supports a Redis backend for exactly this. Most teams do not need it, so know which camp you are in before you add the dependency.</p>\n<p><strong>Half-open can turn into a thundering herd.</strong> The pattern allows a single probe. Under concurrent load, though, a naive implementation lets <em>every</em> request through the instant the cooldown expires, because they all read the state as half-open at once. You then slam the recovering provider with the full backlog. A correct breaker admits one trial and holds the rest. This is a real reason to use a maintained library, because the sketch above has this exact bug.</p>\n<p><strong>A breaker plus retries can multiply.</strong> These two patterns compose, but think about the order. Retry <em>inside</em> the breaker, and a single logical call can become several provider hits before it is finally counted as one failure. The cleaner arrangement is retries for the transient case, wrapped by the breaker for the sustained case, with the retries kept small: a couple of quick retries with backoff, and if the whole thing still fails, that is one failure the breaker records. Do not stack aggressive retries under a lenient breaker and expect either to help.</p>\n<p><strong>An open breaker can hide the outage from your dashboards.</strong> Failing fast is good for users and bad for whoever is on call, because the provider’s errors stop showing up as your errors. Make “breaker opened” a first-class, alertable event. Otherwise the first sign of trouble is a user asking why the assistant keeps saying it is unavailable.</p>\n<h2>What I would do differently</h2>\n<p>The first time I added a breaker to a service, I put it in the wrong place: around a function that did retrieval <em>and</em> the model call. When the vector store hiccuped, the breaker tripped and cut off the model too, even though the model was fine. A breaker should wrap exactly one dependency, so that tripping tells you something specific and the fallback can be specific too. If a call fans out to several dependencies, that means several breakers, not one around the group.</p>\n<p>The other thing I would build in from the start is a real fallback, not just an error. An open breaker is an opportunity: it is exactly the moment to serve a cached response, a smaller or self-hosted model, or a plainly worded degraded answer. Decide that in advance, per endpoint, and an outage turns from a wall of 503s into something the product mostly rides through.</p>\n<h2>Closing</h2>\n<p>A circuit breaker does not make your LLM provider more reliable. It makes <em>your</em> service reliable in the face of a provider that is not, which is the only kind of reliability you actually control when the model lives behind someone else’s API. It pairs with the <a href=\"/blog/2026-07-11-llm-api-rate-limits-retries-backoff/\">retry and backoff logic</a> from the earlier post and with keeping the <a href=\"/blog/2026-07-10-fastapi-blocking-event-loop/\">event loop unblocked</a> so a slow dependency cannot stall everything: retries for the blip, the breaker for the outage, async so neither one holds a worker hostage. That combination is most of what keeps the answer path of a copilot like <a href=\"/project/archi/\">Archi</a> responsive on a bad day for the upstream.</p>","frontmatter":{"title":"Circuit Breakers for LLM API Calls","date":"2026-08-29T00:00:00.000Z","description":"When an LLM provider degrades, retries make it worse. A practical guide to adding a circuit breaker in Python: the three states, tuning, and failure modes.","thumbnail":{"childImageSharp":{"fluid":{"base64":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAABQAAAALCAIAAADwazoUAAAACXBIWXMAAAsSAAALEgHS3X78AAABwklEQVQoz12Qa2+bMBSG+VA1gDE2NxuMsWkIxJARwnKBNd32odqXqept+/+/ZYYom1bpkXWO/T7Hso1cPmikOJ+Ob+Pw67B/Oeyfh9Pb8fDa756G03u3/TlvvtyPv3WtkxdFY5iIXSg2g+ofyk/juruvd+eyHXO1r/vzajM0/df19kt7+L6sjwuX/VUMx5OOJxxfAk/YKLOxuBRBVHO+R345tTizULaATJ9Cf85PSAPg7ArXIRMLE3GARc7Hdv2DRBsb85Cobf2Uys4jpTWd6jsmxZin8guUFDJZ0UAAnN6AkNKuXT87WPhRVeTfAloGZAlgAlEKtY+4geN1Vg3LzTnItllSNrIuM4WDvFw+9vW71qCXW4jdOAS5iWL5rmp3RSNDLadaVrwaZX0Os86LK5I1IVdx2qrVo4Vi22VBpNxAwuDOx1zL/WrT3VUiTIGbzm+ev0SvFkwWgNyCyIacUAX9zHRiE1I9Ql/uoCT0Uh+zEDPPY5abGBaklkMtOIV0j0iJk4YXp0I9YKr8pNHTpwyMNYsr5twal90rFEdVJHZx/pkt91T0AWsBSi1IZv8jhumQ/4kWILq1Q40uTM3HwD/+AE3HRJtRV0csAAAAAElFTkSuQmCC","aspectRatio":1.899441340782123,"src":"/static/0df538f83ed1b79f46a4c41b7c7bb925/40a76/hero.png","srcSet":"/static/0df538f83ed1b79f46a4c41b7c7bb925/c972b/hero.png 340w,\n/static/0df538f83ed1b79f46a4c41b7c7bb925/27625/hero.png 680w,\n/static/0df538f83ed1b79f46a4c41b7c7bb925/40a76/hero.png 1360w,\n/static/0df538f83ed1b79f46a4c41b7c7bb925/ed396/hero.png 2000w","sizes":"(max-width: 1360px) 100vw, 1360px"}}}}}},"pageContext":{"slug":"/2026-08-29-circuit-breaker-llm-api-calls/","previous":"blog/2026-08-30-prompt-caching-claude-api/","next":"blog/2026-09-01-metadata-filtering-rag-vector-search/"}},"staticQueryHashes":["32046230"]}