{"componentChunkName":"component---src-templates-blog-post-js","path":"/blog/2026-07-11-llm-api-rate-limits-retries-backoff/","result":{"data":{"site":{"siteMetadata":{"title":"M.Hassan Ahmed","author":"Hassan11196"}},"markdownRemark":{"id":"32ce6b7d-30eb-52d4-aba7-70e05e325d44","excerpt":"The first time a rate limit bites, it never looks like a rate limit. A demo that ran fine all week suddenly throws  in the middle of a live session. An agent…","html":"<p>The first time a rate limit bites, it never looks like a rate limit. A demo that ran fine all week suddenly throws <code class=\"language-text\">429</code> in the middle of a live session. An agent that was happily calling tools stalls halfway through a task. A batch of documents you are indexing gets through the first forty, and then every call comes back rejected. Nothing in your code changed. What changed is that more requests arrived at once than the provider is willing to serve you in a minute.</p>\n<p>This post is for engineers running an LLM-backed service, especially an <a href=\"https://fastapi.tiangolo.com/async/\">async FastAPI</a> one, who have hit that wall or want to stop hitting it. I’ve run into it in the backends behind <a href=\"/project/cloud-canvas-ai/\">CloudCanvasAI</a>, <a href=\"/project/archi/\">Archi</a>, and <a href=\"/project/gemini-alchemy/\">Gemini Alchemy</a>. In all three, a single user request fans out into many model calls, and a burst of those calls trips the provider’s per-minute budget.</p>\n<p>The fix is not one thing. It is a retry policy that knows what to retry, a backoff that does not make the problem worse, and a way to pace your own traffic so you rarely reach the limit at all. The sections below take those in order.</p>\n<p>\n  <a\n    class=\"gatsby-resp-image-link\"\n    href=\"/static/22b16db611c0d1406928cbdd55d9ba22/60356/rate-limit-429s.png\"\n    style=\"display: block\"\n    target=\"_blank\"\n    rel=\"noopener\"\n  >\n  \n  <span\n    class=\"gatsby-resp-image-wrapper\"\n    style=\"position: relative; display: block; margin: 7vw 0; max-width: 1360px; margin-left: auto; margin-right: auto;\"\n  >\n    <span\n      class=\"gatsby-resp-image-background-image\"\n      style=\"padding-bottom: 48.529411764705884%; position: relative; bottom: 0; left: 0; background-image: url('data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAABQAAAAKCAIAAAA7N+mxAAAACXBIWXMAAAsSAAALEgHS3X78AAABkklEQVQoz2WRy27bMBBFtQ5qq+L7pSdJkRLlyEmcODZidNEgiwJd5P9/pkO1iwAFDgbcHNw7wwLzhqmeyp5rC489Uhms99iU2HwnTYa2VabLsB5AbEDMFm56PD7flofrerql9YJ5T5Uj0hIJ01Ws++fQftOGvxriDihKYkpSQw4ACURYgCkvdKAgf9W43cga5h4oMGsJ76i0FanLSkPJO2yCHt/t/Z52WFiQIQpzaLFpYsR8zFOMBbWzWJ9hJdJG+XRBYvgmbKrjp3soTRBdQsIxMxt7dssPm27woDqBT0QoKtGTLsrTBSmH20i030k3mfjLHnfKExWwHlm9NOHtdPs8Xn7X45WZA+GByAnkQTy+ErfgJsr1hYe1lD616aM77Djs7xFzTCdtz814rf1VDy9ZFpHIuYCeWDgwxXqGhlSPyMS5mX+2acfyYRDcJm8YMA8wIRCgcqYqFXAPzAYhPeNOCK9VYCp4Pb3WsG3gKjI1bTnxi5ZNqpZi+7cBjgkNAami1FHpSWZmQJlE/9MApg9/AKAPRcqfraQwAAAAAElFTkSuQmCC'); background-size: cover; display: block;\"\n    >\n      <picture>\n        <source\n          srcset=\"/static/22b16db611c0d1406928cbdd55d9ba22/51a8e/rate-limit-429s.webp 340w,\n/static/22b16db611c0d1406928cbdd55d9ba22/713b7/rate-limit-429s.webp 680w,\n/static/22b16db611c0d1406928cbdd55d9ba22/54376/rate-limit-429s.webp 1360w\"\n          sizes=\"(max-width: 1360px) 100vw, 1360px\"\n          type=\"image/webp\"\n        />\n        <source\n          srcset=\"/static/22b16db611c0d1406928cbdd55d9ba22/ad208/rate-limit-429s.png 340w,\n/static/22b16db611c0d1406928cbdd55d9ba22/a5a26/rate-limit-429s.png 680w,\n/static/22b16db611c0d1406928cbdd55d9ba22/60356/rate-limit-429s.png 1360w\"\n          sizes=\"(max-width: 1360px) 100vw, 1360px\"\n          type=\"image/png\"\n        />\n        <img\n          class=\"gatsby-resp-image-image\"\n          style=\"width: 100%; height: 100%; margin: 0; vertical-align: middle; position: absolute; top: 0; left: 0; box-shadow: inset 0px 0px 0px 400px white;\"\n          src=\"/static/22b16db611c0d1406928cbdd55d9ba22/60356/rate-limit-429s.png\"\n          alt=\"Requests from an agent and a RAG fan-out hit a wall labelled RPM and TPM; some pass through as 200 OK while others bounce back as 429 and loop around after a delay, on their way to the Claude API which bills tokens on every accepted request including retries\"\n          title=\"\"\n          src=\"/static/22b16db611c0d1406928cbdd55d9ba22/60356/rate-limit-429s.png\"\n        />\n      </picture>\n      </span>\n  </span>\n  \n  </a>\n    </p>\n<h2>Two limits, not one: requests and tokens</h2>\n<p>Most people picture a rate limit as “requests per minute” and stop there. LLM providers enforce at least two limits at once, and the second one is the one that surprises you.</p>\n<ul>\n<li><strong>Requests per minute (RPM)</strong> is the obvious one. Too many calls in a window, and new ones are rejected.</li>\n<li><strong>Tokens per minute (TPM)</strong> counts the tokens flowing through, both input and output. Anthropic splits this further into input tokens per minute and output tokens per minute, and also enforces a daily ceiling. The <a href=\"https://platform.claude.com/docs/en/api/rate-limits\">Anthropic rate limits reference</a> has the exact tiers; <a href=\"https://platform.openai.com/docs/guides/rate-limits\">OpenAI’s rate limit guide</a> describes the same RPM-plus-TPM shape.</li>\n</ul>\n<p>The token limit is why counting requests alone misleads you. A request counter sees a small chat turn and a call that stuffs a 100k-token document into a long-context model as the same thing: one request each. But the second one can blow your token budget on its own. If you pace only by request count, you will still get <code class=\"language-text\">429</code>s from the token side, and you will wonder why, because your dashboard shows you nowhere near the request ceiling.</p>\n<p>Both limits report back the same way: an HTTP <code class=\"language-text\">429</code> with headers telling you where you stand. The response carries <code class=\"language-text\">retry-after</code> (how many seconds to wait) and a set of <code class=\"language-text\">x-ratelimit-*</code> headers showing your limit and how much of it is left. Those headers are not decoration. They are the provider telling you exactly how long to back off, and the single most common mistake is ignoring them.</p>\n<h2>Retry only what can succeed on a second try</h2>\n<p>Before writing any backoff code, sort the failures. Retrying blindly is how a slow minute becomes an outage.</p>\n<p><img src=\"/10186e254c7300e74cb1f561b240de8e/retry-flow.svg\" alt=\"A decision diagram: an API response splits three ways. A 2xx success is used and the backoff resets. A 400, 401, 403, or 404 is a permanent client error that should not be retried because a second attempt fails the same way. A 429, 500, 503, 529, timeout, or connection error is retryable: check for a Retry-After header and sleep exactly that long if present, otherwise sleep a full-jitter exponential delay, capped, up to a maximum attempt count. When attempts run out, stop and surface the error rather than looping forever\"></p>\n<p>The status code decides which side of the line a failure falls on:</p>\n<ul>\n<li><strong><code class=\"language-text\">400</code>, <code class=\"language-text\">401</code>, <code class=\"language-text\">403</code>, <code class=\"language-text\">404</code>: do not retry.</strong> These are your fault, not a timing problem: a malformed request, a bad key, a missing permission, or a wrong model ID. A retry sends the identical broken request and gets the identical rejection. Fail fast and surface it.</li>\n<li><strong><code class=\"language-text\">429</code>, <code class=\"language-text\">500</code>, <code class=\"language-text\">502</code>, <code class=\"language-text\">503</code>, <code class=\"language-text\">504</code>, <code class=\"language-text\">529</code>, plus timeouts and dropped connections: retry.</strong> The <code class=\"language-text\">429</code> is a rate limit. <code class=\"language-text\">500</code> and <code class=\"language-text\">502</code>/<code class=\"language-text\">503</code>/<code class=\"language-text\">504</code> are transient server faults. <code class=\"language-text\">529</code> is Anthropic’s “overloaded” signal, which means the service is busy rather than your account being over budget. All of these can succeed if you wait and try again.</li>\n</ul>\n<p>The <a href=\"https://platform.claude.com/docs/en/api/errors\">Anthropic error reference</a> marks which codes are retryable, and the mapping is close enough across providers that you can write one classifier. In the Anthropic Python SDK, the exception types make this clean. <code class=\"language-text\">anthropic.RateLimitError</code> (429), <code class=\"language-text\">anthropic.InternalServerError</code> (5xx), and <code class=\"language-text\">anthropic.APIConnectionError</code> (the network dropped before a response) are the retryable branch. An <code class=\"language-text\">anthropic.BadRequestError</code> or <code class=\"language-text\">anthropic.NotFoundError</code> is not.</p>\n<h2>Back off exponentially, and add jitter</h2>\n<p>When a call is retryable and there is no <code class=\"language-text\">retry-after</code> to obey, you wait before trying again, and you wait longer each time. That is exponential backoff: one second, then two, then four. Most people know that part. The part they skip is jitter (randomness in the delay), and skipping it turns a recoverable blip into a self-inflicted stampede.</p>\n<p>Think about what caused the <code class=\"language-text\">429</code> in the first place: a burst of requests all landed together. If every one of them backs off by exactly the same formula, they all wake up at the same instant and hit the API together again. You have rebuilt the original burst, just shifted a few seconds later. AWS wrote the reference piece on this, <a href=\"https://aws.amazon.com/blogs/architecture/exponential-backoff-and-jitter/\">Exponential Backoff And Jitter</a>, and its conclusion is blunt: add randomness to the delay so the retries spread out instead of synchronizing.</p>\n<p>The version I reach for is full jitter. Instead of sleeping the full exponential delay, sleep a random amount between zero and that delay:</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\"><span class=\"token keyword\">import</span> random\n\nBASE <span class=\"token operator\">=</span> <span class=\"token number\">1.0</span>     <span class=\"token comment\"># seconds</span>\nCAP <span class=\"token operator\">=</span> <span class=\"token number\">60.0</span>     <span class=\"token comment\"># never wait longer than this between tries</span>\n\n<span class=\"token keyword\">def</span> <span class=\"token function\">backoff_seconds</span><span class=\"token punctuation\">(</span>attempt<span class=\"token punctuation\">:</span> <span class=\"token builtin\">int</span><span class=\"token punctuation\">,</span> retry_after<span class=\"token punctuation\">:</span> <span class=\"token builtin\">float</span> <span class=\"token operator\">|</span> <span class=\"token boolean\">None</span> <span class=\"token operator\">=</span> <span class=\"token boolean\">None</span><span class=\"token punctuation\">)</span> <span class=\"token operator\">-</span><span class=\"token operator\">></span> <span class=\"token builtin\">float</span><span class=\"token punctuation\">:</span>\n    <span class=\"token comment\"># The server told us how long to wait. Obey it exactly.</span>\n    <span class=\"token keyword\">if</span> retry_after <span class=\"token keyword\">is</span> <span class=\"token keyword\">not</span> <span class=\"token boolean\">None</span><span class=\"token punctuation\">:</span>\n        <span class=\"token keyword\">return</span> retry_after\n    <span class=\"token comment\"># Otherwise: full jitter over an exponentially growing window.</span>\n    window <span class=\"token operator\">=</span> <span class=\"token builtin\">min</span><span class=\"token punctuation\">(</span>CAP<span class=\"token punctuation\">,</span> BASE <span class=\"token operator\">*</span> <span class=\"token number\">2</span> <span class=\"token operator\">**</span> attempt<span class=\"token punctuation\">)</span>\n    <span class=\"token keyword\">return</span> random<span class=\"token punctuation\">.</span>uniform<span class=\"token punctuation\">(</span><span class=\"token number\">0</span><span class=\"token punctuation\">,</span> window<span class=\"token punctuation\">)</span></code></pre></div>\n<p>Two rules do the work here:</p>\n<ul>\n<li><strong>If <code class=\"language-text\">retry-after</code> is present, it wins outright.</strong> The provider has told you the exact wait. Guessing a longer exponential delay wastes time, and guessing a shorter one earns another <code class=\"language-text\">429</code>. The <a href=\"https://developer.mozilla.org/en-US/docs/Web/HTTP/Headers/Retry-After\"><code class=\"language-text\">Retry-After</code> header</a> exists precisely so you do not have to guess.</li>\n<li><strong>Only when it is absent do you fall back to jittered exponential growth.</strong> The growth is capped, so a run of failures cannot schedule a ten-minute sleep.</li>\n</ul>\n<p>Next, wrap that function around a call. The loop below adds error classification and a hard attempt ceiling, and it pulls <code class=\"language-text\">retry-after</code> from the error’s response when there is one:</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\"><span class=\"token keyword\">import</span> asyncio\n<span class=\"token keyword\">import</span> anthropic\n\nRETRYABLE <span class=\"token operator\">=</span> <span class=\"token punctuation\">(</span>\n    anthropic<span class=\"token punctuation\">.</span>RateLimitError<span class=\"token punctuation\">,</span>        <span class=\"token comment\"># 429</span>\n    anthropic<span class=\"token punctuation\">.</span>InternalServerError<span class=\"token punctuation\">,</span>   <span class=\"token comment\"># 500, 502, 503, 504, 529</span>\n    anthropic<span class=\"token punctuation\">.</span>APIConnectionError<span class=\"token punctuation\">,</span>    <span class=\"token comment\"># network dropped before a response</span>\n<span class=\"token punctuation\">)</span>\n\n<span class=\"token keyword\">async</span> <span class=\"token keyword\">def</span> <span class=\"token function\">call_with_retry</span><span class=\"token punctuation\">(</span>make_call<span class=\"token punctuation\">,</span> <span class=\"token operator\">*</span><span class=\"token punctuation\">,</span> max_attempts<span class=\"token punctuation\">:</span> <span class=\"token builtin\">int</span> <span class=\"token operator\">=</span> <span class=\"token number\">6</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n    <span class=\"token keyword\">for</span> attempt <span class=\"token keyword\">in</span> <span class=\"token builtin\">range</span><span class=\"token punctuation\">(</span>max_attempts<span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n        <span class=\"token keyword\">try</span><span class=\"token punctuation\">:</span>\n            <span class=\"token keyword\">return</span> <span class=\"token keyword\">await</span> make_call<span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span>\n        <span class=\"token keyword\">except</span> RETRYABLE <span class=\"token keyword\">as</span> err<span class=\"token punctuation\">:</span>\n            <span class=\"token keyword\">if</span> attempt <span class=\"token operator\">==</span> max_attempts <span class=\"token operator\">-</span> <span class=\"token number\">1</span><span class=\"token punctuation\">:</span>\n                <span class=\"token keyword\">raise</span>  <span class=\"token comment\"># out of tries; let the caller decide what to do</span>\n            retry_after <span class=\"token operator\">=</span> <span class=\"token boolean\">None</span>\n            resp <span class=\"token operator\">=</span> <span class=\"token builtin\">getattr</span><span class=\"token punctuation\">(</span>err<span class=\"token punctuation\">,</span> <span class=\"token string\">\"response\"</span><span class=\"token punctuation\">,</span> <span class=\"token boolean\">None</span><span class=\"token punctuation\">)</span>\n            <span class=\"token keyword\">if</span> resp <span class=\"token keyword\">is</span> <span class=\"token keyword\">not</span> <span class=\"token boolean\">None</span><span class=\"token punctuation\">:</span>\n                header <span class=\"token operator\">=</span> resp<span class=\"token punctuation\">.</span>headers<span class=\"token punctuation\">.</span>get<span class=\"token punctuation\">(</span><span class=\"token string\">\"retry-after\"</span><span class=\"token punctuation\">)</span>\n                retry_after <span class=\"token operator\">=</span> <span class=\"token builtin\">float</span><span class=\"token punctuation\">(</span>header<span class=\"token punctuation\">)</span> <span class=\"token keyword\">if</span> header <span class=\"token keyword\">else</span> <span class=\"token boolean\">None</span>\n            <span class=\"token keyword\">await</span> asyncio<span class=\"token punctuation\">.</span>sleep<span class=\"token punctuation\">(</span>backoff_seconds<span class=\"token punctuation\">(</span>attempt<span class=\"token punctuation\">,</span> retry_after<span class=\"token punctuation\">)</span><span class=\"token punctuation\">)</span></code></pre></div>\n<p>The ceiling matters as much as the backoff. A retry loop with no limit does not heal an overloaded endpoint. It keeps throwing load at the one service that is already past its budget. Meanwhile, every request waiting on that loop holds a connection, a task, and a slice of your event loop. Bound the attempts, and bound the total wall-clock time too if the call sits on a request path where a user is waiting.</p>\n<h2>Let the SDK do the boring part</h2>\n<p>You rarely need to write that loop by hand, because the official SDKs already do it. The Anthropic SDK retries <code class=\"language-text\">408</code>, <code class=\"language-text\">409</code>, <code class=\"language-text\">429</code>, and <code class=\"language-text\">5xx</code> responses plus connection errors, with exponential backoff, and it reads <code class=\"language-text\">retry-after</code> for you. The only catch is the default: <strong>two</strong> retries. The snippet below raises that budget for a whole client, or for a single call:</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\"><span class=\"token keyword\">from</span> anthropic <span class=\"token keyword\">import</span> Anthropic\n\n<span class=\"token comment\"># Raise the retry budget for the whole client.</span>\nclient <span class=\"token operator\">=</span> Anthropic<span class=\"token punctuation\">(</span>max_retries<span class=\"token operator\">=</span><span class=\"token number\">5</span><span class=\"token punctuation\">)</span>\n\n<span class=\"token comment\"># Or override per request, e.g. on a call you know is bursty.</span>\nclient<span class=\"token punctuation\">.</span>with_options<span class=\"token punctuation\">(</span>max_retries<span class=\"token operator\">=</span><span class=\"token number\">8</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">.</span>messages<span class=\"token punctuation\">.</span>create<span class=\"token punctuation\">(</span>\n    model<span class=\"token operator\">=</span><span class=\"token string\">\"claude-opus-4-8\"</span><span class=\"token punctuation\">,</span>\n    max_tokens<span class=\"token operator\">=</span><span class=\"token number\">1024</span><span class=\"token punctuation\">,</span>\n    messages<span class=\"token operator\">=</span><span class=\"token punctuation\">[</span><span class=\"token punctuation\">{</span><span class=\"token string\">\"role\"</span><span class=\"token punctuation\">:</span> <span class=\"token string\">\"user\"</span><span class=\"token punctuation\">,</span> <span class=\"token string\">\"content\"</span><span class=\"token punctuation\">:</span> <span class=\"token string\">\"...\"</span><span class=\"token punctuation\">}</span><span class=\"token punctuation\">]</span><span class=\"token punctuation\">,</span>\n<span class=\"token punctuation\">)</span></code></pre></div>\n<p>Two retries is fine for a quiet service and thin under real fan-out, so set it deliberately rather than inheriting the default. Watch the timeout as well. The SDK’s client timeout defaults to ten minutes, and timeouts are themselves retried. So the worst-case wall-clock time for a call is roughly the timeout times the retry count, which is worth knowing before you set <code class=\"language-text\">max_retries</code> to something large on a user-facing path.</p>\n<p>The same knob exists elsewhere. If you call the API over raw <a href=\"https://www.python-httpx.org/\"><code class=\"language-text\">httpx</code></a> instead of the SDK, <a href=\"https://tenacity.readthedocs.io/\"><code class=\"language-text\">tenacity</code></a> gives you the same retry-with-jittered-backoff behavior as a decorator. OpenAI’s SDK exposes the identical <code class=\"language-text\">max_retries</code> setting.</p>\n<p>Reaching for the SDK’s retry is the right default. Write your own loop only when you need behavior it does not give you: a shared budget across many concurrent calls, a circuit breaker, or logic that switches to a fallback model after N failures instead of retrying the same one.</p>\n<h2>Pace your own traffic: the cheapest 429 is the one you never send</h2>\n<p>Retries are damage control. The better lever is pacing your own traffic so you stay under the limit on purpose. That matters most exactly where I hit rate limits: fan-out, where one job launches many model calls. Think of indexing a document set, enriching every chunk with a model-generated summary, or running an <a href=\"/blog/2026-06-30-llm-agent-tool-loop/\">agent</a> that issues many tool calls. The naive version launches all of them at once and lets the provider sort it out. The provider sorts it out with a wave of <code class=\"language-text\">429</code>s.</p>\n<p><img src=\"/66924860f5c6bd2771922fd08b7ea3ce/concurrency-gate.svg\" alt=\"Two rows. Without a gate, 200 chunks fire at once, trip the limit, and come back as a wave of 429s that turns into a retry storm and wasted latency. With a concurrency and token gate, the same 200 chunks are queued through a semaphore of size k plus a per-minute token budget, leave as a steady trickle, and reach the LLM API while staying under RPM and TPM\"></p>\n<p>A bounded concurrency gate is the smallest thing that helps. A semaphore caps how many requests are in flight at once, so the burst becomes a steady stream. In the code below, each chunk waits for a free slot in the gate, then makes its call through <code class=\"language-text\">call_with_retry</code>:</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\"><span class=\"token keyword\">import</span> asyncio\n\n<span class=\"token comment\"># Tune this to your requests-per-minute budget, not to your CPU count.</span>\ngate <span class=\"token operator\">=</span> asyncio<span class=\"token punctuation\">.</span>Semaphore<span class=\"token punctuation\">(</span><span class=\"token number\">8</span><span class=\"token punctuation\">)</span>\n\n<span class=\"token keyword\">async</span> <span class=\"token keyword\">def</span> <span class=\"token function\">summarize_chunk</span><span class=\"token punctuation\">(</span>client<span class=\"token punctuation\">,</span> text<span class=\"token punctuation\">:</span> <span class=\"token builtin\">str</span><span class=\"token punctuation\">)</span> <span class=\"token operator\">-</span><span class=\"token operator\">></span> <span class=\"token builtin\">str</span><span class=\"token punctuation\">:</span>\n    <span class=\"token keyword\">async</span> <span class=\"token keyword\">with</span> gate<span class=\"token punctuation\">:</span>  <span class=\"token comment\"># at most 8 model calls in flight at any moment</span>\n        resp <span class=\"token operator\">=</span> <span class=\"token keyword\">await</span> call_with_retry<span class=\"token punctuation\">(</span><span class=\"token keyword\">lambda</span><span class=\"token punctuation\">:</span> client<span class=\"token punctuation\">.</span>messages<span class=\"token punctuation\">.</span>create<span class=\"token punctuation\">(</span>\n            model<span class=\"token operator\">=</span><span class=\"token string\">\"claude-opus-4-8\"</span><span class=\"token punctuation\">,</span>\n            max_tokens<span class=\"token operator\">=</span><span class=\"token number\">256</span><span class=\"token punctuation\">,</span>\n            messages<span class=\"token operator\">=</span><span class=\"token punctuation\">[</span><span class=\"token punctuation\">{</span><span class=\"token string\">\"role\"</span><span class=\"token punctuation\">:</span> <span class=\"token string\">\"user\"</span><span class=\"token punctuation\">,</span> <span class=\"token string\">\"content\"</span><span class=\"token punctuation\">:</span> <span class=\"token string-interpolation\"><span class=\"token string\">f\"Summarize:\\n</span><span class=\"token interpolation\"><span class=\"token punctuation\">{</span>text<span class=\"token punctuation\">}</span></span><span class=\"token string\">\"</span></span><span class=\"token punctuation\">}</span><span class=\"token punctuation\">]</span><span class=\"token punctuation\">,</span>\n        <span class=\"token punctuation\">)</span><span class=\"token punctuation\">)</span>\n        <span class=\"token keyword\">return</span> resp<span class=\"token punctuation\">.</span>content<span class=\"token punctuation\">[</span><span class=\"token number\">0</span><span class=\"token punctuation\">]</span><span class=\"token punctuation\">.</span>text\n\n<span class=\"token keyword\">async</span> <span class=\"token keyword\">def</span> <span class=\"token function\">summarize_all</span><span class=\"token punctuation\">(</span>client<span class=\"token punctuation\">,</span> chunks<span class=\"token punctuation\">:</span> <span class=\"token builtin\">list</span><span class=\"token punctuation\">[</span><span class=\"token builtin\">str</span><span class=\"token punctuation\">]</span><span class=\"token punctuation\">)</span> <span class=\"token operator\">-</span><span class=\"token operator\">></span> <span class=\"token builtin\">list</span><span class=\"token punctuation\">[</span><span class=\"token builtin\">str</span><span class=\"token punctuation\">]</span><span class=\"token punctuation\">:</span>\n    <span class=\"token keyword\">return</span> <span class=\"token keyword\">await</span> asyncio<span class=\"token punctuation\">.</span>gather<span class=\"token punctuation\">(</span><span class=\"token operator\">*</span><span class=\"token punctuation\">(</span>summarize_chunk<span class=\"token punctuation\">(</span>client<span class=\"token punctuation\">,</span> c<span class=\"token punctuation\">)</span> <span class=\"token keyword\">for</span> c <span class=\"token keyword\">in</span> chunks<span class=\"token punctuation\">)</span><span class=\"token punctuation\">)</span></code></pre></div>\n<p>Pick the semaphore size from your request budget, not from how many tasks you happen to have. And remember the token side. Eight concurrent calls sounds safe until each one carries a large prompt, and then you are under the request limit and over the token limit.</p>\n<p>For that, the honest fix is a per-minute token allowance on top of the concurrency cap. A small async rate limiter such as <a href=\"https://aiolimiter.readthedocs.io/\"><code class=\"language-text\">aiolimiter</code></a> gives you the request-per-window bucket, and you track token spend yourself against the TPM figure from the rate-limit headers. Staying comfortably under the ceiling costs you a little throughput. It buys you a service that does not spend half its time retrying.</p>\n<h2>Failure modes that bite in production</h2>\n<p><strong>Ignoring <code class=\"language-text\">retry-after</code>.</strong> The provider hands you the exact wait, and your code guesses a different number. Guess too short and you earn another <code class=\"language-text\">429</code> immediately; guess too long and you sit idle. Read the header first, always.</p>\n<p><strong>No jitter.</strong> Synchronized retries reproduce the original burst. Jitter is the single change that separates a retry policy that recovers from one that oscillates.</p>\n<p><strong>Retrying the non-retryable.</strong> A <code class=\"language-text\">400</code> for a malformed request or a <code class=\"language-text\">401</code> for a bad key will fail identically on every attempt. All you have added is latency and, on a paid endpoint, cost. Classify before you loop.</p>\n<p><strong>Expecting a retry to save a broken stream.</strong> Once you are <a href=\"/blog/2026-06-30-fastapi-sse-streaming-llm/\">streaming a response over SSE</a>, an automatic retry cannot silently paper over a failure mid-stream, and the tokens already generated were already billed. Rate limits are checked when the request starts. So the durable answer for streaming under load is to keep concurrency low enough that streams start cleanly, and to fall back to another model on persistent <code class=\"language-text\">529</code> overload rather than hammering the busy one.</p>\n<p><strong>Paying for retries.</strong> Every accepted request bills its tokens. A retry that succeeds on the third try also billed the two failed attempts, if they got far enough to generate output. On a tight budget, aggressive retries quietly inflate spend. For work that is not latency-sensitive, the <a href=\"https://platform.claude.com/docs/en/build-with-claude/batch-processing\">Message Batches API</a> sidesteps the whole problem: you submit the jobs, the provider schedules them under its own limits at half the price, and you stop managing <code class=\"language-text\">429</code>s by hand.</p>\n<h2>What I would do differently</h2>\n<p>Early on, I treated rate limits as something to survive at the edge: wrap calls in a retry decorator and move on. That handles the symptom. What actually stabilized the fan-out paths in CloudCanvasAI was moving the control one level up, to a bounded gate in front of the model calls. Now the service shapes its own traffic instead of discovering the limit by bouncing off it. Retries still sit underneath as the safety net, but they fire rarely, because the gate keeps me under the ceiling to begin with.</p>\n<p>The other habit worth forming early is watching the <code class=\"language-text\">x-ratelimit-*</code> headers as a live signal, rather than waiting for the <code class=\"language-text\">429</code>. They tell you how much budget is left this minute. If you log the remaining-tokens header and see it trending toward zero, you can slow down before the wall instead of after it. That turns rate limiting from an exception you catch into a number you steer by. It is the same instinct behind keeping the <a href=\"/blog/2026-07-10-fastapi-blocking-event-loop/\">event loop unblocked</a>: the smooth-looking service is the one that never lets a queue build up in the first place.</p>\n<h2>Closing</h2>\n<p>Rate limits are not an edge case to catch and forget. They are a budget the provider hands you every minute, in two dimensions, and a well-behaved client spends that budget on purpose. Obey <code class=\"language-text\">retry-after</code> when it is there. Add jitter so your retries spread out instead of firing in unison. Retry only what a second attempt can fix, cap the attempts, and put a bounded gate in front of your fan-out so the retry path stays the exception rather than the norm.</p>\n<p>I have shipped every one of these behind the model calls in <a href=\"/project/cloud-canvas-ai/\">CloudCanvasAI</a>, <a href=\"/project/gemini-alchemy/\">Gemini Alchemy</a>, and the <a href=\"/project/archi/\">RAG copilot for CMS operations</a> at CERN. The pattern that held up is the dull one: stay under the limit deliberately, and keep retries as the net you hope not to need. The services that survive a traffic spike are not the ones with the cleverest retry loop. They are the ones that saw the wall coming and slowed down first.</p>\n<hr>\n<p><em>Diagrams by M. Hassan Ahmed, released under CC0. Image credit: original work by the author.</em></p>","frontmatter":{"title":"Handling LLM API Rate Limits: Retries and Backoff","date":"2026-07-11T00:00:00.000Z","description":"Your LLM backend returns 429s the moment traffic bursts. How to retry with backoff and jitter, respect Retry-After, and pace fan-out to stay under the limit.","thumbnail":{"childImageSharp":{"fluid":{"base64":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAABQAAAALCAIAAADwazoUAAAACXBIWXMAAAsSAAALEgHS3X78AAABvElEQVQoz22Q2XLbMAxF9WQt3EktpERSu+TYdVwvrSeebE7+/58K2X3MzCHmcoBLEAieJrsBZjf15nrZfn+er5dfry+7z7cD8Pby/NC3tyPE9+v+fJjWYwkWMAYrWoe0AUD48fJ0+Oq3r8D6921z/F7vb8P2fXP4Gncf0+5j3t9Md1qRGgBLgKUj0i+oGgkfMou4S4SLuY2ZTRbtE9Dc4qLDqlFZy/Oepg2WPoDSmJYhNhGtuPA266SsU9VI1UAKMUuEp3ekGXnaclmzfIBXEmYDOFW7Gzd/lZnytJ7tAA2pcEL6wY5zs9V5K6VTyj/ignTatFR5MFci7zI7c9WKesr3zwgbqhyipflzEkWLhQUQt0hYIi1llenmbJgRKYOEljExEdYAjIqzJqEVfAexCmVNDHHhfr0LqE+Ux2mdgDmmJsR6gWjJqiZzLNVcwQpKSP1IRHSINLQMoKFvBuuHiJS5bsZ+TEhW6pzxggqNmVng5r+4a55a51umqiAmBVOOZzWF2UxbHLf9rtuchvV5KnsfwyzMLNBH1FhaYbq0GmAv0DlfJRkQoixCeZjkLK2KyqfaRbiA7A9AGVrEPzB4P5h5opdNAAAAAElFTkSuQmCC","aspectRatio":1.899441340782123,"src":"/static/30ff3f68b1dcfb1160e7300152411747/40a76/hero.png","srcSet":"/static/30ff3f68b1dcfb1160e7300152411747/c972b/hero.png 340w,\n/static/30ff3f68b1dcfb1160e7300152411747/27625/hero.png 680w,\n/static/30ff3f68b1dcfb1160e7300152411747/40a76/hero.png 1360w,\n/static/30ff3f68b1dcfb1160e7300152411747/ed396/hero.png 2000w","sizes":"(max-width: 1360px) 100vw, 1360px"}}}}}},"pageContext":{"slug":"/2026-07-11-llm-api-rate-limits-retries-backoff/","previous":"blog/2026-07-12-llm-kv-cache-gpu-memory/","next":"blog/2026-07-14-reliable-json-from-llms/"}},"staticQueryHashes":["32046230"]}