{"componentChunkName":"component---src-templates-blog-post-js","path":"/blog/2026-08-14-fastapi-token-bucket-rate-limiting/","result":{"data":{"site":{"siteMetadata":{"title":"M.Hassan Ahmed","author":"Hassan11196"}},"markdownRemark":{"id":"5ce0839d-0e24-5d40-aa41-f3a3b56c2974","excerpt":"Most of my FastAPI services sit in front of something expensive: a model call that costs real money per request, a vector search that pins a CPU, or a document…","html":"<p>Most of my FastAPI services sit in front of something expensive: a model call that costs real money per request, a vector search that pins a CPU, or a document render inside a sandbox. One enthusiastic client with a <code class=\"language-text\">for</code> loop can run up a bill or starve everyone else before you notice. The fix is boring and old: cap how fast any single caller can hit the endpoint.</p>\n<p>I wrote earlier about <a href=\"/blog/2026-07-11-llm-api-rate-limits-retries-backoff/\">surviving an LLM provider’s rate limits</a> from the client side, which means backing off when <em>they</em> send you a 429 (“Too Many Requests”). This post covers the other end of that pipe. Here you are the API, and you decide who gets a 429 and when. The algorithm I keep reaching for is the token bucket, because it handles the case that matters in practice: a client that is usually quiet but occasionally needs to fire a short burst.</p>\n<p>This is for engineers running a FastAPI backend who want per-user throttling that works across more than one worker process. I build it in-process first, because that version is easy to reason about. Then I move the state into <a href=\"https://redis.io/\">Redis</a> once we hit the wall that in-process state always hits, wire it into FastAPI, and go through the failure modes.</p>\n<h2>Why a token bucket and not a simple counter</h2>\n<p>The naive approach is a fixed window: count each user’s requests per minute and reset the counter at the top of the minute. It works until you look at the boundary between windows. A client can send a whole minute’s allowance in the last second of one window, then the whole next allowance in the first second of the next. So 100/minute becomes 200 requests in two seconds across the seam, and the counter never notices.</p>\n<p>The <a href=\"https://en.wikipedia.org/wiki/Token_bucket\">token bucket</a> fixes that with two independent knobs:</p>\n<ul>\n<li>A bucket holds up to <code class=\"language-text\">C</code> tokens (its capacity).</li>\n<li>Tokens refill at <code class=\"language-text\">r</code> per second, up to that cap (the refill rate).</li>\n<li>Every request takes one token. If the bucket is empty, the request is refused.</li>\n</ul>\n<p>Capacity controls how big a burst you tolerate, and the refill rate controls the sustained pace. They are separate on purpose, which is what makes the algorithm worth its small amount of extra code.</p>\n<p><img src=\"/ba0822f2353f6049ee93e21caa456b0b/token-bucket.svg\" alt=\"Token bucket mechanics: tokens refill at r per second into a bucket of capacity C, each request removes a token, and an empty bucket returns a 429 with Retry-After\"></p>\n<p>A client sending requests at or below <code class=\"language-text\">r</code> per second never empties the bucket and never sees a limit. A client that spikes drains the bucket to zero, gets served for the length of that burst, and is then throttled to the refill rate until it eases off. That shape matches real traffic far better than a hard per-minute cap. It is why the token bucket shows up everywhere, from network shapers to the rate limiters that big APIs run in front of their edge.</p>\n<h2>The in-process version</h2>\n<p>Start with a single class. It stores a token count and the timestamp of the last refill. Instead of running a background timer to add tokens, it refills lazily: when a request arrives, it adds the tokens earned since the last refill. Lazy refill keeps this cheap, because you only do arithmetic when a request actually shows up.</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\"><span class=\"token keyword\">import</span> time\n<span class=\"token keyword\">from</span> dataclasses <span class=\"token keyword\">import</span> dataclass\n\n<span class=\"token decorator annotation punctuation\">@dataclass</span>\n<span class=\"token keyword\">class</span> <span class=\"token class-name\">TokenBucket</span><span class=\"token punctuation\">:</span>\n    capacity<span class=\"token punctuation\">:</span> <span class=\"token builtin\">float</span>      <span class=\"token comment\"># burst size, C</span>\n    refill_rate<span class=\"token punctuation\">:</span> <span class=\"token builtin\">float</span>   <span class=\"token comment\"># tokens per second, r</span>\n    tokens<span class=\"token punctuation\">:</span> <span class=\"token builtin\">float</span> <span class=\"token operator\">=</span> <span class=\"token boolean\">None</span>\n    updated<span class=\"token punctuation\">:</span> <span class=\"token builtin\">float</span> <span class=\"token operator\">=</span> <span class=\"token boolean\">None</span>\n\n    <span class=\"token keyword\">def</span> <span class=\"token function\">allow</span><span class=\"token punctuation\">(</span>self<span class=\"token punctuation\">,</span> cost<span class=\"token punctuation\">:</span> <span class=\"token builtin\">float</span> <span class=\"token operator\">=</span> <span class=\"token number\">1.0</span><span class=\"token punctuation\">)</span> <span class=\"token operator\">-</span><span class=\"token operator\">></span> <span class=\"token builtin\">bool</span><span class=\"token punctuation\">:</span>\n        now <span class=\"token operator\">=</span> time<span class=\"token punctuation\">.</span>monotonic<span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span>\n        <span class=\"token keyword\">if</span> self<span class=\"token punctuation\">.</span>tokens <span class=\"token keyword\">is</span> <span class=\"token boolean\">None</span><span class=\"token punctuation\">:</span>\n            self<span class=\"token punctuation\">.</span>tokens<span class=\"token punctuation\">,</span> self<span class=\"token punctuation\">.</span>updated <span class=\"token operator\">=</span> self<span class=\"token punctuation\">.</span>capacity<span class=\"token punctuation\">,</span> now\n\n        <span class=\"token comment\"># refill for the time that passed, capped at capacity</span>\n        elapsed <span class=\"token operator\">=</span> now <span class=\"token operator\">-</span> self<span class=\"token punctuation\">.</span>updated\n        self<span class=\"token punctuation\">.</span>tokens <span class=\"token operator\">=</span> <span class=\"token builtin\">min</span><span class=\"token punctuation\">(</span>self<span class=\"token punctuation\">.</span>capacity<span class=\"token punctuation\">,</span> self<span class=\"token punctuation\">.</span>tokens <span class=\"token operator\">+</span> elapsed <span class=\"token operator\">*</span> self<span class=\"token punctuation\">.</span>refill_rate<span class=\"token punctuation\">)</span>\n        self<span class=\"token punctuation\">.</span>updated <span class=\"token operator\">=</span> now\n\n        <span class=\"token keyword\">if</span> self<span class=\"token punctuation\">.</span>tokens <span class=\"token operator\">>=</span> cost<span class=\"token punctuation\">:</span>\n            self<span class=\"token punctuation\">.</span>tokens <span class=\"token operator\">-=</span> cost\n            <span class=\"token keyword\">return</span> <span class=\"token boolean\">True</span>\n        <span class=\"token keyword\">return</span> <span class=\"token boolean\">False</span></code></pre></div>\n<p>Two decisions in there are deliberate:</p>\n<ul>\n<li><strong>Monotonic time.</strong> I use <a href=\"https://docs.python.org/3/library/time.html#time.monotonic\"><code class=\"language-text\">time.monotonic()</code></a> rather than wall-clock time. The wall clock can jump backwards on an NTP (network time) correction, which can hand a client free tokens or, worse, produce a negative elapsed value.</li>\n<li><strong>A <code class=\"language-text\">cost</code> parameter.</strong> Not every request is equal. A cheap health check and a call that triggers a 30-second model generation should not draw the same single token. Passing a cost lets you charge the expensive route more.</li>\n</ul>\n<p>This class is genuinely useful for a single-process service or a quick local guard. It is also a trap, and the trap has a specific shape.</p>\n<h2>Where in-process state falls apart: more than one worker</h2>\n<p>The moment you run more than one worker, that dictionary of buckets is per-process. Uvicorn with four workers, or three replicas behind a load balancer, means three or four independent copies of every client’s bucket. Each copy enforces the limit you set. So your real limit is the configured limit times the number of workers, and it drifts every time you scale.</p>\n<p><img src=\"/89ecb52f31f37dfeda04eb82a56cc177/inprocess-vs-shared.svg\" alt=\"In-process buckets multiply the limit by replica count, while a shared Redis bucket keeps one true count using an atomic Lua script\"></p>\n<p>The fix is to move the bucket out of the process and into a store every worker shares. Redis is the usual choice. It is fast, and more importantly here, it can run a script atomically. That atomicity is the whole game: the check (“do I have a token?”), the refill, and the decrement have to happen as one indivisible step. If two workers both read “1 token left” and both decide they can spend it, you have handed out two tokens where you had one. Under real concurrency, that race fires constantly.</p>\n<h2>Moving the bucket into Redis with a Lua script</h2>\n<p>A <a href=\"https://redis.io/docs/latest/develop/interact/programmability/eval-intro/\">Redis Lua script</a> runs on the server without interruption from other commands, so one worker’s read-refill-write cycle can’t interleave with another’s. Redis even documents a token bucket limiter as a <a href=\"https://redis.io/docs/latest/develop/use-cases/rate-limiter/redis-py/\">recommended pattern</a>.</p>\n<p>Here is the script I use. It reads the stored tokens and timestamp, refills against the server’s own clock, and decrements if it can (or works out how long the caller must wait). It also sets a TTL (time to live) on the key, so idle buckets evict themselves instead of leaking keys forever.</p>\n<div class=\"gatsby-highlight\" data-language=\"lua\"><pre class=\"language-lua\"><code class=\"language-lua\"><span class=\"token comment\">-- KEYS[1] = bucket key (e.g. \"rl:user:42\")</span>\n<span class=\"token comment\">-- ARGV[1] = capacity, ARGV[2] = refill_rate/sec, ARGV[3] = cost</span>\n<span class=\"token keyword\">local</span> capacity <span class=\"token operator\">=</span> <span class=\"token function\">tonumber</span><span class=\"token punctuation\">(</span>ARGV<span class=\"token punctuation\">[</span><span class=\"token number\">1</span><span class=\"token punctuation\">]</span><span class=\"token punctuation\">)</span>\n<span class=\"token keyword\">local</span> rate     <span class=\"token operator\">=</span> <span class=\"token function\">tonumber</span><span class=\"token punctuation\">(</span>ARGV<span class=\"token punctuation\">[</span><span class=\"token number\">2</span><span class=\"token punctuation\">]</span><span class=\"token punctuation\">)</span>\n<span class=\"token keyword\">local</span> cost     <span class=\"token operator\">=</span> <span class=\"token function\">tonumber</span><span class=\"token punctuation\">(</span>ARGV<span class=\"token punctuation\">[</span><span class=\"token number\">3</span><span class=\"token punctuation\">]</span><span class=\"token punctuation\">)</span>\n\n<span class=\"token comment\">-- one clock for every worker: Redis' own TIME, not the app server's</span>\n<span class=\"token keyword\">local</span> t   <span class=\"token operator\">=</span> redis<span class=\"token punctuation\">.</span><span class=\"token function\">call</span><span class=\"token punctuation\">(</span><span class=\"token string\">'TIME'</span><span class=\"token punctuation\">)</span>\n<span class=\"token keyword\">local</span> now <span class=\"token operator\">=</span> <span class=\"token function\">tonumber</span><span class=\"token punctuation\">(</span>t<span class=\"token punctuation\">[</span><span class=\"token number\">1</span><span class=\"token punctuation\">]</span><span class=\"token punctuation\">)</span> <span class=\"token operator\">+</span> <span class=\"token function\">tonumber</span><span class=\"token punctuation\">(</span>t<span class=\"token punctuation\">[</span><span class=\"token number\">2</span><span class=\"token punctuation\">]</span><span class=\"token punctuation\">)</span> <span class=\"token operator\">/</span> <span class=\"token number\">1000000</span>\n\n<span class=\"token keyword\">local</span> state  <span class=\"token operator\">=</span> redis<span class=\"token punctuation\">.</span><span class=\"token function\">call</span><span class=\"token punctuation\">(</span><span class=\"token string\">'HMGET'</span><span class=\"token punctuation\">,</span> KEYS<span class=\"token punctuation\">[</span><span class=\"token number\">1</span><span class=\"token punctuation\">]</span><span class=\"token punctuation\">,</span> <span class=\"token string\">'tokens'</span><span class=\"token punctuation\">,</span> <span class=\"token string\">'ts'</span><span class=\"token punctuation\">)</span>\n<span class=\"token keyword\">local</span> tokens <span class=\"token operator\">=</span> <span class=\"token function\">tonumber</span><span class=\"token punctuation\">(</span>state<span class=\"token punctuation\">[</span><span class=\"token number\">1</span><span class=\"token punctuation\">]</span><span class=\"token punctuation\">)</span>\n<span class=\"token keyword\">local</span> ts     <span class=\"token operator\">=</span> <span class=\"token function\">tonumber</span><span class=\"token punctuation\">(</span>state<span class=\"token punctuation\">[</span><span class=\"token number\">2</span><span class=\"token punctuation\">]</span><span class=\"token punctuation\">)</span>\n<span class=\"token keyword\">if</span> tokens <span class=\"token operator\">==</span> <span class=\"token keyword\">nil</span> <span class=\"token keyword\">then</span>\n  tokens<span class=\"token punctuation\">,</span> ts <span class=\"token operator\">=</span> capacity<span class=\"token punctuation\">,</span> now\n<span class=\"token keyword\">end</span>\n\n<span class=\"token keyword\">local</span> tokens <span class=\"token operator\">=</span> math<span class=\"token punctuation\">.</span><span class=\"token function\">min</span><span class=\"token punctuation\">(</span>capacity<span class=\"token punctuation\">,</span> tokens <span class=\"token operator\">+</span> math<span class=\"token punctuation\">.</span><span class=\"token function\">max</span><span class=\"token punctuation\">(</span><span class=\"token number\">0</span><span class=\"token punctuation\">,</span> now <span class=\"token operator\">-</span> ts<span class=\"token punctuation\">)</span> <span class=\"token operator\">*</span> rate<span class=\"token punctuation\">)</span>\n\n<span class=\"token keyword\">local</span> allowed<span class=\"token punctuation\">,</span> retry_after <span class=\"token operator\">=</span> <span class=\"token number\">0</span><span class=\"token punctuation\">,</span> <span class=\"token number\">0.0</span>\n<span class=\"token keyword\">if</span> tokens <span class=\"token operator\">>=</span> cost <span class=\"token keyword\">then</span>\n  tokens  <span class=\"token operator\">=</span> tokens <span class=\"token operator\">-</span> cost\n  allowed <span class=\"token operator\">=</span> <span class=\"token number\">1</span>\n<span class=\"token keyword\">else</span>\n  retry_after <span class=\"token operator\">=</span> <span class=\"token punctuation\">(</span>cost <span class=\"token operator\">-</span> tokens<span class=\"token punctuation\">)</span> <span class=\"token operator\">/</span> rate     <span class=\"token comment\">-- seconds until enough refills</span>\n<span class=\"token keyword\">end</span>\n\nredis<span class=\"token punctuation\">.</span><span class=\"token function\">call</span><span class=\"token punctuation\">(</span><span class=\"token string\">'HSET'</span><span class=\"token punctuation\">,</span> KEYS<span class=\"token punctuation\">[</span><span class=\"token number\">1</span><span class=\"token punctuation\">]</span><span class=\"token punctuation\">,</span> <span class=\"token string\">'tokens'</span><span class=\"token punctuation\">,</span> tokens<span class=\"token punctuation\">,</span> <span class=\"token string\">'ts'</span><span class=\"token punctuation\">,</span> now<span class=\"token punctuation\">)</span>\nredis<span class=\"token punctuation\">.</span><span class=\"token function\">call</span><span class=\"token punctuation\">(</span><span class=\"token string\">'EXPIRE'</span><span class=\"token punctuation\">,</span> KEYS<span class=\"token punctuation\">[</span><span class=\"token number\">1</span><span class=\"token punctuation\">]</span><span class=\"token punctuation\">,</span> math<span class=\"token punctuation\">.</span><span class=\"token function\">ceil</span><span class=\"token punctuation\">(</span>capacity <span class=\"token operator\">/</span> rate<span class=\"token punctuation\">)</span> <span class=\"token operator\">+</span> <span class=\"token number\">1</span><span class=\"token punctuation\">)</span>\n<span class=\"token keyword\">return</span> <span class=\"token punctuation\">{</span>allowed<span class=\"token punctuation\">,</span> <span class=\"token function\">tostring</span><span class=\"token punctuation\">(</span>tokens<span class=\"token punctuation\">)</span><span class=\"token punctuation\">,</span> <span class=\"token function\">tostring</span><span class=\"token punctuation\">(</span>retry_after<span class=\"token punctuation\">)</span><span class=\"token punctuation\">}</span></code></pre></div>\n<p>Reading the clock with <code class=\"language-text\">redis.call(&#39;TIME&#39;)</code>, instead of passing the time in from Python, is a small choice that removes a whole class of bug. Every worker now measures elapsed time against the same clock, so a few seconds of skew between app servers can’t leak or eat tokens. Redis returns integers cleanly but not floats, which is why the token count and retry hint come back as strings for the caller to parse.</p>\n<p>The Python side registers the script once and then calls it. <a href=\"https://redis-py.readthedocs.io/en/stable/\"><code class=\"language-text\">register_script</code></a> handles the <code class=\"language-text\">EVALSHA</code> caching: it sends the script body over the wire once and references it by hash after that. <code class=\"language-text\">take_token</code> then parses the three results back into a boolean and two floats.</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\"><span class=\"token keyword\">import</span> redis\n\npool <span class=\"token operator\">=</span> redis<span class=\"token punctuation\">.</span>Redis<span class=\"token punctuation\">(</span>host<span class=\"token operator\">=</span><span class=\"token string\">\"localhost\"</span><span class=\"token punctuation\">,</span> port<span class=\"token operator\">=</span><span class=\"token number\">6379</span><span class=\"token punctuation\">,</span> decode_responses<span class=\"token operator\">=</span><span class=\"token boolean\">True</span><span class=\"token punctuation\">)</span>\n_take <span class=\"token operator\">=</span> pool<span class=\"token punctuation\">.</span>register_script<span class=\"token punctuation\">(</span>LUA<span class=\"token punctuation\">)</span>   <span class=\"token comment\"># LUA = the script above, as a string</span>\n\n<span class=\"token keyword\">def</span> <span class=\"token function\">take_token</span><span class=\"token punctuation\">(</span>key<span class=\"token punctuation\">:</span> <span class=\"token builtin\">str</span><span class=\"token punctuation\">,</span> capacity<span class=\"token punctuation\">:</span> <span class=\"token builtin\">float</span><span class=\"token punctuation\">,</span> rate<span class=\"token punctuation\">:</span> <span class=\"token builtin\">float</span><span class=\"token punctuation\">,</span> cost<span class=\"token punctuation\">:</span> <span class=\"token builtin\">float</span> <span class=\"token operator\">=</span> <span class=\"token number\">1.0</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n    allowed<span class=\"token punctuation\">,</span> tokens<span class=\"token punctuation\">,</span> retry_after <span class=\"token operator\">=</span> _take<span class=\"token punctuation\">(</span>keys<span class=\"token operator\">=</span><span class=\"token punctuation\">[</span>key<span class=\"token punctuation\">]</span><span class=\"token punctuation\">,</span> args<span class=\"token operator\">=</span><span class=\"token punctuation\">[</span>capacity<span class=\"token punctuation\">,</span> rate<span class=\"token punctuation\">,</span> cost<span class=\"token punctuation\">]</span><span class=\"token punctuation\">)</span>\n    <span class=\"token keyword\">return</span> <span class=\"token builtin\">bool</span><span class=\"token punctuation\">(</span><span class=\"token builtin\">int</span><span class=\"token punctuation\">(</span>allowed<span class=\"token punctuation\">)</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">,</span> <span class=\"token builtin\">float</span><span class=\"token punctuation\">(</span>tokens<span class=\"token punctuation\">)</span><span class=\"token punctuation\">,</span> <span class=\"token builtin\">float</span><span class=\"token punctuation\">(</span>retry_after<span class=\"token punctuation\">)</span></code></pre></div>\n<h2>Wiring it into FastAPI with a dependency</h2>\n<p>FastAPI’s <a href=\"https://fastapi.tiangolo.com/tutorial/dependencies/\">dependency system</a> is the natural place for this. A dependency runs before the handler, can read the request, and can abort with an <a href=\"https://fastapi.tiangolo.com/tutorial/handling-errors/\"><code class=\"language-text\">HTTPException</code></a> before any expensive work starts. That last part matters: you want to reject a throttled request before it reaches the model call, not after.</p>\n<p>The dependency below keys each client by API key (or IP as a fallback), takes a token, and either raises a 429 or reports the remaining budget in a header:</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\"><span class=\"token keyword\">import</span> math\n<span class=\"token keyword\">from</span> fastapi <span class=\"token keyword\">import</span> Depends<span class=\"token punctuation\">,</span> FastAPI<span class=\"token punctuation\">,</span> Request<span class=\"token punctuation\">,</span> Response<span class=\"token punctuation\">,</span> HTTPException\n\napp <span class=\"token operator\">=</span> FastAPI<span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span>\n\nBURST <span class=\"token operator\">=</span> <span class=\"token number\">20</span>        <span class=\"token comment\"># capacity: allow a short spike of 20</span>\nRATE  <span class=\"token operator\">=</span> <span class=\"token number\">5.0</span>       <span class=\"token comment\"># 5 tokens/sec sustained -> 300 requests/min</span>\n\n<span class=\"token keyword\">def</span> <span class=\"token function\">rate_limit</span><span class=\"token punctuation\">(</span>request<span class=\"token punctuation\">:</span> Request<span class=\"token punctuation\">,</span> response<span class=\"token punctuation\">:</span> Response<span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n    client <span class=\"token operator\">=</span> request<span class=\"token punctuation\">.</span>headers<span class=\"token punctuation\">.</span>get<span class=\"token punctuation\">(</span><span class=\"token string\">\"x-api-key\"</span><span class=\"token punctuation\">)</span> <span class=\"token keyword\">or</span> request<span class=\"token punctuation\">.</span>client<span class=\"token punctuation\">.</span>host\n    key <span class=\"token operator\">=</span> <span class=\"token string-interpolation\"><span class=\"token string\">f\"rl:</span><span class=\"token interpolation\"><span class=\"token punctuation\">{</span>client<span class=\"token punctuation\">}</span></span><span class=\"token string\">\"</span></span>\n    <span class=\"token keyword\">try</span><span class=\"token punctuation\">:</span>\n        allowed<span class=\"token punctuation\">,</span> tokens<span class=\"token punctuation\">,</span> retry_after <span class=\"token operator\">=</span> take_token<span class=\"token punctuation\">(</span>key<span class=\"token punctuation\">,</span> BURST<span class=\"token punctuation\">,</span> RATE<span class=\"token punctuation\">)</span>\n    <span class=\"token keyword\">except</span> redis<span class=\"token punctuation\">.</span>RedisError<span class=\"token punctuation\">:</span>\n        <span class=\"token keyword\">return</span>  <span class=\"token comment\"># Redis down: fail open, don't take the whole API down with it</span>\n\n    <span class=\"token keyword\">if</span> <span class=\"token keyword\">not</span> allowed<span class=\"token punctuation\">:</span>\n        <span class=\"token keyword\">raise</span> HTTPException<span class=\"token punctuation\">(</span>\n            status_code<span class=\"token operator\">=</span><span class=\"token number\">429</span><span class=\"token punctuation\">,</span>\n            detail<span class=\"token operator\">=</span><span class=\"token string\">\"rate limit exceeded\"</span><span class=\"token punctuation\">,</span>\n            headers<span class=\"token operator\">=</span><span class=\"token punctuation\">{</span><span class=\"token string\">\"Retry-After\"</span><span class=\"token punctuation\">:</span> <span class=\"token builtin\">str</span><span class=\"token punctuation\">(</span>math<span class=\"token punctuation\">.</span>ceil<span class=\"token punctuation\">(</span>retry_after<span class=\"token punctuation\">)</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">}</span><span class=\"token punctuation\">,</span>\n        <span class=\"token punctuation\">)</span>\n    response<span class=\"token punctuation\">.</span>headers<span class=\"token punctuation\">[</span><span class=\"token string\">\"X-RateLimit-Remaining\"</span><span class=\"token punctuation\">]</span> <span class=\"token operator\">=</span> <span class=\"token builtin\">str</span><span class=\"token punctuation\">(</span><span class=\"token builtin\">int</span><span class=\"token punctuation\">(</span>tokens<span class=\"token punctuation\">)</span><span class=\"token punctuation\">)</span>\n\n<span class=\"token decorator annotation punctuation\">@app<span class=\"token punctuation\">.</span>post</span><span class=\"token punctuation\">(</span><span class=\"token string\">\"/generate\"</span><span class=\"token punctuation\">,</span> dependencies<span class=\"token operator\">=</span><span class=\"token punctuation\">[</span>Depends<span class=\"token punctuation\">(</span>rate_limit<span class=\"token punctuation\">)</span><span class=\"token punctuation\">]</span><span class=\"token punctuation\">)</span>\n<span class=\"token keyword\">async</span> <span class=\"token keyword\">def</span> <span class=\"token function\">generate</span><span class=\"token punctuation\">(</span>prompt<span class=\"token punctuation\">:</span> <span class=\"token builtin\">str</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n    <span class=\"token keyword\">return</span> <span class=\"token keyword\">await</span> run_expensive_model<span class=\"token punctuation\">(</span>prompt<span class=\"token punctuation\">)</span></code></pre></div>\n<p>Two response details are worth getting right:</p>\n<ul>\n<li><strong>A 429 with <code class=\"language-text\">Retry-After</code>.</strong> Sending <a href=\"https://developer.mozilla.org/en-US/docs/Web/HTTP/Status/429\"><code class=\"language-text\">429 Too Many Requests</code></a> with a <a href=\"https://developer.mozilla.org/en-US/docs/Web/HTTP/Headers/Retry-After\"><code class=\"language-text\">Retry-After</code></a> header, which is standardized in <a href=\"https://www.rfc-editor.org/rfc/rfc9110#field.retry-after\">RFC 9110</a>, tells a well-behaved client exactly when to come back. Otherwise it has to guess, and it will hammer you.</li>\n<li><strong><code class=\"language-text\">X-RateLimit-Remaining</code> on success.</strong> This header is a convention, not a standard, but it is a common one. It lets a client throttle itself before it trips the limit at all.</li>\n</ul>\n<p>A request you never have to reject is cheaper than the fastest 429.</p>\n<h2>Failure modes I’ve actually hit</h2>\n<p><strong>Redis being down.</strong> The dependency above fails open: if Redis is unreachable, requests pass. That is the right default for most public APIs, where a rate-limiter outage taking down the whole service is worse than briefly not enforcing limits. But it is a real decision. If you are protecting something where an unbounded burst is dangerous, fail closed instead and return 429 or 503 when the limiter can’t answer. Pick one on purpose and write down why, because the default you fall into by accident is rarely the one you’d choose.</p>\n<p><strong>Choosing the wrong key.</strong> <code class=\"language-text\">request.client.host</code> is easy and often wrong. Behind a load balancer or proxy, it may be the proxy’s IP, so every client shares one bucket and one noisy user throttles everyone. If you terminate TLS at a proxy, read the real client from a trusted forwarded header, and only one you set yourself. For authenticated traffic, key on the API key or user id rather than IP. That is the thing you actually want to limit, and it doesn’t break for users behind a shared NAT.</p>\n<p><strong>Counting requests when the cost is tokens.</strong> For an LLM endpoint, one request is not one unit of load. A 200-token completion and a 4,000-token one hit your budget very differently. This is the same trap I described in the <a href=\"/blog/2026-07-11-llm-api-rate-limits-retries-backoff/\">provider rate-limit post</a>, just on the serving side. If you limit request count while your real constraint is tokens per minute, a few large requests blow the budget while the request counter looks calm. The <code class=\"language-text\">cost</code> parameter is the hook for this. Charge each route its rough expected cost, and cap the genuinely heavy work with a separate, smaller bucket.</p>\n<p><strong>Blocking the event loop.</strong> <code class=\"language-text\">redis-py</code> is synchronous. Calling it straight from an async endpoint parks the whole event loop on a network round trip, which is exactly the <a href=\"/blog/2026-07-10-fastapi-blocking-event-loop/\">blocking-the-event-loop problem</a> I’ve written about. Use <a href=\"https://redis.readthedocs.io/en/stable/examples/asyncio_examples.html\"><code class=\"language-text\">redis.asyncio</code></a> so the limiter check is awaited like everything else on the path.</p>\n<h2>How it compares with other rate-limiting algorithms</h2>\n<p>Token bucket is not the only option, and it is not always the right one. Two alternatives come up most often:</p>\n<ul>\n<li><strong>Sliding-window log.</strong> It stores a timestamp per request and counts the ones inside the window. It is exact and makes no compromise at window boundaries, but it stores O(requests) data per client, which gets expensive under load.</li>\n<li><strong>Leaky bucket.</strong> It enforces a perfectly smooth output rate with no burst at all. That is what you want when feeding a downstream system that truly can’t spike, and the wrong choice when a bursty client is normal and fine.</li>\n</ul>\n<p>The <a href=\"https://redis.io/tutorials/howtos/ratelimiting/\">Redis rate-limiting guide</a> walks through several of these side by side.</p>\n<p>I default to token bucket because most of my traffic is bursty but bounded: a user opens a page, fires a handful of requests, then goes quiet. Capacity absorbs the handful, the refill rate holds the long-run average, and the memory cost is two numbers per client. When I’ve needed strict smoothing into a fragile downstream, I’ve switched to leaky bucket for that hop specifically and kept token bucket at the edge.</p>\n<h2>What I’d do differently</h2>\n<p>The mistake I made the first time was hard-coding one limit for the whole service. Real endpoints don’t deserve the same budget. A login route and a document-generation route have completely different cost profiles, and one global bucket either throttles the cheap route too hard or lets the expensive one run free. Now I parameterize the dependency, so <code class=\"language-text\">Depends(rate_limit)</code> becomes a small factory that takes capacity and rate per route group. Same script, same Redis, different knobs.</p>\n<p>I’d also add the <code class=\"language-text\">X-RateLimit-*</code> response headers from the start rather than bolting them on later. Once clients can see how much budget they have left, the well-behaved ones pace themselves, and your 429 rate drops on its own. It’s a few lines that pay for themselves in support tickets you never get.</p>\n<h2>Closing</h2>\n<p>This limiter is a small piece of plumbing in front of the expensive parts of my projects: the FastAPI backends behind <a href=\"/project/archi/\">Archi</a>, <a href=\"/project/cloud-canvas-ai/\">CloudCanvasAI</a>, and <a href=\"/project/gemini-alchemy/\">Gemini Alchemy</a>. There, a request can mean a model call, a sandboxed render, or a retrieval over a large index, and none of those should discover their load profile from a runaway client. Two knobs, one atomic script, and a 429 that tells the caller when to try again cover the case surprisingly well. The whole thing is short enough to read in one sitting.</p>\n<hr>\n<p><em>Image credit: token bucket diagrams by M. Hassan Ahmed, created for this post, released under <a href=\"https://creativecommons.org/publicdomain/zero/1.0/\">CC0 1.0</a> (public domain).</em></p>","frontmatter":{"title":"Rate Limiting a FastAPI Service with a Token Bucket","date":"2026-08-14T00:00:00.000Z","description":"Add per-user rate limiting to a FastAPI backend with the token bucket algorithm: an in-process version, an atomic Redis script, 429s, and the failure modes.","thumbnail":{"childImageSharp":{"fluid":{"base64":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAABQAAAALCAIAAADwazoUAAAACXBIWXMAAAsSAAALEgHS3X78AAACEElEQVQozzWQ2W7UMBSGI3WZ7IkTO7FjZ0+afWmmdDrT0lbttIhFhUIFXCKehEseg0VccQGPiKct0nfx+0j/8WcLmJ4QeuKSVV48G6a33fimG1+Pcx5uhum2HW7q9tX+wbt+fzPvhg1JfuV6x7wlSKYnAzrTMEv7/aPLalhV43EzPc3bBQ/1eJI3h0W/LIdVOSz3ukU7PyVRI+pENqmggIAjm75kMNFgkuErgEomEQ13pqNdDXJ2VFvUPRXEupUoRqiCUDEDjqCYvmpFAO9xdCfT7Lhu7obxc1l9zIv3cXqbpLdpdsclTZKxZo6yxkpKg6WKFQiSTnWUsnyBo8lmveFkhwdfrs9+rU+/XZ//vDr7fnX+4/nF77L8pDDmTB0aGrstzSpXnFCQDSqbTN4Is3t5yq00K1ZBpFkJf9gDXJBfw+U50mOFCpJGdJiE5TGOD1A4AVysjr6+WP95uf6bVx9QMLnRkRsvIZ2bUeYtBjLvUF+BOleRz7UJ36HBWLVCGYR8JfGWQXDBAU7Nv8OECcewIhFQ0fEfQdzRE0TVVUGAw8HxOxr0AKWYVHBDUbKy8ivX61w6OKSmNiucIEd+Bv3UpqZBeNlRrQCy1iZVnj+Bbk5IjdmIaV+FbR22mHaYbcoeZA1NCxLt4SiGzHgoiyraVeBMQVuidR+geM+2grZlJP4/7ijOlryZ8DnPM9X5B06KUziMKam8AAAAAElFTkSuQmCC","aspectRatio":1.899441340782123,"src":"/static/71976257c3eef0cc9cb8c8710a4c6c3c/40a76/hero.png","srcSet":"/static/71976257c3eef0cc9cb8c8710a4c6c3c/c972b/hero.png 340w,\n/static/71976257c3eef0cc9cb8c8710a4c6c3c/27625/hero.png 680w,\n/static/71976257c3eef0cc9cb8c8710a4c6c3c/40a76/hero.png 1360w,\n/static/71976257c3eef0cc9cb8c8710a4c6c3c/ed396/hero.png 2000w","sizes":"(max-width: 1360px) 100vw, 1360px"}}}}}},"pageContext":{"slug":"/2026-08-14-fastapi-token-bucket-rate-limiting/","previous":"blog/2026-08-11-how-flashattention-speeds-up-attention/","next":"blog/2026-08-13-zero-downtime-reindexing-opensearch/"}},"staticQueryHashes":["32046230"]}