{"componentChunkName":"component---src-templates-blog-post-js","path":"/blog/2026-08-19-concurrent-llm-calls-asyncio-gather/","result":{"data":{"site":{"siteMetadata":{"title":"M.Hassan Ahmed","author":"Hassan11196"}},"markdownRemark":{"id":"a37228c9-b46e-5bd9-ba71-1b923a94d004","excerpt":"A retrieval request in Archi does not touch one source. To answer an operator’s question, it searches the logbook, runs a keyword pass over JIRA tickets, and…","html":"<p>A retrieval request in <a href=\"/project/archi/\">Archi</a> does not touch one source. To answer an operator’s question, it searches the logbook, runs a keyword pass over JIRA tickets, and pulls the recent portal pages, then hands the merged set to a reranker. That is four calls, all over the network, none of them depending on the others. The first version I wrote awaited them one after another. The request took as long as all four added together, even though the CPU sat idle waiting on the second call while the first one had already come back.</p>\n<p>That is the bug this post is about, and it is not a blocking-the-loop bug. It is the opposite mistake: code that is correctly async, with every call properly awaited, that still runs slow because it never overlaps the waits.</p>\n<p>This post is for engineers building an <a href=\"https://fastapi.tiangolo.com/async/\">async FastAPI</a> backend that makes more than one independent network call per request: retrieval fan-out, a batch of embeddings, several tool calls in an agent turn. The pattern here gets those seconds back. I will show the naive version and the <a href=\"https://docs.python.org/3/library/asyncio-task.html#asyncio.gather\"><code class=\"language-text\">asyncio.gather</code></a> fix. Most of the post is about the parts that bite: bounding how many calls run at once, what happens when one of them fails, and timeouts.</p>\n<h2>await does not mean concurrent</h2>\n<p>Here is the shape of the slow version: a BM25 keyword search (BM25 is a standard keyword-ranking function), a vector search, and a rerank, each an async function that awaits the network.</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\"><span class=\"token keyword\">async</span> <span class=\"token keyword\">def</span> <span class=\"token function\">retrieve</span><span class=\"token punctuation\">(</span>query<span class=\"token punctuation\">:</span> <span class=\"token builtin\">str</span><span class=\"token punctuation\">)</span> <span class=\"token operator\">-</span><span class=\"token operator\">></span> <span class=\"token builtin\">list</span><span class=\"token punctuation\">[</span>Hit<span class=\"token punctuation\">]</span><span class=\"token punctuation\">:</span>\n    bm25    <span class=\"token operator\">=</span> <span class=\"token keyword\">await</span> search_bm25<span class=\"token punctuation\">(</span>query<span class=\"token punctuation\">)</span>      <span class=\"token comment\"># ~180 ms</span>\n    vectors <span class=\"token operator\">=</span> <span class=\"token keyword\">await</span> search_vectors<span class=\"token punctuation\">(</span>query<span class=\"token punctuation\">)</span>   <span class=\"token comment\"># ~240 ms</span>\n    rerank  <span class=\"token operator\">=</span> <span class=\"token keyword\">await</span> rerank_hits<span class=\"token punctuation\">(</span>query<span class=\"token punctuation\">,</span> bm25 <span class=\"token operator\">+</span> vectors<span class=\"token punctuation\">)</span>  <span class=\"token comment\"># ~300 ms</span>\n    <span class=\"token keyword\">return</span> rerank</code></pre></div>\n<p>This code is correct. It does not block the <a href=\"/blog/2026-07-10-fastapi-blocking-event-loop/\">event loop</a>, other requests interleave fine, and the server stays responsive under load. It is just slow for this one request, because each <code class=\"language-text\">await</code> fully finishes before the next one starts. The <code class=\"language-text\">await</code> on <code class=\"language-text\">search_bm25</code> says “park me until BM25 answers”, and only when BM25 answers does the code reach the <code class=\"language-text\">await</code> on <code class=\"language-text\">search_vectors</code>. The waits happen back to back, so the latencies add up: 180 plus 240 plus 300, about 720 ms, most of it spent doing nothing.</p>\n<p>The reranker genuinely depends on the first two (it needs their hits), so it has to come last. But <code class=\"language-text\">search_bm25</code> and <code class=\"language-text\">search_vectors</code> do not depend on each other at all, and there is no reason for the second to wait for the first. That is the waste, and it is exactly the case <code class=\"language-text\">gather</code> exists for.</p>\n<p><img src=\"/f32ce67588c2d93a12deb219e4be4bec/sequential-vs-gather.svg\" alt=\"Two timelines compared. On top, three calls awaited one after another run back to back and their latencies add up to 720 ms. On the bottom, the two independent calls started together with asyncio.gather overlap, so the total collapses to about 300 ms, the length of the single slowest call.\"></p>\n<h2>gather starts them together</h2>\n<p><code class=\"language-text\">asyncio.gather</code> takes several awaitables, schedules them all at once, and returns their results in the order you passed them in, not the order they finished. Here is the same function rewritten:</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\"><span class=\"token keyword\">import</span> asyncio\n\n<span class=\"token keyword\">async</span> <span class=\"token keyword\">def</span> <span class=\"token function\">retrieve</span><span class=\"token punctuation\">(</span>query<span class=\"token punctuation\">:</span> <span class=\"token builtin\">str</span><span class=\"token punctuation\">)</span> <span class=\"token operator\">-</span><span class=\"token operator\">></span> <span class=\"token builtin\">list</span><span class=\"token punctuation\">[</span>Hit<span class=\"token punctuation\">]</span><span class=\"token punctuation\">:</span>\n    <span class=\"token comment\"># these two do not depend on each other, so run them together</span>\n    bm25<span class=\"token punctuation\">,</span> vectors <span class=\"token operator\">=</span> <span class=\"token keyword\">await</span> asyncio<span class=\"token punctuation\">.</span>gather<span class=\"token punctuation\">(</span>\n        search_bm25<span class=\"token punctuation\">(</span>query<span class=\"token punctuation\">)</span><span class=\"token punctuation\">,</span>\n        search_vectors<span class=\"token punctuation\">(</span>query<span class=\"token punctuation\">)</span><span class=\"token punctuation\">,</span>\n    <span class=\"token punctuation\">)</span>\n    <span class=\"token comment\"># the reranker needs both, so it waits for the gather</span>\n    <span class=\"token keyword\">return</span> <span class=\"token keyword\">await</span> rerank_hits<span class=\"token punctuation\">(</span>query<span class=\"token punctuation\">,</span> bm25 <span class=\"token operator\">+</span> vectors<span class=\"token punctuation\">)</span></code></pre></div>\n<p>Both searches are now in flight at the same time. The <code class=\"language-text\">await</code> on the <code class=\"language-text\">gather</code> returns when the slower of the two is done, so the pair costs about 240 ms instead of 420 ms. The whole request drops from roughly 720 ms to 540 ms. Fan out three or four independent sources instead of two and the gap grows, because the sequential version keeps adding while the concurrent one stays pinned to its slowest branch.</p>\n<p>Two details are worth internalizing:</p>\n<ul>\n<li><strong>Results come back positional.</strong> <code class=\"language-text\">gather(a, b)</code> always gives <code class=\"language-text\">[result_a, result_b]</code>, even if <code class=\"language-text\">b</code> finished first. That is why the tuple unpacking above is safe.</li>\n<li><strong>It only helps for work that is genuinely independent.</strong> If call two needs call one’s output, no amount of <code class=\"language-text\">gather</code> changes that. The dependency is real, and those calls have to run in sequence.</li>\n</ul>\n<p>The whole skill is spotting which calls in a request are actually independent. In a <a href=\"/blog/2026-07-02-hybrid-search-rag-bm25-vectors/\">hybrid search pipeline</a>, the lexical and vector passes are the obvious pair. In an agent turn, it is often several tool calls the model asked for in one step.</p>\n<h2>The trap: unbounded fan-out</h2>\n<p>The first time <code class=\"language-text\">gather</code> works, the temptation is to point it at everything. Re-embedding a thousand documents becomes a <code class=\"language-text\">gather</code> of a thousand coroutines; fifty queued questions become a <code class=\"language-text\">gather</code> of fifty LLM calls. This falls over, and it fails in a way that looks unrelated to the code you changed.</p>\n<p><code class=\"language-text\">gather</code> does not throttle. It starts every awaitable you hand it right now, all at once, so five hundred coroutines means five hundred HTTP requests opening at the same instant. That drains the <a href=\"/blog/2026-07-10-fastapi-blocking-event-loop/\">connection pool</a>. Once the pool is empty the rest queue behind it anyway, so you paid the memory for five hundred pending tasks and got none of the parallelism. Worse, if these are calls to a model provider, you just sent five hundred requests in one burst and tripped the <a href=\"/blog/2026-07-11-llm-api-rate-limits-retries-backoff/\">rate limit</a>. Now a chunk of them come back as 429s (HTTP “Too Many Requests”).</p>\n<p>The fix is to cap how many run concurrently with an <a href=\"https://docs.python.org/3/library/asyncio-sync.html#asyncio.Semaphore\"><code class=\"language-text\">asyncio.Semaphore</code></a>. A semaphore is a counter with a fixed number of permits. Each task takes a permit before doing its work and returns it afterwards, so no more than N tasks are ever inside at the same time. Wrap each call, then <code class=\"language-text\">gather</code> the wrappers:</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\"><span class=\"token keyword\">import</span> asyncio\n\n<span class=\"token keyword\">async</span> <span class=\"token keyword\">def</span> <span class=\"token function\">gather_bounded</span><span class=\"token punctuation\">(</span>limit<span class=\"token punctuation\">:</span> <span class=\"token builtin\">int</span><span class=\"token punctuation\">,</span> <span class=\"token operator\">*</span>coros<span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n    sem <span class=\"token operator\">=</span> asyncio<span class=\"token punctuation\">.</span>Semaphore<span class=\"token punctuation\">(</span>limit<span class=\"token punctuation\">)</span>\n\n    <span class=\"token keyword\">async</span> <span class=\"token keyword\">def</span> <span class=\"token function\">run</span><span class=\"token punctuation\">(</span>coro<span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n        <span class=\"token keyword\">async</span> <span class=\"token keyword\">with</span> sem<span class=\"token punctuation\">:</span>          <span class=\"token comment\"># waits here if all permits are taken</span>\n            <span class=\"token keyword\">return</span> <span class=\"token keyword\">await</span> coro\n\n    <span class=\"token keyword\">return</span> <span class=\"token keyword\">await</span> asyncio<span class=\"token punctuation\">.</span>gather<span class=\"token punctuation\">(</span><span class=\"token operator\">*</span><span class=\"token punctuation\">(</span>run<span class=\"token punctuation\">(</span>c<span class=\"token punctuation\">)</span> <span class=\"token keyword\">for</span> c <span class=\"token keyword\">in</span> coros<span class=\"token punctuation\">)</span><span class=\"token punctuation\">)</span>\n\n<span class=\"token comment\"># embed 500 documents, but never more than 8 requests in flight</span>\nvectors <span class=\"token operator\">=</span> <span class=\"token keyword\">await</span> gather_bounded<span class=\"token punctuation\">(</span><span class=\"token number\">8</span><span class=\"token punctuation\">,</span> <span class=\"token operator\">*</span><span class=\"token punctuation\">(</span>embed<span class=\"token punctuation\">(</span>doc<span class=\"token punctuation\">)</span> <span class=\"token keyword\">for</span> doc <span class=\"token keyword\">in</span> documents<span class=\"token punctuation\">)</span><span class=\"token punctuation\">)</span></code></pre></div>\n<p>All five hundred tasks still get created, but the <code class=\"language-text\">async with sem</code> line is a gate. Only eight get through at a time; the rest suspend cheaply at that line until a permit frees up. You keep the overlap and stop it from becoming a stampede. The right value for <code class=\"language-text\">limit</code> is not a guess. It is whatever your provider’s rate limit and your connection pool can actually sustain, usually a small number like 5 to 20.</p>\n<p><img src=\"/354bd7cfe2549ce227056af828eb8b32/bounded-concurrency.svg\" alt=\"A bounded fan-out. Eight tasks are all scheduled by gather at once, but a Semaphore with four permits sits in the middle as a gate. Four tasks are busy in slots, the other four wait. On the right, at most four HTTP or LLM calls are ever in flight, which is what the API and the connection pool actually see. A note reads: without the bound, 500 tasks would open 500 connections and trip the rate limit.\"></p>\n<h2>When one of them fails</h2>\n<p>Concurrency makes error handling less obvious, because now several things can go wrong at the same time. By default, <code class=\"language-text\">gather</code> is unforgiving about this. The moment one awaitable raises, that exception propagates straight up to whoever awaited the <code class=\"language-text\">gather</code>. You get the one error, and the results of everything else, including the calls that succeeded, are gone.</p>\n<p>That is often the wrong behavior for retrieval fan-out. If the JIRA search times out but the logbook and the vector store both answered, I would rather rerank the two good sets than fail the whole request over one flaky source.</p>\n<p>Pass <code class=\"language-text\">return_exceptions=True</code> and <code class=\"language-text\">gather</code> stops short-circuiting. Instead of raising, it puts each exception into the results list in that call’s slot. You get one entry per input, some values and some exceptions, and you decide what to do with each. In this version, a failed source is logged and skipped:</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\">results <span class=\"token operator\">=</span> <span class=\"token keyword\">await</span> asyncio<span class=\"token punctuation\">.</span>gather<span class=\"token punctuation\">(</span>\n    search_bm25<span class=\"token punctuation\">(</span>query<span class=\"token punctuation\">)</span><span class=\"token punctuation\">,</span>\n    search_vectors<span class=\"token punctuation\">(</span>query<span class=\"token punctuation\">)</span><span class=\"token punctuation\">,</span>\n    search_jira<span class=\"token punctuation\">(</span>query<span class=\"token punctuation\">)</span><span class=\"token punctuation\">,</span>\n    return_exceptions<span class=\"token operator\">=</span><span class=\"token boolean\">True</span><span class=\"token punctuation\">,</span>\n<span class=\"token punctuation\">)</span>\n\nhits <span class=\"token operator\">=</span> <span class=\"token punctuation\">[</span><span class=\"token punctuation\">]</span>\n<span class=\"token keyword\">for</span> source<span class=\"token punctuation\">,</span> r <span class=\"token keyword\">in</span> <span class=\"token builtin\">zip</span><span class=\"token punctuation\">(</span><span class=\"token punctuation\">(</span><span class=\"token string\">\"bm25\"</span><span class=\"token punctuation\">,</span> <span class=\"token string\">\"vectors\"</span><span class=\"token punctuation\">,</span> <span class=\"token string\">\"jira\"</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">,</span> results<span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n    <span class=\"token keyword\">if</span> <span class=\"token builtin\">isinstance</span><span class=\"token punctuation\">(</span>r<span class=\"token punctuation\">,</span> Exception<span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n        log<span class=\"token punctuation\">.</span>warning<span class=\"token punctuation\">(</span><span class=\"token string\">\"source %s failed: %r\"</span><span class=\"token punctuation\">,</span> source<span class=\"token punctuation\">,</span> r<span class=\"token punctuation\">)</span>   <span class=\"token comment\"># degrade, don't die</span>\n    <span class=\"token keyword\">else</span><span class=\"token punctuation\">:</span>\n        hits<span class=\"token punctuation\">.</span>extend<span class=\"token punctuation\">(</span>r<span class=\"token punctuation\">)</span></code></pre></div>\n<p>One thing about the default mode surprises people. When <code class=\"language-text\">gather</code> short-circuits on the first exception, the other tasks are not cancelled. They keep running in the background, detached, and if one of them later fails too, you get a scary “exception was never retrieved” warning from a task nobody is awaiting.</p>\n<p>If you want cleaner semantics, where one failure cancels its siblings and the errors are collected together, use <a href=\"https://docs.python.org/3/library/asyncio-task.html#asyncio.TaskGroup\"><code class=\"language-text\">asyncio.TaskGroup</code></a> (Python 3.11+). It cancels the rest of the group on the first error and raises an <a href=\"https://peps.python.org/pep-0654/\"><code class=\"language-text\">ExceptionGroup</code></a>:</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\"><span class=\"token keyword\">async</span> <span class=\"token keyword\">with</span> asyncio<span class=\"token punctuation\">.</span>TaskGroup<span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span> <span class=\"token keyword\">as</span> tg<span class=\"token punctuation\">:</span>\n    t_bm25    <span class=\"token operator\">=</span> tg<span class=\"token punctuation\">.</span>create_task<span class=\"token punctuation\">(</span>search_bm25<span class=\"token punctuation\">(</span>query<span class=\"token punctuation\">)</span><span class=\"token punctuation\">)</span>\n    t_vectors <span class=\"token operator\">=</span> tg<span class=\"token punctuation\">.</span>create_task<span class=\"token punctuation\">(</span>search_vectors<span class=\"token punctuation\">(</span>query<span class=\"token punctuation\">)</span><span class=\"token punctuation\">)</span>\n<span class=\"token comment\"># both are awaited at the end of the block; if either raised,</span>\n<span class=\"token comment\"># the other is cancelled and the block raises an ExceptionGroup</span></code></pre></div>\n<p>The rule of thumb I use:</p>\n<ul>\n<li><code class=\"language-text\">gather(..., return_exceptions=True)</code> when I want partial results and can degrade around a missing source.</li>\n<li><code class=\"language-text\">TaskGroup</code> when the calls are all-or-nothing, and a failure in one means the others are wasted work worth cancelling.</li>\n</ul>\n<h2>Don’t forget the timeout</h2>\n<p>Overlapping the calls means the request is now only as fast as its slowest branch, which is a problem if one branch can hang. A single retrieval source that stalls for thirty seconds pins the whole <code class=\"language-text\">gather</code> for thirty seconds, and the concurrency you added buys nothing. Put a deadline on each call, so a slow one fails fast and becomes a handled exception instead of a hang. On Python 3.11+, the <code class=\"language-text\">asyncio.timeout</code> context manager reads cleanly:</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\"><span class=\"token keyword\">async</span> <span class=\"token keyword\">def</span> <span class=\"token function\">with_deadline</span><span class=\"token punctuation\">(</span>coro<span class=\"token punctuation\">,</span> seconds<span class=\"token operator\">=</span><span class=\"token number\">2.0</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n    <span class=\"token keyword\">async</span> <span class=\"token keyword\">with</span> asyncio<span class=\"token punctuation\">.</span>timeout<span class=\"token punctuation\">(</span>seconds<span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>   <span class=\"token comment\"># asyncio.wait_for on older Pythons</span>\n        <span class=\"token keyword\">return</span> <span class=\"token keyword\">await</span> coro\n\nresults <span class=\"token operator\">=</span> <span class=\"token keyword\">await</span> asyncio<span class=\"token punctuation\">.</span>gather<span class=\"token punctuation\">(</span>\n    with_deadline<span class=\"token punctuation\">(</span>search_bm25<span class=\"token punctuation\">(</span>query<span class=\"token punctuation\">)</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">,</span>\n    with_deadline<span class=\"token punctuation\">(</span>search_vectors<span class=\"token punctuation\">(</span>query<span class=\"token punctuation\">)</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">,</span>\n    with_deadline<span class=\"token punctuation\">(</span>search_jira<span class=\"token punctuation\">(</span>query<span class=\"token punctuation\">)</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">,</span>\n    return_exceptions<span class=\"token operator\">=</span><span class=\"token boolean\">True</span><span class=\"token punctuation\">,</span>\n<span class=\"token punctuation\">)</span></code></pre></div>\n<p>Now a stuck source raises <code class=\"language-text\">TimeoutError</code> after two seconds and lands in the results as an exception, which the loop in the previous section already knows how to skip. Bounded concurrency, per-call deadlines, and partial-failure handling are the three things that turn a <code class=\"language-text\">gather</code> from a demo into something you can put on a hot path.</p>\n<h2>Failure modes worth knowing</h2>\n<p><strong><code class=\"language-text\">gather</code> does nothing for CPU-bound work.</strong> It overlaps <em>waiting</em>, not computing. “Calls” that actually parse files or run a local embedding model on the CPU hold the <a href=\"https://docs.python.org/3/glossary.html#term-global-interpreter-lock\">GIL</a> (Python’s global interpreter lock), so they run one at a time no matter how many you gather. That is a process pool problem, covered in the <a href=\"/blog/2026-07-10-fastapi-blocking-event-loop/\">event-loop post</a>, not a <code class=\"language-text\">gather</code> one.</p>\n<p><strong>Results are ordered, completion is not.</strong> <code class=\"language-text\">gather</code> hands results back in input order, so you cannot use it to stream the first answer as soon as it lands. To react to each result the moment it finishes, reach for <a href=\"https://docs.python.org/3/library/asyncio-task.html#asyncio.as_completed\"><code class=\"language-text\">asyncio.as_completed</code></a> instead, which yields futures in completion order.</p>\n<p><strong>Share one client, not one per call.</strong> Creating a fresh <code class=\"language-text\">httpx.AsyncClient</code> inside each coroutine defeats connection pooling, because every call pays for a new TLS handshake. Build one client and pass it in, so the whole fan-out shares one pool. That client’s own pool limit is a second, quieter bound on concurrency, worth setting to match the semaphore.</p>\n<p><strong>A bare <code class=\"language-text\">gather</code> on unbounded input is a latent incident.</strong> It passes every test with ten items and falls over the day production hands it ten thousand. If the width of the fan-out comes from data rather than a fixed list, it needs the semaphore from the start, not after the first rate-limit page.</p>\n<h2>What I would do differently</h2>\n<p>Early on, I treated <code class=\"language-text\">gather</code> as an optimization to sprinkle on later, once something felt slow. That was backwards. The useful habit is to look at each request while writing it and ask which calls actually depend on each other. The independent ones should overlap by default, and it is easier to write the <code class=\"language-text\">gather</code> up front than to untangle a chain of sequential awaits after the fact. The dependency graph of a request is usually shallow: a couple of parallel fetches, then a step that needs all of them. That shape maps straight onto one bounded <code class=\"language-text\">gather</code> feeding one final <code class=\"language-text\">await</code>.</p>\n<p>The other thing I would tell my earlier self: never ship a <code class=\"language-text\">gather</code> without deciding the two policies that come with it, namely how many run at once and what happens when one fails. The happy path is one line. The two seconds you save are real only if a single slow or broken source cannot take the whole request down with it. That is the difference between the fan-out in <a href=\"/project/archi/\">Archi</a> and <a href=\"/project/cloud-canvas-ai/\">CloudCanvasAI</a> feeling fast and it becoming the thing that pages you. That difference lives entirely in the bound and the error handling, not in the <code class=\"language-text\">gather</code> itself.</p>\n<hr>\n<p><em>Diagrams by M. Hassan Ahmed, released under CC0. No external image was used for this post; the figures are original work by the author.</em></p>","frontmatter":{"title":"Fan Out Concurrent LLM Calls with asyncio.gather","date":"2026-08-19T00:00:00.000Z","description":"Awaiting retrieval and LLM calls one by one wastes seconds per request. Here's how to fan them out with asyncio.gather, bound it, and handle partial failures.","thumbnail":{"childImageSharp":{"fluid":{"base64":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAABQAAAALCAIAAADwazoUAAAACXBIWXMAAAsSAAALEgHS3X78AAAB3UlEQVQozyWR6Y6bMBSF+dVJ2IzNDsY2S8gEJguEkDSTqIumI02lblP1/V+lh0T6ZB0f+9x7DZqSF5Cr67D/dTy+H4bfffdjPPzZ9z+xwgQQx/H9dPyLIyme7xGgzUgCSFDU69Oquyx3z1V7bPpr1Y5tf11uz/X6Y9NdALZlO9p+/mDH95RmUgkspgwnY35BvUJ3sjnJdEdAYDWovGkB03DEdJMKc0JqyABsiF9VRafUFsJ0pcGEecNgqJXqDp8c+Lf7dzSdcN1OLa8gUZMtT1l9grDcXCep4fApQwVvuqhsUYJVDYSXLacU4Qinup1YrmLphlfHtBwhTKZg6iTBRDRpfbnxxYbGjeVh5tSgfEqRVMOluRVRt5LpuRCXUlwVPzushKlbMV7I+LZ6+qRWZyduUUu346nuDW2OT2eFrlsL7/BUvjT5VxGMqDWzIhwZNLOjJU2aMBtCvrdZ/sHw70dAQ4eZGTB3UfPPrXpt5AsEZSVMFLWIWOCXJqdd/X1TvfK0T0TrePnMDBFEOJyZPsJZMHSrt3X9LQsP9Db23I6Ik593/7b126P6UstrEm+9sKJejhTQpg5Tk8iPHmOxjrI1CxbTg6fnRDYTqhjycpB5L1VvOvzBDB7uETP4Dyi0SscBzjsMAAAAAElFTkSuQmCC","aspectRatio":1.899441340782123,"src":"/static/8562546a0ec941aaf484051b6f70cfca/40a76/hero.png","srcSet":"/static/8562546a0ec941aaf484051b6f70cfca/c972b/hero.png 340w,\n/static/8562546a0ec941aaf484051b6f70cfca/27625/hero.png 680w,\n/static/8562546a0ec941aaf484051b6f70cfca/40a76/hero.png 1360w,\n/static/8562546a0ec941aaf484051b6f70cfca/ed396/hero.png 2000w","sizes":"(max-width: 1360px) 100vw, 1360px"}}}}}},"pageContext":{"slug":"/2026-08-19-concurrent-llm-calls-asyncio-gather/","previous":"blog/2026-08-15-grouped-query-attention-kv-cache/","next":"blog/2026-08-18-product-quantization-vector-search/"}},"staticQueryHashes":["32046230"]}