{"componentChunkName":"component---src-templates-blog-post-js","path":"/blog/2026-08-31-continuous-batching-llm-serving/","result":{"data":{"site":{"siteMetadata":{"title":"M.Hassan Ahmed","author":"Hassan11196"}},"markdownRemark":{"id":"4eb7370e-f70d-59d2-8466-2debeb48fdcb","excerpt":"You put a model behind an API, send it a handful of requests, and watch GPU utilization sit at something embarrassing like 30%. The card is expensive and mostly…","html":"<p>You put a model behind an API, send it a handful of requests, and watch GPU utilization sit at something embarrassing like 30%. The card is expensive and mostly idle. The first instinct is to batch requests together. That helps, but only up to a point. The reason it stalls is subtle, and fixing it is where most of the throughput on a serving GPU actually comes from.</p>\n<p>This post is for engineers running an LLM behind a real endpoint. It explains why a simple batch loop leaves so much performance on the table, what continuous batching does instead, and which settings control it. I ran into this on the GPU side of the <a href=\"/project/cms-workflow-operations/\">CMS workflow tooling</a> I maintained at CERN. I also hit it on <a href=\"/project/archi/\">Archi</a>, our RAG (retrieval-augmented generation) copilot, where answers vary wildly in length and requests arrive whenever an operator has a question.</p>\n<h2>Why batching whole requests stalls</h2>\n<p>Text generation is autoregressive: the model produces one token, appends it to the sequence, and feeds the whole thing back in to produce the next. So a request is not one GPU call. It is hundreds, one per output token. The very first pass, which processes the entire prompt at once, is called <em>prefill</em>. Every step after that produces a single token and is called <em>decode</em>.</p>\n<p>Now imagine the obvious batching scheme: collect a few requests, run them as a batch, and return the results when the batch finishes. The catch is that the requests in a batch almost never finish together. One user asks for a yes/no answer and is done in 20 tokens. Another asks for a full explanation that runs for 800.</p>\n<p>In this scheme, the short request can’t leave the batch early. Its slot stays pinned until the longest request in the batch is done, and for those hundreds of extra steps it contributes nothing. The GPU is doing work, but a chunk of the batch is just padding.</p>\n<p><img src=\"/db510d33194c930b0c99d03b73b96968/static-vs-continuous.svg\" alt=\"The same four GPU slots over the same wall-clock window. In static batching, requests A, B and C finish early but their slots stay idle (hatched) until the longest request D drains the batch, so only four requests complete. In continuous batching each slot is refilled the moment a request finishes, so nine requests complete in the same window\"></p>\n<p>This is head-of-line blocking: requests stuck behind the slowest one in the batch. It gets worse as the spread in output lengths grows. You can’t batch your way around it by grouping similar requests, because output length is unpredictable. You don’t know how long a request will run until it stops.</p>\n<h2>Iteration-level scheduling: reschedule between every token</h2>\n<p>The fix came from a 2022 paper, <a href=\"https://www.usenix.org/conference/osdi22/presentation/yu\">Orca</a>, which changed where the scheduler sits. Instead of scheduling a whole request at a time, it schedules one <em>iteration</em> at a time, where an iteration is a single forward pass of the model. The scheduler runs between every token.</p>\n<p>That sounds like a small change, and it is not. Once the scheduler can decide on every step, three things happen the moment a request emits its stop token:</p>\n<ul>\n<li>the request leaves the batch,</li>\n<li>its memory frees up,</li>\n<li>a request waiting in the queue takes the empty slot on the very next step.</li>\n</ul>\n<p>Nobody waits for the slowest sequence any more. The batch’s membership changes constantly, which is why the technique is usually called <em>continuous batching</em>. (Orca’s own term is iteration-level scheduling; the industry mostly says continuous batching.)</p>\n<p><img src=\"/09563ff330211b36e08a090b7ecff662/scheduler-loop.svg\" alt=\"One iteration of an iteration-level scheduler. Between tokens the scheduler evicts finished requests, frees their KV blocks, admits waiting requests if memory allows, and builds the next batch; the GPU then runs exactly one forward pass, and the loop repeats before the next token\"></p>\n<p>The loop underneath is simple to picture. Each pass evicts finished requests and gives back their memory, admits waiting requests while they fit, and then runs exactly one forward pass:</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\">running <span class=\"token operator\">=</span> <span class=\"token punctuation\">[</span><span class=\"token punctuation\">]</span>          <span class=\"token comment\"># requests currently in the batch</span>\nwaiting <span class=\"token operator\">=</span> deque<span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span>     <span class=\"token comment\"># requests that have arrived but not started</span>\n\n<span class=\"token keyword\">while</span> running <span class=\"token keyword\">or</span> waiting<span class=\"token punctuation\">:</span>\n    <span class=\"token comment\"># 1. drop anything that hit its stop token or max length</span>\n    finished <span class=\"token operator\">=</span> <span class=\"token punctuation\">[</span>r <span class=\"token keyword\">for</span> r <span class=\"token keyword\">in</span> running <span class=\"token keyword\">if</span> r<span class=\"token punctuation\">.</span>is_done<span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">]</span>\n    <span class=\"token keyword\">for</span> r <span class=\"token keyword\">in</span> finished<span class=\"token punctuation\">:</span>\n        r<span class=\"token punctuation\">.</span>release_kv_blocks<span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span>     <span class=\"token comment\"># give the memory back</span>\n        r<span class=\"token punctuation\">.</span>return_to_client<span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span>\n    running <span class=\"token operator\">=</span> <span class=\"token punctuation\">[</span>r <span class=\"token keyword\">for</span> r <span class=\"token keyword\">in</span> running <span class=\"token keyword\">if</span> <span class=\"token keyword\">not</span> r<span class=\"token punctuation\">.</span>is_done<span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">]</span>\n\n    <span class=\"token comment\"># 2. pull in new work while there is room on the card</span>\n    <span class=\"token keyword\">while</span> waiting <span class=\"token keyword\">and</span> can_admit<span class=\"token punctuation\">(</span>waiting<span class=\"token punctuation\">[</span><span class=\"token number\">0</span><span class=\"token punctuation\">]</span><span class=\"token punctuation\">,</span> running<span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n        running<span class=\"token punctuation\">.</span>append<span class=\"token punctuation\">(</span>waiting<span class=\"token punctuation\">.</span>popleft<span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">)</span>\n\n    <span class=\"token comment\"># 3. run exactly ONE forward pass over the current batch</span>\n    step<span class=\"token punctuation\">(</span>running<span class=\"token punctuation\">)</span>     <span class=\"token comment\"># prefill for new members, one decode token for the rest</span></code></pre></div>\n<p>Notice that step 3 mixes two kinds of work in one pass: new members get their prefill while everyone else gets one decode token. Real engines like <a href=\"https://blog.vllm.ai/2023/06/20/vllm.html\">vLLM</a> and Hugging Face’s Text Generation Inference (TGI) are far more careful than this, but the shape is right: evict, admit, run one step, repeat.</p>\n<p>Anyscale benchmarked this against the request-at-a-time approach. They reported up to <a href=\"https://www.anyscale.com/blog/continuous-batching-llm-inference\">23x more throughput</a> on some workloads, with lower median latency at the same time, because requests stop queueing behind unrelated long ones.</p>\n<h2>Two details that make it harder than the loop suggests</h2>\n<p>Two details make the naive picture wrong. Both are worth knowing before you tune anything.</p>\n<h3>Attention can’t be batched like the rest of the model</h3>\n<p>Batching a transformer isn’t uniform across the model. Attention for each request depends on that request’s own past tokens, which live in its own <a href=\"/blog/2026-07-12-llm-kv-cache-gpu-memory/\">KV cache</a> (the stored keys and values from earlier tokens). So you can’t naively stack sequences of different lengths into one matrix for the attention step, the way you can for the big feed-forward layers.</p>\n<p>Orca’s answer was <em>selective batching</em>: batch the operations that don’t care about sequence shape, and handle attention per sequence. That is what lets requests at completely different positions share a single forward pass at all.</p>\n<h3>Memory has to be managed underneath</h3>\n<p>The second detail is memory, and this is the one that bites in production. When you admit a new request, you commit to holding its KV cache for as long as it runs, and that cache <a href=\"/blog/2026-07-12-llm-kv-cache-gpu-memory/\">grows with every token</a>. Admit too many requests and a batch that fit a moment ago runs out of VRAM mid-generation.</p>\n<p>So continuous batching needs a memory manager underneath it. The manager has to pack caches tightly and, when the card fills, temporarily evict a request and bring it back later. In vLLM that manager is <a href=\"https://arxiv.org/abs/2309.06180\">PagedAttention</a>, which stores the cache in fixed-size blocks so space can be handed out and reclaimed on demand. Continuous batching is the scheduler; paged memory is what makes the scheduler’s decisions safe.</p>\n<h2>The prefill-versus-decode collision</h2>\n<p>This is the failure mode that surprises people most. Prefill and decode have very different shapes. A decode step does one token per request and is cheap. A prefill processes the entire prompt at once and, for a long prompt, is a heavy burst of compute.</p>\n<p>Suppose your scheduler runs a new request’s full prefill in a single iteration. That iteration takes a long time, and every request already streaming tokens to a user goes quiet while it runs. The user watching text appear sees it freeze mid-sentence. On a busy server with long prompts, the stream stutters instead of flowing smoothly.</p>\n<p><img src=\"/cc098455ed15233b7e9a8c796d0017ca/chunked-prefill.svg\" alt=\"Why a new request can stutter everyone already streaming. When an 8k-token prefill runs in one step, no decode tokens are emitted and every open stream freezes. Chunked prefill splits that prefill across several iterations, interleaving it with decode steps so tokens keep flowing at a steady rate\"></p>\n<p>The common fix is <em>chunked prefill</em>. It splits a long prompt’s prefill into smaller pieces and interleaves them with everyone else’s decode steps, so ongoing streams keep getting tokens instead of stalling for one giant step. You trade a slightly slower prefill for a much steadier time-between-tokens across all active requests. For an interactive product, that is usually the right trade. vLLM exposes this directly, and the <a href=\"https://arxiv.org/abs/2403.02310\">Sarathi-Serve paper</a> explains the design, framing it as an explicit throughput-versus-latency knob.</p>\n<h2>The two vLLM settings that matter</h2>\n<p>If you run vLLM, two <a href=\"https://docs.vllm.ai/en/latest/serving/engine_args.html\">engine arguments</a> control most of the behavior above:</p>\n<ul>\n<li><code class=\"language-text\">max_num_seqs</code> caps how many requests can be in the running batch at once. Higher means more throughput and more KV memory pressure.</li>\n<li><code class=\"language-text\">max_num_batched_tokens</code> caps how many tokens (prefill plus decode) one iteration may process. This is the lever that tames prefill bursts. Lower it and long prefills get chunked more aggressively, which smooths latency at some cost to raw prefill speed.</li>\n</ul>\n<p>There is no single correct setting. An offline batch job wants both cranked up for throughput. An interactive chat endpoint wants <code class=\"language-text\">max_num_batched_tokens</code> low enough that no single prefill can freeze the streams. Either way, these are the two numbers to reach for, and they trade the same two quantities, throughput and latency, against each other every time.</p>\n<h2>Failure modes worth knowing</h2>\n<ul>\n<li><strong>Out-of-memory (OOM) errors under load, not at startup.</strong> The batch is dynamic, so a server can pass every smoke test on short prompts and then fall over when a burst of long requests gets admitted together. Load-test with the length distribution you actually expect, not toy prompts.</li>\n<li><strong>Latency spikes from unchunked prefill.</strong> If time-between-tokens looks fine in isolation but users complain about stutter under load, a long prefill sharing the batch is the usual cause. Cap the batched-token budget.</li>\n<li><strong>Throughput that collapses at high concurrency.</strong> Past a point, admitting more requests just means more of them get preempted and swapped in and out of memory. More concurrency is not always more throughput; the sweet spot is set by how much KV cache fits on the card.</li>\n<li><strong>Fairness.</strong> A stream of short requests can keep grabbing freed slots and starve a long one, or the reverse. Most engines schedule roughly first-come-first-served for this reason. If you write your own admission logic, starvation is easy to introduce by accident.</li>\n</ul>\n<h2>What I would check first</h2>\n<p>When a serving GPU is underused, I don’t start by buying a bigger card. I check whether the engine is actually doing continuous batching. A surprising number of homegrown loops still batch a whole request at a time and pin short requests behind long ones. Fixing that is usually the largest win available, the same way <a href=\"/blog/2026-07-12-llm-kv-cache-gpu-memory/\">paging the KV cache</a> is the largest memory win.</p>\n<p>After that, in order:</p>\n<ol>\n<li>Confirm KV memory is paged, so admission is safe.</li>\n<li>Set <code class=\"language-text\">max_num_seqs</code> from the length distribution you measured, not one you guessed.</li>\n<li>Turn on chunked prefill if streams stutter under load.</li>\n</ol>\n<p>None of it is exotic. Continuous batching is one idea (move the scheduler inside the token loop), and almost everything else is bookkeeping to make that idea safe on a finite card.</p>\n<p>I apply the same reasoning to the <a href=\"/blog/2026-06-30-llm-agent-tool-loop/\">tool loops</a> and <a href=\"/blog/2026-06-30-fastapi-sse-streaming-llm/\">streaming endpoints</a> behind Archi. The model is rarely the bottleneck; the scheduling around it usually is. Getting a fully loaded GPU is mostly a matter of never letting a slot sit idle while work is waiting for it.</p>\n<hr>\n<p><em>Diagrams by M. Hassan Ahmed, released under CC0. No external image was used in this post.</em></p>","frontmatter":{"title":"How Continuous Batching Speeds Up LLM Serving","date":"2026-08-31T00:00:00.000Z","description":"A naive LLM server holds finished slots idle until the slowest request in the batch drains. Continuous batching refills them every step. How it works.","thumbnail":{"childImageSharp":{"fluid":{"base64":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAABQAAAALCAIAAADwazoUAAAACXBIWXMAAAsSAAALEgHS3X78AAAB80lEQVQozz2QWW+jMBSFeZnOJBAWgzEYzGawzRoITZu0TZqlM9K8jDTz2v//P+Y2lSp9ujqWfHyPj9YkXZv0QB23x+l4ub/AfFkfrg/X8/ZyWB8v2+v5/nKaT6+b02k+77snFTdwH4zancXurPibGRmY581OTce83fFuX29eQfP2SY3HangW6wPManjh/ZOOC7AAmuEkKze3cYkCYXnc9LiFSx1l3y22sNKlcwNlSyf7nDrKV5jrbm6gTNNtZuHKizov7hBtbqIngUrjPmV9wvrAl4EvKJFARBTwKRxcaUsrNtzCcDk8qaMPAQjazMVmKqYh6Qsn5y6P7CxyCmpnofUBQwUlAsyR6ZWWL5ywdmlr+sIN6iHur6H8zcf33dsfOZ+pmmgTp2Mz/+y3v5r5jcZDiEttaVL4pAXZ8k1UPthEWkTyQF0j9U/dv0+Hv2o+hUL60qVdUD4wuafVzqcdBfPCDFduYROF2YDZGvZDitQXPS47XL6y7jFUI5ElrmKvyoliWGDIjHjoV9piFRgoBwMKGygMIkAQ5lXULSMsGJYZqauwHbOppm1BVOZLIPGqj9iLFYHCVl51q6oETCwSIjv+OMrn9Y1e7FW+FWxUyQTU2aaKBgLmH4YPnelO8sXSYSZKkZs6KAEQ6BuOk3wBR92O/wOC1lN8rolhEwAAAABJRU5ErkJggg==","aspectRatio":1.899441340782123,"src":"/static/3113e5e31d17506d62c29098723d1894/40a76/hero.png","srcSet":"/static/3113e5e31d17506d62c29098723d1894/c972b/hero.png 340w,\n/static/3113e5e31d17506d62c29098723d1894/27625/hero.png 680w,\n/static/3113e5e31d17506d62c29098723d1894/40a76/hero.png 1360w,\n/static/3113e5e31d17506d62c29098723d1894/ed396/hero.png 2000w","sizes":"(max-width: 1360px) 100vw, 1360px"}}}}}},"pageContext":{"slug":"/2026-08-31-continuous-batching-llm-serving/","previous":"blog/2026-09-01-metadata-filtering-rag-vector-search/","next":"blog/2026-09-04-idempotency-keys-fastapi-safe-retries/"}},"staticQueryHashes":["32046230"]}