Continuous Batching for LLM Inference

Static batching leaves the GPU idle when requests finish at different steps. How continuous batching schedules LLM inference per token to raise throughput.

Here is a number that looks wrong the first time you see it. You put a GPU behind an LLM, send it one request at a time, and watch utilization sit somewhere around 10 percent. The card is expensive, it is warm, and it is mostly doing nothing. Batch a few requests together and the number climbs, but not as much as you’d hope. The requests don’t finish at the same time, so the short ones end up waiting on the long ones.

That waiting is the whole problem, and continuous batching is the fix most modern serving engines settle on. This post is for engineers who run or serve LLMs. It covers:

  • why a naive batch wastes the GPU,
  • what “batch at the iteration instead of the request” actually means,
  • where the technique still bites.

It is the throughput companion to an earlier post on why the KV cache runs your GPU out of memory. The same serving loop sits under both.

Why one request at a time wastes the card

A GPU is built to do a lot of arithmetic in parallel. A single decode step for one sequence (generating its next token) is a small matrix multiply against a large set of weights. That work doesn’t come close to filling the hardware. The GPU reads the weights from memory, uses them once, and then idles until the next step. For a single stream, you spend most of your time moving weights, not multiplying them, which makes this a memory-bandwidth problem, not a compute one.

The way out is to run many sequences through the same weight read. If eight requests each need their next token, you can stack them and do one bigger multiply that reuses the weights you already pulled from memory. Same weight traffic, eight times the useful output. That is why inference uses batching at all, and why serving one request at a time leaves most of the card on the floor.

Static batching, and where it stalls

The obvious way to batch is to collect a group of requests, run them together until they’re all done, then take the next group. Call it static batching. It works, and it beats one-at-a-time, but it has a failure built in: generation lengths are not equal, and you don’t know them in advance.

Consider two users. One asks for a yes-or-no answer that finishes in six tokens. Another pastes a stack trace and asks for a rewrite that runs eight hundred. In a static batch they start together, and the batch can’t return anyone or admit anyone new until its longest sequence emits its stop token (the token that ends generation). The short request finished long ago. Its slot is still there, still computed every step, producing padding nobody wants.

Static batching timeline. Four requests start together in one batch. Request A finishes at step 3, C at step 5, D at step 7, but B runs to step 10. Because the batch only frees when the slowest request is done, the slots for A, C, and D keep taking GPU work as padding after they finish, shown as wasted grey area, until step 10 when all slots free at once.

The grey in that picture is the cost, and it scales with how uneven your traffic is. LLM traffic is very uneven.

Static batching also hurts latency. A request that could have been admitted at step 3 has to wait for the current batch to fully drain, so your queue backs up behind the longest generation in flight. High-variance workloads are exactly where static batching disappoints, and a chat or copilot workload is nothing but high variance.

The fix: choose the batch every iteration

The fix came out of a serving system called Orca (Yu et al., “Orca: A Distributed Serving System for Transformer-Based Generative Models,” OSDI 2022), which calls it iteration-level scheduling. The industry mostly calls it continuous batching, or in-flight batching.

The change is small to state and large in effect. Instead of picking a batch once and running it to completion, the scheduler decides the batch at the start of every decode step:

  1. After each step, any sequence that just emitted its stop token leaves the batch and goes back to its caller.
  2. A waiting request from the queue fills the slot it vacated.

The batch becomes a living set that changes shape every iteration, not a fixed group you commit to.

Continuous batching timeline. Same four requests in four slots. When request A finishes at step 3, queued request E immediately takes its slot for the rest of the run. When C finishes at step 5, request F takes over. When D finishes at step 7, request G takes over. Slot 2 runs request B the whole time. There is no grey wasted area; every slot does real work at every step.

Nothing is padded. A finished request leaves the moment it is done, not at the end of whatever batch it happened to share, so its latency drops. Its slot goes straight to work on the next request instead of grinding out padding. The GPU stays full because the scheduler keeps it full, step by step.

What the serving loop actually does

Stripped of the CUDA and the attention kernels, the scheduler loops over steps, not over requests. Here it is in rough Python. Each pass admits new work, advances every running sequence by one token, and retires whatever finished:

running = []          # sequences currently in the batch
waiting = Queue()     # arrived, not yet admitted

while True:
    # 1. Admit new work if there is room in the batch and KV budget.
    while len(running) < MAX_BATCH and can_fit_kv(waiting.peek()):
        running.append(waiting.get())

    # 2. One forward pass over the whole current batch: each
    #    sequence advances by exactly one token this iteration.
    logits = model.step(running)          # decode step for all at once
    for seq in running:
        seq.append(sample(logits[seq.id]))

    # 3. Retire anyone who just finished; free their KV blocks.
    done = [s for s in running if s.is_finished()]
    for seq in done:
        respond(seq)                      # return to the caller now
        free_kv(seq)
    running = [s for s in running if not s.is_finished()]

Two details in that loop do most of the work.

Step 1 is admission control: deciding whether a waiting request may join. It doesn’t only check a batch-size cap. It also checks whether there is enough KV cache memory to hold the new sequence, because every admitted request reserves cache that grows with its length. This is exactly where continuous batching meets paged memory management, which is why the two show up together in the same engines.

Step 3 is where the win lands. The loop answers a finished sequence and frees its cache inside the same iteration. Both the response and the memory come back without waiting on the rest of the batch.

Why it pairs with PagedAttention

Continuous batching admits and retires sequences constantly, so the KV cache is allocated and freed in small pieces at a high rate. Suppose each sequence’s cache had to be one contiguous reservation sized for the maximum possible length. That churn would fragment memory badly, and you couldn’t pack the batch tightly. The two techniques need each other.

vLLM’s PagedAttention is the answer that stuck. It stores the KV cache in small fixed-size blocks that need not be contiguous, and hands them out on demand, the way an operating system pages virtual memory. The vLLM team measured older serving systems wasting 60 to 80 percent of KV memory to fragmentation and over-reservation. Paging drops that to under 4 percent (Kwon et al., “Efficient Memory Management for LLM Serving with PagedAttention,” SOSP 2023).

That reclaimed memory is more than tidiness. It means a bigger batch, and a bigger batch means more throughput. I went through the per-token cache math in the KV cache post; that number governs how many sequences you can keep running at once.

How much throughput you get, and why

The gain is real and large, but it depends heavily on how uneven your generation lengths are. Treat any single figure as workload-specific, not a guarantee. The Anyscale team benchmarked continuous batching against naive and static batching. They reported up to a 23x throughput improvement over a naive Hugging Face pipeline, with the gap widening as the variance in output length grew (Anyscale, “How continuous batching enables 23x throughput in LLM inference”). The Orca paper reported large throughput gains over FasterTransformer at matched latency, for the same reason.

Every one of those numbers has the same mechanism behind it: you stop paying for padding, and you keep the GPU busy with real tokens. If your outputs were all the same length, continuous batching would barely beat static batching, because there would be no stragglers to wait on. It looks impressive precisely because production traffic is lumpy.

Every major serving stack now does this by default. vLLM, Hugging Face’s Text Generation Inference (TGI), and NVIDIA’s TensorRT-LLM all run continuous or in-flight batching. In practice you get it by picking one of them, not by writing the scheduler yourself (vLLM, TGI docs).

Failure modes and tradeoffs

Turning it on is easy. Understanding what it does to your tail latency (the slowest few percent of requests) is the part people skip.

  • Prefill stalls decode. Prefill is the big, compute-heavy pass over a new request’s prompt, and when the scheduler admits a request mid-flight, that pass can hold up the decode step for every sequence already running. Users who were happily streaming see a stutter the instant someone with a huge prompt joins. The mitigation is chunked prefill, which breaks a long prompt into pieces interleaved with ongoing decodes so no single admission freezes the batch. Most engines expose it as a flag, so know it exists before you debug a mysterious latency spike.
  • Throughput and per-request latency pull apart. A bigger batch gives more tokens per second overall but makes each stream slightly slower, since every sequence shares the step. Tuning the max batch size means choosing a point on that curve. The right point for a batch job is not the right point for an interactive copilot.
  • KV memory, not batch count, limits admission. You can set a generous max batch size and still admit far fewer sequences, because the cache filled up first. When throughput plateaus below the batch limit you configured, memory is usually the real ceiling, and the same shrink-the-cache levers from the KV post apply.
  • Preemption is real. When memory runs tight, some engines evict a running sequence, drop its cache, and recompute it later. That keeps the system alive under load, but it shows up as the occasional request that mysteriously took much longer. If your p99 (99th-percentile) latency has a fat tail under pressure, preemption is a prime suspect.
  • It doesn’t fix your rate limits. Continuous batching raises the throughput of one replica. It does nothing for an upstream model API you may be calling, where the discipline is retries and backoff instead. Different layer, different problem.

What I would check first

If I were standing up a serving deployment today, I wouldn’t hand-roll any of this. I’d pick an engine that already does continuous batching plus paged KV memory, then spend my time on two numbers:

  • the longest context I will actually admit;
  • the max batch size that keeps p99 latency inside my budget.

Together those set the memory ceiling and the latency floor, and almost every throughput surprise traces back to one of them.

This matters to me because of two workloads. One is the GPU side of the CMS workflow tooling I maintained. The other is the long-context answers behind Archi, the RAG (retrieval-augmented generation) copilot for CMS operations, where a single reply can pull thousands of tokens of logs and docs. On workloads like these, generation lengths are all over the place. That is exactly the case where scheduling the batch every iteration instead of every request is the difference between one busy GPU and four idle ones.


Diagrams by M. Hassan Ahmed, released under CC0. No external image was used in this post.