LLM KV Cache: Why GPU Memory Runs Out
A long prompt or a long agent run can hit CUDA out of memory, and the KV cache is usually why. How it grows, the per-token math, and how to shrink it.
The first time it happens, it makes no sense. A 7B model that served fine all week starts throwing CUDA out of memory in the middle of a request, not at load. The weights fit. They fit yesterday, and nothing about the model changed. What changed is the request: someone pasted a long document, or an agent got twelve tool calls deep into a task, and the request that broke was longer than the ones before it.
The thing eating your VRAM (the GPU’s own memory) is not the weights. It is the KV cache, and unlike the weights it is not a fixed cost. It grows with every token in the conversation, and it grows again with every request you batch together.
This post is for engineers who serve or run LLMs and want to know exactly what that memory is, why it exists, and which levers actually shrink it. I’ve run into it in two places. One is the long-context side of Archi, the RAG copilot for CMS operations, where a single answer can pull in thousands of tokens of logs and docs. The other is agent loops, where the transcript only ever gets longer.
Why the KV cache exists
A transformer generates text one token at a time. To produce the next token, self-attention compares the current token against every token before it. That comparison runs through three projections of the sequence: queries, keys, and values (Vaswani et al., “Attention Is All You Need”). Roughly, the current token’s query is matched against the earlier tokens’ keys, and the best matches’ values are blended into the output.
Here is the part that matters for memory. The keys and values for tokens 1 through N-1 do not change when you generate token N, because they were computed from tokens that are already fixed. Recomputing them at every step is pure waste. Instead, you compute each token’s key and value once, keep them, and reuse them at every later step. That store is the KV cache.
The trade is explicit:
- Without the cache, each decode step (each step that produces one new token) redoes work for every previous token, so a sequence of length N costs on the order of N² over the full generation.
- With the cache, each step computes only the new column, which is close to constant extra compute.
You stop repeating work. In return, you have to hold every past key and value in GPU memory for the whole life of the request. That held memory is the bill.
How big the cache gets: the per-token math
The cache size is not mysterious. For a single token it is:
bytes/token = 2 x num_layers x num_kv_heads x head_dim x dtype_bytesThe 2 is there because you store both keys and values. For plain multi-head attention, num_kv_heads x head_dim is usually just the model’s hidden size (the width of each token’s vector), so a cleaner form is:
bytes/token = 2 x num_layers x hidden_size x dtype_bytesNow put in real numbers. Llama-2-7B has 32 layers and a hidden size of 4096, and in fp16 (16-bit floating point) each value takes 2 bytes:
num_layers = 32
hidden_size = 4096
dtype_bytes = 2 # fp16
per_token = 2 * num_layers * hidden_size * dtype_bytes
# = 2 * 32 * 4096 * 2 = 524,288 bytes -> 512 KB per tokenThat is half a megabyte per token. It looks small until you multiply it by a context length and a batch size, and both of those are linear multipliers. Here is a batch of 8 sequences of 4096 tokens each:
seq_len = 4096
batch = 8
total = per_token * seq_len * batch
# 512 KB * 4096 * 8 = 16 GBThat is sixteen gigabytes of cache, on top of roughly 14 GB of fp16 weights. On a 24 GB card that request does not fit, and the failure lands at generation time, not at startup. The chart shows the cache stacking on top of the fixed weights as the batch grows.
Two things follow from this:
- The weights are the fixed cost and the cache is the variable one. Capacity planning that only counts weights is planning for the wrong number.
- Throughput and context length pull against each other on the same budget. Every extra token of context is memory you cannot spend on a bigger batch, and vice versa.
Fragmentation wastes even more of the cache
Knowing the math for one sequence is not enough. A real serving engine runs many sequences at once, each with a different, unpredictable length. If you reserve a contiguous block big enough for the maximum context of every request, most of that reservation sits empty for the requests that finish early or never grow that long.
The vLLM team measured this and found that older serving systems wasted 60–80% of KV cache memory to fragmentation and over-reservation (Kwon et al., “Efficient Memory Management for LLM Serving with PagedAttention,” SOSP 2023). Their fix, PagedAttention, borrows the idea of paged virtual memory from operating systems. It stores the cache in small fixed-size blocks that do not have to be contiguous, and hands them out on demand. That drops the waste to under 4%. Because the freed memory becomes a bigger batch, it also raises throughput substantially (vLLM launch post). If you serve models and are not on an engine that pages the cache, this is usually the single biggest win available to you.
Four levers that shrink the cache
Every term in 2 x num_layers x num_kv_heads x head_dim x dtype_bytes is a place to cut. The useful levers each attack a different term, and they stack.
Grouped-query attention (GQA) cuts num_kv_heads. Instead of every query head owning its own keys and values, groups of query heads share one set. Llama-2-70B uses 64 query heads but only 8 KV heads, an 8x reduction in cache size, and the GQA paper reports under 0.5 perplexity loss at that ratio. GQA is a property of the model, not something you turn on at serving time, so it is a reason to prefer a GQA model when memory is tight. Most current open models (Llama 3, Mistral, Qwen) already use it.
Quantizing the cache cuts dtype_bytes. Storing keys and values in int8 instead of fp16 halves the cache, for a modest and usually acceptable quality cost. Serving engines expose this as a flag. It is independent of GQA, so you get both factors.
Paging cuts waste rather than the theoretical size, as described above. For a busy multi-tenant service it is usually the biggest single win.
Shorter context cuts seq_len, and it is the one lever that lives in your code rather than in the model or the engine. In a RAG system, it means retrieving fewer, tighter chunks instead of stuffing the window. (I wrote about measuring retrieval quality and reranking partly for this reason: better ranking lets you send less.) In an agent loop, it means summarising or trimming old turns instead of replaying the entire transcript on every step. The cache grows linearly with what you put in the window, so discipline in the window is discipline in VRAM.
Failure modes worth knowing
- OOM (out of memory) at generation, not at load. The cache is allocated as tokens arrive, so a service can pass every startup check and every short-prompt test, then fall over on the first genuinely long request. Load-test with realistic context lengths, not toy prompts.
- Batch size that only works on average. If you size the batch for typical requests, a batch that happens to be full of long conversations exceeds the budget. The safe batch size is set by the worst case you admit, not the mean.
- Streaming does not save you. Sending tokens out over SSE as they generate is a UX win, but the cache still holds the whole sequence server-side until the request ends. A long stream is a long-lived allocation.
- Agent loops leak context. Each tool call appends to the transcript, so the cache for a single agentic request climbs the deeper the tool loop goes. The request that OOMs is often not a big input; it is a long run.
- Prefix caching cuts prefill, not footprint. Reusing the cache for a shared system prompt speeds up prefill (processing the prompt before the first token), but those tokens still occupy memory. It is a latency optimisation, not a capacity one.
Tradeoffs, and what I would check first
If I were sizing a deployment again, I would start from the variable cost, not the weights. Decide the longest context you will actually admit, multiply by the batch size you need for throughput, and check that weights + KV fits with headroom before picking hardware. That single calculation catches most surprises.
The order of levers I reach for:
- Pick a GQA model.
- Serve it on an engine that pages the cache.
- Turn on int8 KV quantization if quality holds.
- Only then, spend effort trimming context in the application.
The first three are close to free. The last one is real work, but it is also where an application engineer has the most direct control. On a long-context RAG product, it is often the difference between fitting on the GPU you have and paying for a bigger one.
None of this is exotic. The KV cache is the transformer trading repeated compute for stored state. The memory it costs is one short formula multiplied by two numbers you set: how long the context is, and how many requests you batch. Once you can see those numbers, CUDA out of memory reads as a budget you overspent rather than a mystery. I did this arithmetic often on the HPC and GPU side of the CMS workflow tooling I maintained, and it is the same check that keeps Archi answering long questions without falling off the card.
Diagrams by M. Hassan Ahmed, released under CC0. No external image was used in this post.