{"componentChunkName":"component---src-templates-blog-post-js","path":"/blog/2026-07-12-llm-kv-cache-gpu-memory/","result":{"data":{"site":{"siteMetadata":{"title":"M.Hassan Ahmed","author":"Hassan11196"}},"markdownRemark":{"id":"8a72594d-1c63-59d4-8297-7642f0c60106","excerpt":"The first time it happens, it makes no sense. A 7B model that served fine all week starts throwing  in the middle of a request, not at load. The weights fit…","html":"<p>The first time it happens, it makes no sense. A 7B model that served fine all week starts throwing <code class=\"language-text\">CUDA out of memory</code> in the middle of a request, not at load. The weights fit. They fit yesterday, and nothing about the model changed. What changed is the request: someone pasted a long document, or an agent got twelve tool calls deep into a task, and the request that broke was longer than the ones before it.</p>\n<p>The thing eating your VRAM (the GPU’s own memory) is not the weights. It is the KV cache, and unlike the weights it is not a fixed cost. It grows with every token in the conversation, and it grows again with every request you batch together.</p>\n<p>This post is for engineers who serve or run LLMs and want to know exactly what that memory is, why it exists, and which levers actually shrink it. I’ve run into it in two places. One is the long-context side of <a href=\"/project/archi/\">Archi</a>, the RAG copilot for CMS operations, where a single answer can pull in thousands of tokens of logs and docs. The other is agent loops, where the transcript only ever gets longer.</p>\n<h2>Why the KV cache exists</h2>\n<p>A transformer generates text one token at a time. To produce the next token, self-attention compares the current token against every token before it. That comparison runs through three projections of the sequence: queries, keys, and values (<a href=\"https://arxiv.org/abs/1706.03762\">Vaswani et al., “Attention Is All You Need”</a>). Roughly, the current token’s query is matched against the earlier tokens’ keys, and the best matches’ values are blended into the output.</p>\n<p>Here is the part that matters for memory. The keys and values for tokens 1 through N-1 do not change when you generate token N, because they were computed from tokens that are already fixed. Recomputing them at every step is pure waste. Instead, you compute each token’s key and value once, keep them, and reuse them at every later step. That store is the KV cache.</p>\n<p><img src=\"/d7524438167843ff4f9dbe454fa08675/attention-recompute.svg\" alt=\"At each decode step, without a cache you recompute the key and value for every earlier token, which is O(N) work per step and O(N squared) over the sequence; with a cache the earlier keys and values are reused and only the new token&#x27;s column is computed, trading repeated compute for memory you must hold in VRAM\"></p>\n<p>The trade is explicit:</p>\n<ul>\n<li><strong>Without the cache</strong>, each decode step (each step that produces one new token) redoes work for every previous token, so a sequence of length N costs on the order of N² over the full generation.</li>\n<li><strong>With the cache</strong>, each step computes only the new column, which is close to constant extra compute.</li>\n</ul>\n<p>You stop repeating work. In return, you have to hold every past key and value in GPU memory for the whole life of the request. That held memory is the bill.</p>\n<h2>How big the cache gets: the per-token math</h2>\n<p>The cache size is not mysterious. For a single token it is:</p>\n<div class=\"gatsby-highlight\" data-language=\"text\"><pre class=\"language-text\"><code class=\"language-text\">bytes/token = 2 x num_layers x num_kv_heads x head_dim x dtype_bytes</code></pre></div>\n<p>The <code class=\"language-text\">2</code> is there because you store both keys <em>and</em> values. For plain multi-head attention, <code class=\"language-text\">num_kv_heads x head_dim</code> is usually just the model’s hidden size (the width of each token’s vector), so a cleaner form is:</p>\n<div class=\"gatsby-highlight\" data-language=\"text\"><pre class=\"language-text\"><code class=\"language-text\">bytes/token = 2 x num_layers x hidden_size x dtype_bytes</code></pre></div>\n<p>Now put in real numbers. Llama-2-7B has 32 layers and a hidden size of 4096, and in fp16 (16-bit floating point) each value takes 2 bytes:</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\">num_layers   <span class=\"token operator\">=</span> <span class=\"token number\">32</span>\nhidden_size  <span class=\"token operator\">=</span> <span class=\"token number\">4096</span>\ndtype_bytes  <span class=\"token operator\">=</span> <span class=\"token number\">2</span>           <span class=\"token comment\"># fp16</span>\n\nper_token <span class=\"token operator\">=</span> <span class=\"token number\">2</span> <span class=\"token operator\">*</span> num_layers <span class=\"token operator\">*</span> hidden_size <span class=\"token operator\">*</span> dtype_bytes\n<span class=\"token comment\"># = 2 * 32 * 4096 * 2 = 524,288 bytes  -> 512 KB per token</span></code></pre></div>\n<p>That is half a megabyte per token. It looks small until you multiply it by a context length and a batch size, and both of those are linear multipliers. Here is a batch of 8 sequences of 4096 tokens each:</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\">seq_len <span class=\"token operator\">=</span> <span class=\"token number\">4096</span>\nbatch   <span class=\"token operator\">=</span> <span class=\"token number\">8</span>\n\ntotal <span class=\"token operator\">=</span> per_token <span class=\"token operator\">*</span> seq_len <span class=\"token operator\">*</span> batch\n<span class=\"token comment\"># 512 KB * 4096 * 8 = 16 GB</span></code></pre></div>\n<p>That is sixteen gigabytes of cache, on top of roughly 14 GB of fp16 weights. On a 24 GB card that request does not fit, and the failure lands at generation time, not at startup. The chart shows the cache stacking on top of the fixed weights as the batch grows.</p>\n<p><img src=\"/21ac5394a875b800bc07a1558409c16e/kv-cache-growth.svg\" alt=\"On a 24 GB card, model weights take a fixed 14 GB; the KV cache is stacked on top and grows with batch and sequence length. One 4k sequence adds 2 GB and fits, batch 4 adds 8 GB and gets tight at 22 GB, batch 8 adds 16 GB and blows past the 24 GB ceiling into CUDA out of memory\"></p>\n<p>Two things follow from this:</p>\n<ol>\n<li><strong>The weights are the <em>fixed</em> cost and the cache is the <em>variable</em> one.</strong> Capacity planning that only counts weights is planning for the wrong number.</li>\n<li><strong>Throughput and context length pull against each other on the same budget.</strong> Every extra token of context is memory you cannot spend on a bigger batch, and vice versa.</li>\n</ol>\n<h2>Fragmentation wastes even more of the cache</h2>\n<p>Knowing the math for one sequence is not enough. A real serving engine runs many sequences at once, each with a different, unpredictable length. If you reserve a contiguous block big enough for the maximum context of every request, most of that reservation sits empty for the requests that finish early or never grow that long.</p>\n<p>The vLLM team measured this and found that older serving systems wasted <strong>60–80% of KV cache memory</strong> to fragmentation and over-reservation (<a href=\"https://arxiv.org/abs/2309.06180\">Kwon et al., “Efficient Memory Management for LLM Serving with PagedAttention,” SOSP 2023</a>). Their fix, PagedAttention, borrows the idea of paged virtual memory from operating systems. It stores the cache in small fixed-size blocks that do not have to be contiguous, and hands them out on demand. That drops the waste to under 4%. Because the freed memory becomes a bigger batch, it also raises throughput substantially (<a href=\"https://blog.vllm.ai/2023/06/20/vllm.html\">vLLM launch post</a>). If you serve models and are not on an engine that pages the cache, this is usually the single biggest win available to you.</p>\n<h2>Four levers that shrink the cache</h2>\n<p>Every term in <code class=\"language-text\">2 x num_layers x num_kv_heads x head_dim x dtype_bytes</code> is a place to cut. The useful levers each attack a different term, and they stack.</p>\n<p><img src=\"/54d1add268d4f06be366fc458348dcbb/shrink-kv.svg\" alt=\"Four levers to shrink the KV cache, each hitting a different term. Grouped-query attention cuts KV heads for up to 8x on Llama-2-70B and is baked into the model. Quantizing the cache from fp16 to int8 halves the bytes per value at serving time. Paged attention reclaims 60 to 80 percent fragmentation waste down to under 4 percent in the serving engine. Shorter context cuts tokens cached, linearly, in your own application code\"></p>\n<p><strong>Grouped-query attention (GQA)</strong> cuts <code class=\"language-text\">num_kv_heads</code>. Instead of every query head owning its own keys and values, groups of query heads share one set. Llama-2-70B uses 64 query heads but only 8 KV heads, an 8x reduction in cache size, and the <a href=\"https://arxiv.org/abs/2305.13245\">GQA paper</a> reports under 0.5 perplexity loss at that ratio. GQA is a property of the model, not something you turn on at serving time, so it is a reason to prefer a GQA model when memory is tight. Most current open models (Llama 3, Mistral, Qwen) already use it.</p>\n<p><strong>Quantizing the cache</strong> cuts <code class=\"language-text\">dtype_bytes</code>. Storing keys and values in int8 instead of fp16 halves the cache, for a modest and usually acceptable quality cost. Serving engines expose this as a flag. It is independent of GQA, so you get both factors.</p>\n<p><strong>Paging</strong> cuts waste rather than the theoretical size, as described above. For a busy multi-tenant service it is usually the biggest single win.</p>\n<p><strong>Shorter context</strong> cuts <code class=\"language-text\">seq_len</code>, and it is the one lever that lives in your code rather than in the model or the engine. In a RAG system, it means retrieving fewer, tighter chunks instead of stuffing the window. (I wrote about <a href=\"/blog/2026-07-04-measuring-rag-retrieval-quality/\">measuring retrieval quality</a> and <a href=\"/blog/2026-07-08-cross-encoder-reranking-rag/\">reranking</a> partly for this reason: better ranking lets you send less.) In an agent loop, it means summarising or trimming old turns instead of replaying the entire transcript on every step. The cache grows linearly with what you put in the window, so discipline in the window is discipline in VRAM.</p>\n<h2>Failure modes worth knowing</h2>\n<ul>\n<li><strong>OOM (out of memory) at generation, not at load.</strong> The cache is allocated as tokens arrive, so a service can pass every startup check and every short-prompt test, then fall over on the first genuinely long request. Load-test with realistic context lengths, not toy prompts.</li>\n<li><strong>Batch size that only works on average.</strong> If you size the batch for typical requests, a batch that happens to be full of long conversations exceeds the budget. The safe batch size is set by the worst case you admit, not the mean.</li>\n<li><strong>Streaming does not save you.</strong> Sending tokens out over <a href=\"/blog/2026-06-30-fastapi-sse-streaming-llm/\">SSE</a> as they generate is a UX win, but the cache still holds the whole sequence server-side until the request ends. A long stream is a long-lived allocation.</li>\n<li><strong>Agent loops leak context.</strong> Each tool call appends to the transcript, so the cache for a single agentic request climbs the deeper the <a href=\"/blog/2026-06-30-llm-agent-tool-loop/\">tool loop</a> goes. The request that OOMs is often not a big input; it is a long run.</li>\n<li><strong>Prefix caching cuts prefill, not footprint.</strong> Reusing the cache for a shared system prompt speeds up prefill (processing the prompt before the first token), but those tokens still occupy memory. It is a latency optimisation, not a capacity one.</li>\n</ul>\n<h2>Tradeoffs, and what I would check first</h2>\n<p>If I were sizing a deployment again, I would start from the variable cost, not the weights. Decide the longest context you will actually admit, multiply by the batch size you need for throughput, and check that <code class=\"language-text\">weights + KV</code> fits with headroom <em>before</em> picking hardware. That single calculation catches most surprises.</p>\n<p>The order of levers I reach for:</p>\n<ol>\n<li>Pick a GQA model.</li>\n<li>Serve it on an engine that pages the cache.</li>\n<li>Turn on int8 KV quantization if quality holds.</li>\n<li>Only then, spend effort trimming context in the application.</li>\n</ol>\n<p>The first three are close to free. The last one is real work, but it is also where an application engineer has the most direct control. On a long-context RAG product, it is often the difference between fitting on the GPU you have and paying for a bigger one.</p>\n<p>None of this is exotic. The KV cache is the transformer trading repeated compute for stored state. The memory it costs is one short formula multiplied by two numbers you set: how long the context is, and how many requests you batch. Once you can see those numbers, <code class=\"language-text\">CUDA out of memory</code> reads as a budget you overspent rather than a mystery. I did this arithmetic often on the HPC and GPU side of the <a href=\"/project/cms-workflow-operations/\">CMS workflow tooling</a> I maintained, and it is the same check that keeps <a href=\"/project/archi/\">Archi</a> answering long questions without falling off the card.</p>\n<hr>\n<p><em>Diagrams by M. Hassan Ahmed, released under CC0. No external image was used in this post.</em></p>","frontmatter":{"title":"LLM KV Cache: Why GPU Memory Runs Out","date":"2026-07-12T00:00:00.000Z","description":"A long prompt or a long agent run can hit CUDA out of memory, and the KV cache is usually why. How it grows, the per-token math, and how to shrink it.","thumbnail":{"childImageSharp":{"fluid":{"base64":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAABQAAAALCAIAAADwazoUAAAACXBIWXMAAAsSAAALEgHS3X78AAAB2UlEQVQoz0WO63KbMBBG+dUA5qogkEASMverb6FxYgjGaTLt+z9R12Y6nTmjWe1+ZyWlEV0j+lb0VVSf6/NtuI27aRlu8/G6vCzT/gOuH4cZ+tfTchs+T8VQsQbyICo/bP70wAzyMDtl/Zh27/VpSbtLvpuy/gIFNPPdWB7mbvik6dHAGeRBVAxXrFjPWwPFG4dvXPFkRaoVwak7HEbQgQLQbLZBMcTWvnJPPyoTumvU4Zw3Iu622x4FGawwHhnAQTHzU+ZnAU7AV3Sb6XYEAynaRDQIbzUrIkEW8zaJ4btHGaSWK3wvDrDwPOFhiX0JJ/xCeZjM8qRqR8hmDS3GpPtdHOe0/y52X2XHCSMBIyH3I+6HMoxzJvNQZJoVKppJTZcjPyvLw9d0/TMt03B52Q+n/dC1Q1IcnLhGsrF5bbDaCguHpm6YWVjqJlU0i+p2aCLB8jfRz0E94nwkxUTyO7SYwnIKkjccD748O0GlGgG8BxYALxO461boxYMnXrH4ycpZtr9EfYWC14tobiS7ePLVT94dUt9lMME3CciwCWSKaIXCBnBI5dLGizqP9c9Ra5PSJRWiNYxsnGqGvyqAohr+ytMGr6gbb+NQEzELcdONVAP/H/0Lr/wFUNZLWBRwh3EAAAAASUVORK5CYII=","aspectRatio":1.899441340782123,"src":"/static/3532ef527ffd15a7acd37770357c598c/40a76/hero.png","srcSet":"/static/3532ef527ffd15a7acd37770357c598c/c972b/hero.png 340w,\n/static/3532ef527ffd15a7acd37770357c598c/27625/hero.png 680w,\n/static/3532ef527ffd15a7acd37770357c598c/40a76/hero.png 1360w,\n/static/3532ef527ffd15a7acd37770357c598c/ed396/hero.png 2000w","sizes":"(max-width: 1360px) 100vw, 1360px"}}}}}},"pageContext":{"slug":"/2026-07-12-llm-kv-cache-gpu-memory/","previous":"blog/2026-07-09-hnsw-vector-search-explained/","next":"blog/2026-07-11-llm-api-rate-limits-retries-backoff/"}},"staticQueryHashes":["32046230"]}