{"componentChunkName":"component---src-templates-blog-post-js","path":"/blog/2026-08-30-prompt-caching-claude-api/","result":{"data":{"site":{"siteMetadata":{"title":"M.Hassan Ahmed","author":"Hassan11196"}},"markdownRemark":{"id":"9f12d0f7-13dc-557e-8a4c-b51c0ba975ce","excerpt":"Most of the tokens in an LLM request are the same tokens you sent last time. Take CloudCanvasAI, where Claude edits  and  files inside a per-user sandbox. Every…","html":"<p>Most of the tokens in an LLM request are the same tokens you sent last time. Take <a href=\"/project/cloud-canvas-ai/\">CloudCanvasAI</a>, where Claude edits <code class=\"language-text\">.docx</code> and <code class=\"language-text\">.pptx</code> files inside a per-user sandbox. Every turn carries the same system prompt, the same tool descriptions, and a growing conversation history. The only genuinely new part is the last thing the user typed. Re-sending the rest on each request means paying to re-read a document that hasn’t changed.</p>\n<p>Prompt caching fixes that. You mark a stable chunk of the request once, and Anthropic keeps its processed form around for a few minutes. The next request that starts with the same bytes skips that work.</p>\n<p>This post is for engineers already calling the <a href=\"https://platform.claude.com/docs/en/api/overview\">Claude API</a> who want to cut cost and time-to-first-token on repeated calls without changing what the model sees. I’ll cover what actually gets cached, the exact price of a hit versus a miss, and the handful of ways people break caching without noticing.</p>\n<h2>What actually gets cached: a prefix</h2>\n<p>The cache is a <strong>prefix match</strong>. Anthropic renders every request into one ordered sequence: <code class=\"language-text\">tools</code> first, then <code class=\"language-text\">system</code>, then <code class=\"language-text\">messages</code>. Caching keys off that sequence from the start up to a marker you place, and everything before the marker is the cached prefix. Change a single byte anywhere inside that prefix and the match fails from the point of the change onward, so the model reprocesses the rest.</p>\n<p>That ordering is the whole game. Stable content has to come first. Volatile content (the new user question, a per-request id, anything with a timestamp) has to come after the last marker. Get the order wrong and you cache nothing.</p>\n<p><img src=\"/5304fcf54278b72694b271a9e6961b69/cache-anatomy.svg\" alt=\"Anatomy of a cached request: blocks render tools, then system, then messages, with a cache_control breakpoint separating the reused prefix from the new turn\"></p>\n<p>The marker is a <code class=\"language-text\">cache_control</code> entry of type <code class=\"language-text\">ephemeral</code>. It doesn’t cache a specific block in isolation. It caches the whole prefix <em>up to and including</em> that block.</p>\n<h2>Turning it on</h2>\n<p>The simplest form marks a stable system prompt. In the example below, the style guide and format rules are the same on every request, so they go in <code class=\"language-text\">system</code> with a <code class=\"language-text\">cache_control</code> marker. The user’s turn stays in <code class=\"language-text\">messages</code>, after the marker, where it belongs.</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\"><span class=\"token keyword\">from</span> anthropic <span class=\"token keyword\">import</span> Anthropic\n\nclient <span class=\"token operator\">=</span> Anthropic<span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span>\n\nSYSTEM_DOCS <span class=\"token operator\">=</span> load_style_guide<span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span>  <span class=\"token comment\"># ~8k tokens, identical every request</span>\n\nresp <span class=\"token operator\">=</span> client<span class=\"token punctuation\">.</span>messages<span class=\"token punctuation\">.</span>create<span class=\"token punctuation\">(</span>\n    model<span class=\"token operator\">=</span><span class=\"token string\">\"claude-opus-5\"</span><span class=\"token punctuation\">,</span>\n    max_tokens<span class=\"token operator\">=</span><span class=\"token number\">4096</span><span class=\"token punctuation\">,</span>\n    system<span class=\"token operator\">=</span><span class=\"token punctuation\">[</span>\n        <span class=\"token punctuation\">{</span>\n            <span class=\"token string\">\"type\"</span><span class=\"token punctuation\">:</span> <span class=\"token string\">\"text\"</span><span class=\"token punctuation\">,</span>\n            <span class=\"token string\">\"text\"</span><span class=\"token punctuation\">:</span> SYSTEM_DOCS<span class=\"token punctuation\">,</span>\n            <span class=\"token string\">\"cache_control\"</span><span class=\"token punctuation\">:</span> <span class=\"token punctuation\">{</span><span class=\"token string\">\"type\"</span><span class=\"token punctuation\">:</span> <span class=\"token string\">\"ephemeral\"</span><span class=\"token punctuation\">}</span><span class=\"token punctuation\">,</span>\n        <span class=\"token punctuation\">}</span>\n    <span class=\"token punctuation\">]</span><span class=\"token punctuation\">,</span>\n    messages<span class=\"token operator\">=</span><span class=\"token punctuation\">[</span><span class=\"token punctuation\">{</span><span class=\"token string\">\"role\"</span><span class=\"token punctuation\">:</span> <span class=\"token string\">\"user\"</span><span class=\"token punctuation\">,</span> <span class=\"token string\">\"content\"</span><span class=\"token punctuation\">:</span> user_turn<span class=\"token punctuation\">}</span><span class=\"token punctuation\">]</span><span class=\"token punctuation\">,</span>\n<span class=\"token punctuation\">)</span></code></pre></div>\n<p>The default cache lifetime, or TTL (time to live), is five minutes, and every hit inside that window resets the clock. If your traffic is bursty enough that requests land more than five minutes apart, you can ask for a one-hour lifetime by adding a <code class=\"language-text\">ttl</code> to the marker:</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\"><span class=\"token string\">\"cache_control\"</span><span class=\"token punctuation\">:</span> <span class=\"token punctuation\">{</span><span class=\"token string\">\"type\"</span><span class=\"token punctuation\">:</span> <span class=\"token string\">\"ephemeral\"</span><span class=\"token punctuation\">,</span> <span class=\"token string\">\"ttl\"</span><span class=\"token punctuation\">:</span> <span class=\"token string\">\"1h\"</span><span class=\"token punctuation\">}</span><span class=\"token punctuation\">,</span></code></pre></div>\n<h3>The minimum prefix length</h3>\n<p>One detail trips people up: there’s a minimum prefix length. Prompts shorter than the model’s floor won’t cache at all, and you get <strong>no error</strong>. The request just runs uncached.</p>\n<p>The floor depends on the model: 512 tokens on Claude Opus 5, 1,024 on Sonnet, and higher on some older models. The current table is in the <a href=\"https://platform.claude.com/docs/en/build-with-claude/prompt-caching\">prompt caching docs</a>. If you’re caching a short prompt and seeing no savings, this is usually why.</p>\n<h3>Caching a growing conversation</h3>\n<p>An agent loop is where caching earns the most, because the history only grows. The trick is to move the breakpoint forward each turn. Put it on the last message that won’t change again, so the cached prefix keeps extending as the conversation does.</p>\n<p>You get up to four breakpoints per request. A common layout uses:</p>\n<ul>\n<li>one on the tool definitions,</li>\n<li>one on the system prompt,</li>\n<li>one that trails the conversation.</li>\n</ul>\n<p>The snippet below sets that trailing breakpoint. It marks the last block of the existing history, then appends the new user turn after it:</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\">messages <span class=\"token operator\">=</span> build_history<span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span>          <span class=\"token comment\"># everything so far</span>\nmessages<span class=\"token punctuation\">[</span><span class=\"token operator\">-</span><span class=\"token number\">1</span><span class=\"token punctuation\">]</span><span class=\"token punctuation\">[</span><span class=\"token string\">\"content\"</span><span class=\"token punctuation\">]</span><span class=\"token punctuation\">[</span><span class=\"token operator\">-</span><span class=\"token number\">1</span><span class=\"token punctuation\">]</span><span class=\"token punctuation\">[</span><span class=\"token string\">\"cache_control\"</span><span class=\"token punctuation\">]</span> <span class=\"token operator\">=</span> <span class=\"token punctuation\">{</span><span class=\"token string\">\"type\"</span><span class=\"token punctuation\">:</span> <span class=\"token string\">\"ephemeral\"</span><span class=\"token punctuation\">}</span>\nmessages<span class=\"token punctuation\">.</span>append<span class=\"token punctuation\">(</span><span class=\"token punctuation\">{</span><span class=\"token string\">\"role\"</span><span class=\"token punctuation\">:</span> <span class=\"token string\">\"user\"</span><span class=\"token punctuation\">,</span> <span class=\"token string\">\"content\"</span><span class=\"token punctuation\">:</span> new_turn<span class=\"token punctuation\">}</span><span class=\"token punctuation\">)</span>  <span class=\"token comment\"># uncached tail</span></code></pre></div>\n<p>Because the new user turn sits after the marker, it never invalidates the cached history behind it.</p>\n<h2>What a hit and a miss actually cost</h2>\n<p>The numbers are what make this worth doing. Relative to the base input token price, the <a href=\"https://platform.claude.com/docs/en/build-with-claude/prompt-caching\">official multipliers</a> are:</p>\n<ul>\n<li>Writing tokens into the five-minute cache costs <strong>1.25×</strong>.</li>\n<li>Reading them back on a hit costs <strong>0.1×</strong>.</li>\n<li>The one-hour cache writes at <strong>2×</strong>, and reads still cost 0.1×.</li>\n</ul>\n<p>Work through the break-even. Without caching, two identical-prefix requests cost <code class=\"language-text\">1× + 1×</code>. With caching, the first writes the prefix (<code class=\"language-text\">1.25×</code>) and the second reads it (<code class=\"language-text\">0.1×</code>), so the same two calls cost <code class=\"language-text\">1.25× + 0.1× = 1.35×</code>.</p>\n<p>A <strong>single</strong> reuse inside the window already comes out ahead, and the gap only widens from there. Ten reads is <code class=\"language-text\">1.25× + (10 × 0.1×) = 2.25×</code> against <code class=\"language-text\">11×</code> uncached. The extra quarter you pay on the write is bought back on the first hit.</p>\n<p>The other half of the payoff is latency. A cache read skips the prefill work (the model’s first pass over the input) for that prefix, so time-to-first-token drops on long prompts. In a streaming UI like CloudCanvasAI’s split-panel view, that’s the difference between the response appearing to start immediately and a visible pause while the model re-reads context it already had.</p>\n<p>If you want the background on why prefill is the expensive part, I wrote about <a href=\"/blog/2026-07-12-llm-kv-cache-gpu-memory/\">the KV cache and GPU memory</a> separately. Prompt caching is essentially Anthropic holding that computed state for you between requests.</p>\n<h2>Where it quietly breaks</h2>\n<p>Every failure mode here has the same symptom: the request works, the answer looks fine, and you’re paying full price because nothing hit the cache. None of these raise an error.</p>\n<p><strong>A timestamp in the prefix.</strong> The classic one. Something like <code class=\"language-text\">f&quot;Current time: {datetime.now()}&quot;</code> at the top of the system prompt changes the first bytes on every request, so the prefix never matches. If the model genuinely needs the time, put it <em>after</em> the last breakpoint, in the user turn.</p>\n<p><strong>Non-deterministic serialization.</strong> Say part of your prefix is JSON built from a dict. Python doesn’t guarantee the same key order across all the code paths that might construct it, and a reordered key is a different byte string. Serialize anything that lands in a cached block with <code class=\"language-text\">json.dumps(obj, sort_keys=True)</code>.</p>\n<p><strong>A tool list that shifts.</strong> Because <code class=\"language-text\">tools</code> renders first, reordering tools or regenerating their descriptions invalidates everything after them, the system prompt and the whole conversation included. Build the tool array in a fixed order and keep the descriptions stable. I ran into a version of this with <a href=\"/project/llm-dev-mate/\">LLM DevMate</a>. A tool set that’s stable across a session caches well. One rebuilt per request from a set or a dict throws the order away and quietly costs you the cache.</p>\n<p><strong>Letting the window lapse.</strong> On the five-minute cache, a gap longer than the TTL means the next request pays for the write again. That’s fine for chat, but wasteful for a batch job that pauses between items. Either keep requests flowing or reach for the one-hour TTL.</p>\n<h2>Confirm it’s working</h2>\n<p>Don’t assume. The <code class=\"language-text\">usage</code> object on every response splits the input tokens into three buckets and tells you exactly what happened:</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\">u <span class=\"token operator\">=</span> resp<span class=\"token punctuation\">.</span>usage\n<span class=\"token keyword\">print</span><span class=\"token punctuation\">(</span>u<span class=\"token punctuation\">.</span>input_tokens<span class=\"token punctuation\">)</span>                  <span class=\"token comment\"># uncached tokens, full price</span>\n<span class=\"token keyword\">print</span><span class=\"token punctuation\">(</span>u<span class=\"token punctuation\">.</span>cache_creation_input_tokens<span class=\"token punctuation\">)</span>   <span class=\"token comment\"># written to cache this request (1.25x)</span>\n<span class=\"token keyword\">print</span><span class=\"token punctuation\">(</span>u<span class=\"token punctuation\">.</span>cache_read_input_tokens<span class=\"token punctuation\">)</span>       <span class=\"token comment\"># served from cache (0.1x)</span></code></pre></div>\n<p>The signal you want is <code class=\"language-text\">cache_read_input_tokens</code> climbing on repeated requests. If it stays at zero across calls you believe share a prefix, one of the invalidators above is at work. Start by diffing the exact bytes of two consecutive prefixes.</p>\n<p>On the first request with a fresh prefix, you’ll see <code class=\"language-text\">cache_creation_input_tokens</code> populated and reads at zero. That’s expected, since something has to write the cache before anything can read it.</p>\n<h2>Tradeoffs, and what I watch</h2>\n<p>Caching isn’t free to reason about. Three things are worth keeping in mind:</p>\n<ul>\n<li><strong>Single-use prefixes are a small loss.</strong> Because of the write premium, caching a prefix you’ll use exactly once costs slightly more than not caching. Be honest about your reuse pattern before marking everything <code class=\"language-text\">ephemeral</code>.</li>\n<li><strong>Low-traffic endpoints may never hit.</strong> If requests arrive minutes apart, they may never see a hit on the default TTL.</li>\n<li><strong>The cache is scoped to a model.</strong> A mid-conversation switch to a different model, or an <a href=\"https://platform.claude.com/docs/en/build-with-claude/extended-thinking\">effort</a> change, starts a new cache namespace. A routing layer that shuffles models per request can forfeit the reuse you were counting on.</li>\n</ul>\n<p>The mental model that keeps me out of trouble is to treat the prefix as immutable and design the request around it: freeze the system prompt, fix the tool order, sort any serialized structure, and shove every varying byte past the last breakpoint. Once the prefix is genuinely stable, caching is close to free money on any workload that repeats it. For the LLM tooling I’ve built around <a href=\"/project/cloud-canvas-ai/\">CloudCanvasAI</a> and <a href=\"/project/archi/\">Archi</a>, that’s nearly all of them.</p>\n<p>If you’re wiring this into a service, it pairs naturally with the <a href=\"/blog/2026-07-11-llm-api-rate-limits-retries-backoff/\">retry and backoff patterns</a> I covered earlier. A cached prefix makes a retried request cheaper too, since the retry reads the same cache the first attempt wrote.</p>","frontmatter":{"title":"Prompt Caching with the Claude API","date":"2026-08-30T00:00:00.000Z","description":"Prompt caching reuses a stable prompt prefix to cut Claude API cost and latency. How breakpoints, TTLs, and cache reads work, and what quietly breaks a hit.","thumbnail":{"childImageSharp":{"fluid":{"base64":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAABQAAAALCAIAAADwazoUAAAACXBIWXMAAAsSAAALEgHS3X78AAABw0lEQVQoz1WP6XKjMBCE+RMbEBK3bonLgMGOTRInqayvXb//Q+1gJ1u11FeqVk83U3KGTo+dHlq1Hezt8vb1sbn82v8+Teev3eW4Px9Bv4ADVxCnuwnJoZ1bzoIUS1I+BRbnfb+71sNxPf1Zbc/1eOq219X20j1fWxCb8zDdms0JBKHrJ2yh5QSJAfzIkLzx49INjRcaOFEMpgYfxF1/mwBOiyCeWw6axwVTtSo6U7Y513FmmNAusTB+pFFkSWrSXIW0r4Z3075ldgqS0vFDyBkuzeeO3j7z6zvdDfx6YEmuXKJRrO/7Z/ZrLnUR8U0qh5j1sBXKyiMqzvSm59PIxo4dnikwDZQk0Ff3gITTGn7YZiRiCyQ8wsFxfCJdLDOqqoJbywvLq4KVlvc1fRlzFAoIRan1Q2nUbL6OOY7l/DsiHY8IF4so1VoJJX9QQgjZNbI2kmVVo/ssMYJL+NpKFlbAi/xQOB7my0AUmsfcBnaFZY15FajaC4QfzQtRKBeYS1qqvPRCeIWBfbAWilBmKORK24SZUFSJbjKzwpl2EYUR4GLmE6FEA+BIusG3DzhwARaILhF1UQ7cBX34/1g+Av+bfwHdJ0MNNm91MgAAAABJRU5ErkJggg==","aspectRatio":1.899441340782123,"src":"/static/ee39d6bd9f6285aa6710ac73b9b54026/40a76/hero.png","srcSet":"/static/ee39d6bd9f6285aa6710ac73b9b54026/c972b/hero.png 340w,\n/static/ee39d6bd9f6285aa6710ac73b9b54026/27625/hero.png 680w,\n/static/ee39d6bd9f6285aa6710ac73b9b54026/40a76/hero.png 1360w,\n/static/ee39d6bd9f6285aa6710ac73b9b54026/ed396/hero.png 2000w","sizes":"(max-width: 1360px) 100vw, 1360px"}}}}}},"pageContext":{"slug":"/2026-08-30-prompt-caching-claude-api/","previous":"blog/2026-08-26-scheduling-gpu-pods-kubernetes/","next":"blog/2026-08-29-circuit-breaker-llm-api-calls/"}},"staticQueryHashes":["32046230"]}