{"componentChunkName":"component---src-templates-blog-post-js","path":"/blog/2026-07-31-graceful-shutdown-fastapi-kubernetes/","result":{"data":{"site":{"siteMetadata":{"title":"M.Hassan Ahmed","author":"Hassan11196"}},"markdownRemark":{"id":"5e20654d-788f-5d13-ac34-605690744acc","excerpt":"You push a new image, ArgoCD rolls the deployment, and for a few seconds a handful of requests come back as  or a reset connection. The app that died logged…","html":"<p>You push a new image, ArgoCD rolls the deployment, and for a few seconds a handful of requests come back as <code class=\"language-text\">502</code> or a reset connection. The app that died logged nothing. Users on a live token stream see it cut off mid-sentence. Roll again and the same thing happens, just to different requests. The code is fine. What is broken is the handoff between Kubernetes tearing a pod down and your FastAPI server noticing.</p>\n<p>This post is for engineers running FastAPI (or any ASGI app) on Kubernetes who lose in-flight requests on every rolling deploy and want to know exactly why. The short version: a pod termination is not one event but several, and they fire in an order most people guess wrong. Get the order right and the drop rate goes to zero. I’ll walk through that sequence, what uvicorn does on shutdown, the <code class=\"language-text\">preStop</code> hook that closes the race, where cleanup belongs, and how to handle long-lived streams.</p>\n<p>I hit this running the <a href=\"/project/cms-workflow-operations/\">FastAPI services behind the CMS workflow operations stack</a> at CERN. There, a rolling update should be invisible to an operator watching a dashboard, not a reason for the request they just submitted to vanish.</p>\n<h2>A pod termination is a sequence, not a moment</h2>\n<p>When you (or a rolling update) delete a pod, the kubelet (the agent on each node that runs pods) <a href=\"https://kubernetes.io/docs/concepts/workloads/pods/pod-lifecycle/#pod-termination\">runs a specific sequence</a>, and two of the steps happen at the same time:</p>\n<ol>\n<li>The pod is marked <code class=\"language-text\">Terminating</code>, and the <a href=\"https://kubernetes.io/docs/concepts/services-networking/endpoint-slices/\">EndpointSlice controller</a> starts removing it from the Service endpoints (the list of pod addresses the Service routes traffic to).</li>\n<li><strong>Concurrently</strong>, the kubelet runs your <code class=\"language-text\">preStop</code> hook if you have one, then sends <code class=\"language-text\">SIGTERM</code> to the container’s main process.</li>\n<li>The <code class=\"language-text\">terminationGracePeriodSeconds</code> clock (default 30) counts down. That clock started at step 1, so it includes the time spent in <code class=\"language-text\">preStop</code>.</li>\n<li>If the process has not exited when the clock hits zero, the kubelet sends <code class=\"language-text\">SIGKILL</code>. That is the <code class=\"language-text\">exit 137</code> you never want to see for a web server.</li>\n</ol>\n<p>The trap is the word <em>concurrently</em> in step 2. Endpoint removal is not instant, and it is not synchronized with <code class=\"language-text\">SIGTERM</code>. The removal has to propagate to every node’s <a href=\"https://kubernetes.io/docs/reference/networking/virtual-ips/\">kube-proxy</a> (or your ingress, or your service mesh) before those data planes, the components that actually forward traffic, stop routing to the dying pod. Meanwhile, <code class=\"language-text\">SIGTERM</code> has already told your app to shut down. So there is a window where the load balancer is still sending new requests to a server that has already decided to stop accepting them.</p>\n<p><img src=\"/0677269e8ec0c012e4568344b830b9db/termination-timeline.svg\" alt=\"Two timelines comparing pod termination without and with a preStop hook. Without one, SIGTERM and endpoint removal fire together, so kube-proxy keeps routing new requests to a server that already stopped accepting, producing resets. With a preStop sleep of fifteen seconds, endpoint removal finishes during the sleep while the app keeps serving; only then does SIGTERM arrive and uvicorn drains in-flight work and exits before the grace period deadline.\"></p>\n<p>That window is the whole problem. Everything below is about closing it.</p>\n<h2>What uvicorn does when it gets SIGTERM</h2>\n<p>The good news is that <a href=\"https://www.uvicorn.org/deployment/\">uvicorn already handles <code class=\"language-text\">SIGTERM</code> gracefully</a>. On the signal, it stops accepting new connections, lets the requests already in flight finish, runs the ASGI lifespan shutdown, and exits. You do not have to write signal handling yourself for the request-draining part.</p>\n<p>Two conditions have to hold for that to actually happen, and both are easy to break.</p>\n<h3>Uvicorn has to be the process that receives the signal</h3>\n<p>If your container starts the server through a shell, the shell is PID 1 (the container’s main process), and a shell does not forward <code class=\"language-text\">SIGTERM</code> to its child by default. For example, <code class=\"language-text\">CMD python -m uvicorn ...</code> is fine, but <code class=\"language-text\">CMD sh -c &quot;uvicorn app:app&quot;</code> is not. The signal lands on the shell and uvicorn never hears it. Nothing drains, and the pod gets <code class=\"language-text\">SIGKILL</code>ed at the end of the grace period every single time. Use the exec form (the JSON-array syntax) so uvicorn is PID 1:</p>\n<div class=\"gatsby-highlight\" data-language=\"dockerfile\"><pre class=\"language-dockerfile\"><code class=\"language-dockerfile\"><span class=\"token comment\"># PID 1 is uvicorn, so it receives SIGTERM directly</span>\n<span class=\"token keyword\">CMD</span> <span class=\"token punctuation\">[</span><span class=\"token string\">\"uvicorn\"</span><span class=\"token punctuation\">,</span> <span class=\"token string\">\"app:app\"</span><span class=\"token punctuation\">,</span> <span class=\"token string\">\"--host\"</span><span class=\"token punctuation\">,</span> <span class=\"token string\">\"0.0.0.0\"</span><span class=\"token punctuation\">,</span> <span class=\"token string\">\"--port\"</span><span class=\"token punctuation\">,</span> <span class=\"token string\">\"8000\"</span><span class=\"token punctuation\">]</span></code></pre></div>\n<p>If you genuinely need a wrapper, run the server with <code class=\"language-text\">exec</code> (<code class=\"language-text\">exec uvicorn ...</code>) so it replaces the shell rather than running underneath it, or add an init like <code class=\"language-text\">tini</code> to forward signals.</p>\n<h3>The drain has to fit inside the grace period</h3>\n<p>Uvicorn will wait for in-flight requests, but the kubelet will not wait for uvicorn past <code class=\"language-text\">terminationGracePeriodSeconds</code>. If a request takes 40 seconds and your grace period is 30, that request gets <code class=\"language-text\">SIGKILL</code>ed no matter how polite uvicorn is being. Cap uvicorn’s own wait so it does not sit forever on a stuck connection:</p>\n<div class=\"gatsby-highlight\" data-language=\"bash\"><pre class=\"language-bash\"><code class=\"language-bash\">uvicorn app:app --host <span class=\"token number\">0.0</span>.0.0 --port <span class=\"token number\">8000</span> <span class=\"token punctuation\">\\</span>\n  --timeout-graceful-shutdown <span class=\"token number\">25</span></code></pre></div>\n<p>Set that value a little below your grace period. Uvicorn then force-closes stragglers and exits cleanly before the kubelet reaches for the bigger hammer.</p>\n<h2>Closing the endpoint-removal race with preStop</h2>\n<p>Uvicorn’s drain handles the requests already in flight. It does nothing about the <em>new</em> requests still arriving during the propagation window, because from the app’s point of view they are legitimate new connections. The fix is not in the app at all. You delay <code class=\"language-text\">SIGTERM</code> long enough for endpoint removal to finish first, using a <a href=\"https://kubernetes.io/docs/concepts/containers/container-lifecycle-hooks/\"><code class=\"language-text\">preStop</code> hook</a> that just sleeps:</p>\n<div class=\"gatsby-highlight\" data-language=\"yaml\"><pre class=\"language-yaml\"><code class=\"language-yaml\"><span class=\"token key atrule\">containers</span><span class=\"token punctuation\">:</span>\n  <span class=\"token punctuation\">-</span> <span class=\"token key atrule\">name</span><span class=\"token punctuation\">:</span> api\n    <span class=\"token key atrule\">lifecycle</span><span class=\"token punctuation\">:</span>\n      <span class=\"token key atrule\">preStop</span><span class=\"token punctuation\">:</span>\n        <span class=\"token key atrule\">exec</span><span class=\"token punctuation\">:</span>\n          <span class=\"token key atrule\">command</span><span class=\"token punctuation\">:</span> <span class=\"token punctuation\">[</span><span class=\"token string\">\"sleep\"</span><span class=\"token punctuation\">,</span> <span class=\"token string\">\"15\"</span><span class=\"token punctuation\">]</span>\n    <span class=\"token key atrule\">terminationGracePeriodSeconds</span><span class=\"token punctuation\">:</span> <span class=\"token number\">45</span>   <span class=\"token comment\"># must exceed preStop + drain</span></code></pre></div>\n<p>The kubelet runs <code class=\"language-text\">preStop</code> and sends <code class=\"language-text\">SIGTERM</code> only after it returns. During those 15 seconds the app keeps serving normally (it has not been told to stop), while the endpoint removal propagates out to every kube-proxy and ingress. By the time <code class=\"language-text\">SIGTERM</code> arrives, nothing is routing new traffic to this pod; new traffic already goes to the other replicas. Uvicorn’s drain only has to finish the requests that were genuinely in flight.</p>\n<p>The number is not magic. It means “longer than your data plane takes to converge.” Five seconds is usually enough for kube-proxy in a small cluster, but a large cluster, an external load balancer, or a mesh like Istio can need more. Measure it: watch the reset rate during a rollout and raise the sleep until the rate hits zero.</p>\n<p>And mind the arithmetic in step 3 above: the sleep is <em>inside</em> the grace period. A 15-second <code class=\"language-text\">preStop</code> under a 30-second grace period leaves uvicorn only 15 seconds to drain. I set the grace period to <code class=\"language-text\">preStop sleep + worst-case drain + a little slack</code>, which is why the example uses 45.</p>\n<h2>The <code class=\"language-text\">sleep</code> binary has to exist</h2>\n<p>One footgun deserves its own section: <code class=\"language-text\">preStop: exec: command: [&quot;sleep&quot;, &quot;15&quot;]</code> needs a <code class=\"language-text\">sleep</code> binary in the image. Distroless and <code class=\"language-text\">scratch</code>-based images often do not have one. The hook then fails silently: the kubelet logs a <code class=\"language-text\">FailedPreStopHook</code> event and proceeds straight to <code class=\"language-text\">SIGTERM</code>, which puts you right back in the race you were trying to close.</p>\n<p>If your base image is minimal, either add a static <code class=\"language-text\">sleep</code>, or move the delay into the app by handling <code class=\"language-text\">SIGTERM</code> yourself and sleeping before you let uvicorn shut down. Check for the <code class=\"language-text\">FailedPreStopHook</code> event after your first deploy; do not assume the hook ran.</p>\n<h2>Readiness probes are not the drain mechanism</h2>\n<p>A common piece of advice is “flip your readiness probe to failing on shutdown so Kubernetes stops sending traffic.” It sounds right, and it is mostly a distraction.</p>\n<p>Failing readiness <em>is</em> one of the signals that triggers endpoint removal. But during termination, the pod is already being removed from endpoints for a more direct reason: it is <code class=\"language-text\">Terminating</code>. Racing your own readiness probe against that does not reliably beat the propagation delay, and it adds a moving part. The <code class=\"language-text\">preStop</code> sleep addresses the real problem, data-plane propagation lag, head on, which is why it is the pattern the <a href=\"https://kubernetes.io/docs/tutorials/services/pods-and-endpoints/\">Kubernetes docs themselves reach for</a>.</p>\n<p>Keep readiness for its actual job, which I wrote about in <a href=\"/blog/2026-07-07-kubernetes-liveness-readiness-startup-probes/\">Kubernetes liveness, readiness, and startup probes</a>: telling Kubernetes when a <em>starting</em> pod is ready, not choreographing a shutdown.</p>\n<h2>Cleanup belongs in the lifespan shutdown</h2>\n<p>Once the drain works, use the graceful path for the cleanup that a <code class=\"language-text\">SIGKILL</code> would have skipped: closing database pools, flushing a metrics buffer, releasing a lease. In FastAPI that goes in the <a href=\"https://fastapi.tiangolo.com/advanced/events/\">lifespan context</a>. The code after <code class=\"language-text\">yield</code> runs during uvicorn’s graceful shutdown:</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\"><span class=\"token keyword\">from</span> contextlib <span class=\"token keyword\">import</span> asynccontextmanager\n<span class=\"token keyword\">from</span> fastapi <span class=\"token keyword\">import</span> FastAPI\n\n<span class=\"token decorator annotation punctuation\">@asynccontextmanager</span>\n<span class=\"token keyword\">async</span> <span class=\"token keyword\">def</span> <span class=\"token function\">lifespan</span><span class=\"token punctuation\">(</span>app<span class=\"token punctuation\">:</span> FastAPI<span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n    pool <span class=\"token operator\">=</span> <span class=\"token keyword\">await</span> create_pool<span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span>\n    app<span class=\"token punctuation\">.</span>state<span class=\"token punctuation\">.</span>pool <span class=\"token operator\">=</span> pool\n    <span class=\"token keyword\">yield</span>\n    <span class=\"token comment\"># runs on graceful shutdown, after in-flight requests drain</span>\n    <span class=\"token keyword\">await</span> pool<span class=\"token punctuation\">.</span>close<span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span>\n\napp <span class=\"token operator\">=</span> FastAPI<span class=\"token punctuation\">(</span>lifespan<span class=\"token operator\">=</span>lifespan<span class=\"token punctuation\">)</span></code></pre></div>\n<p>Be honest about the guarantee here: this runs on a graceful shutdown, not on <code class=\"language-text\">SIGKILL</code>. If the pod blows past its grace period, or the node dies, this code does not execute. So lifespan cleanup is for being tidy, not for correctness.</p>\n<p>Anything that must not be lost (a half-finished write, an unacknowledged job) needs to be safe against a hard kill anyway. Use a transaction, an idempotent retry, or a queue ack you send only after the work is durable. Treat graceful shutdown as an optimization on top of a design that already survives a kill, not as the thing that makes the kill safe.</p>\n<h2>Long-lived streams: SSE, WebSockets, and agent runs</h2>\n<p>Here is where the tidy “drain in-flight requests” model stops being enough, and it is the case I care about most. A Server-Sent Events (SSE) stream or a WebSocket does not <em>finish</em> on its own; it stays open as long as the client is listening. If you are streaming LLM tokens to a browser, a “request” can last minutes. Uvicorn’s graceful drain will politely wait for it, hit <code class=\"language-text\">--timeout-graceful-shutdown</code>, and cut it anyway. So a pure drain either stalls your whole rollout on the longest-lived connection or chops it mid-stream. Neither is what you want.</p>\n<p>The fix is to design the stream to survive being interrupted, not to hold the deploy hostage to it. On <code class=\"language-text\">SIGTERM</code>:</p>\n<ul>\n<li>Stop starting <em>new</em> streams. The <code class=\"language-text\">preStop</code> window plus endpoint removal already handles that.</li>\n<li>For the streams still open, send an explicit end-of-stream event and close, rather than letting them get guillotined at the timeout.</li>\n</ul>\n<p>In this generator, a set <code class=\"language-text\">shutting_down</code> event (or a client disconnect) makes the stream send an <code class=\"language-text\">interrupted</code> event and return:</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\"><span class=\"token keyword\">import</span> asyncio<span class=\"token punctuation\">,</span> contextlib\n\nshutting_down <span class=\"token operator\">=</span> asyncio<span class=\"token punctuation\">.</span>Event<span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span>\n\n<span class=\"token keyword\">async</span> <span class=\"token keyword\">def</span> <span class=\"token function\">token_stream</span><span class=\"token punctuation\">(</span>request<span class=\"token punctuation\">,</span> prompt<span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n    <span class=\"token keyword\">async</span> <span class=\"token keyword\">for</span> token <span class=\"token keyword\">in</span> generate<span class=\"token punctuation\">(</span>prompt<span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n        <span class=\"token keyword\">if</span> shutting_down<span class=\"token punctuation\">.</span>is_set<span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span> <span class=\"token keyword\">or</span> <span class=\"token keyword\">await</span> request<span class=\"token punctuation\">.</span>is_disconnected<span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n            <span class=\"token comment\"># tell the client this stream ended on our terms</span>\n            <span class=\"token keyword\">yield</span> <span class=\"token punctuation\">{</span><span class=\"token string\">\"event\"</span><span class=\"token punctuation\">:</span> <span class=\"token string\">\"interrupted\"</span><span class=\"token punctuation\">,</span> <span class=\"token string\">\"data\"</span><span class=\"token punctuation\">:</span> <span class=\"token string\">\"reconnect\"</span><span class=\"token punctuation\">}</span>\n            <span class=\"token keyword\">return</span>\n        <span class=\"token keyword\">yield</span> <span class=\"token punctuation\">{</span><span class=\"token string\">\"event\"</span><span class=\"token punctuation\">:</span> <span class=\"token string\">\"token\"</span><span class=\"token punctuation\">,</span> <span class=\"token string\">\"data\"</span><span class=\"token punctuation\">:</span> token<span class=\"token punctuation\">}</span></code></pre></div>\n<p>The client watches for that <code class=\"language-text\">interrupted</code> event and reopens the stream. The new stream lands on a healthy replica, because the dying pod is already out of the endpoints. I leaned on the same reconnect-and-resume design for the streaming backends in <a href=\"/project/cloud-canvas-ai/\">CloudCanvasAI</a> and <a href=\"/project/archi/\">Archi</a>. Once a stream can recover from <em>any</em> disconnect, a deploy is just one more disconnect it already knows how to handle.</p>\n<p>If SSE mechanics are new to you, I covered the server side in <a href=\"/blog/2026-06-30-fastapi-sse-streaming-llm/\">FastAPI SSE streaming for LLMs</a> and the browser side in <a href=\"/blog/2026-07-16-cancel-llm-stream-react-abortcontroller/\">cancelling an LLM stream in React</a>.</p>\n<h2>Failure modes that waste an afternoon</h2>\n<p><strong>The shell-wrapper signal trap.</strong> This is the single most common cause, so it is worth repeating. If uvicorn is not PID 1 and nothing forwards the signal, <em>every</em> pod gets <code class=\"language-text\">SIGKILL</code>ed at the grace-period deadline and no drain ever runs. The symptom is that every rollout drops the same number of connections and shutdown always takes exactly <code class=\"language-text\">terminationGracePeriodSeconds</code>; that fixed duration is the tell.</p>\n<p><strong>A blocking call that starves the shutdown.</strong> If a request handler is doing synchronous work on the event loop, uvicorn cannot process the shutdown promptly because the loop is busy. The drain stalls, the grace period expires, and <code class=\"language-text\">SIGKILL</code> follows. This is the same problem I dug into in <a href=\"/blog/2026-07-10-fastapi-blocking-event-loop/\">FastAPI event loop blocking</a>, and it shows up during shutdown too, not just under load.</p>\n<p><strong>Grace period shorter than the work.</strong> If your slowest legitimate request takes 20 seconds and your grace period is 30 with a 15-second <code class=\"language-text\">preStop</code>, the request has 15 seconds and gets killed. The <code class=\"language-text\">preStop</code> sleep, the uvicorn timeout, and the grace period are three numbers that have to be set together, not independently.</p>\n<p><strong><code class=\"language-text\">SIGKILL</code> mistaken for OOM.</strong> A graceful-shutdown timeout and an out-of-memory kill both surface as <code class=\"language-text\">exit 137</code>, because both are <code class=\"language-text\">128 + 9</code>. Read the pod’s <code class=\"language-text\">lastState.terminated.reason</code>: <code class=\"language-text\">OOMKilled</code> is memory, and a plain timeout kill is not. I pulled that thread apart in <a href=\"/blog/2026-07-15-kubernetes-oomkilled-requests-vs-limits/\">Kubernetes OOMKilled: requests vs limits</a>; do not spend an afternoon tuning memory limits for what is actually a drain that ran out of time.</p>\n<h2>What I would do differently</h2>\n<p>Early on, I treated graceful shutdown as an app problem and went looking for the right signal handler to write. Most of the fix actually lived in the pod spec: a <code class=\"language-text\">preStop</code> sleep and a grace period that added up. The app-side work that mattered was smaller than I expected: make uvicorn PID 1, cap its shutdown timeout, and put cleanup in lifespan.</p>\n<p>The larger lesson was to stop treating streams as requests that happen to be long. A request wants to <em>finish</em>; a stream wants to <em>resume</em>. Once I designed the streaming endpoints to reconnect cleanly from any interruption, deploys stopped being a special case, and so did flaky networks and closed laptops. The deploy was never really the hard part. A stream that could not recover was.</p>\n<h2>Where this runs</h2>\n<p>This is the shutdown discipline behind the <a href=\"/project/cms-workflow-operations/\">FastAPI operator console for CMS workflow operations</a>, deployed on Kubernetes through ArgoCD. There, a rolling update lands during the working day, and an operator should never notice one happened. Getting the termination sequence right makes that possible: endpoint removal before <code class=\"language-text\">SIGTERM</code>, a drain that fits the grace period, and streams that resume instead of snapping. It is the difference between a deploy you can run at 2pm without thinking and one you schedule for a quiet window and watch nervously.</p>\n<p>It pairs with the rollout ordering I wrote about in <a href=\"/blog/2026-07-03-argocd-sync-waves-ordering-rollout/\">ArgoCD sync waves</a>. Sync waves decide what comes up in what order; graceful shutdown makes sure that what goes down does so without dropping anyone mid-request.</p>\n<p><em>Diagram by M. Hassan Ahmed, made for this post.</em></p>\n<p>Image credit: Diagram created by M. Hassan Ahmed for this post, released under CC0 (public domain).</p>","frontmatter":{"title":"Graceful Shutdown for FastAPI on Kubernetes","date":"2026-07-31T00:00:00.000Z","description":"A rolling deploy sends SIGTERM and kills your FastAPI pod mid-request, dropping live SSE streams. How to catch it, drain connections, and shut down cleanly.","thumbnail":{"childImageSharp":{"fluid":{"base64":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAABQAAAALCAIAAADwazoUAAAACXBIWXMAAAsSAAALEgHS3X78AAAB/UlEQVQoz02P3W7aMBzFLbWFOP4KBIjtOHEgxIS0fGltByEB+jG6Vdplpe1iL7LH2NXu9yy73ivNYdVU6ae/jo/P+VsGU6OKE9MsrMv5oV5Vm9nD4Xq3XT4erqvNvLy9rDezfbWsN/NdudjcFHkmm7xRoE21pUVi7GfZ1W66uh8VlZ1pUeWrezPbW9POyeJufFlPFgedl8gfn1oJwN3IYVbFrpdYC7IIMgWZxJ0IktC1MOV6Eeo0WI1ohL0Ye80ENqSjIIkDrQbDOOgNBI8K1p9euIqGIzY0RI5c3GxBTBGesCSleoSVJmECSEcet+rTPn3cJtuVb4Ye8ounOitf1t2PFf+w5s9V8FRSnTlEDmb51XE9O5ZymVETAuxJk4a50ROjTSo8P7i/6f75eX67Hp8hhQiHVHaKIny5o8OMBlpMpyIvelmGoxi4VLaxhTtYXCDZ8eXvH+3Dux4AAlHRtn+moXPRZ6kRn2vXOog3YSLsRwCkwlp2PfYEOA+/Hr1f36EVDglGqdnv3g9E0rbvw2Bw3FAzcWDgMgmbigDQXpxAjIMz/u2Zf3ni4JxTP4oSs1zOw9iQrm45fb2cxcVl2+lD+lqx5eA/LuWkP+7wnPqxLSMvamGJPEW6EfPjUGRaTdym+ZoHDh68hXQV68Wncsx6mp22/DuSk4Bvwn8B/DZLUAfiRtMAAAAASUVORK5CYII=","aspectRatio":1.899441340782123,"src":"/static/0ed827f19ec7b4a87eeef897fe2b7026/40a76/hero.png","srcSet":"/static/0ed827f19ec7b4a87eeef897fe2b7026/c972b/hero.png 340w,\n/static/0ed827f19ec7b4a87eeef897fe2b7026/27625/hero.png 680w,\n/static/0ed827f19ec7b4a87eeef897fe2b7026/40a76/hero.png 1360w,\n/static/0ed827f19ec7b4a87eeef897fe2b7026/ed396/hero.png 2000w","sizes":"(max-width: 1360px) 100vw, 1360px"}}}}}},"pageContext":{"slug":"/2026-07-31-graceful-shutdown-fastapi-kubernetes/","previous":"blog/2026-07-27-prompt-caching-cut-llm-cost/","next":"blog/2026-07-30-managing-context-window-llm-agents/"}},"staticQueryHashes":["32046230"]}