Graceful Shutdown for FastAPI on Kubernetes
A rolling deploy sends SIGTERM and kills your FastAPI pod mid-request, dropping live SSE streams. How to catch it, drain connections, and shut down cleanly.
You push a new image, ArgoCD rolls the deployment, and for a few seconds a handful of requests come back as 502 or a reset connection. The app that died logged nothing. Users on a live token stream see it cut off mid-sentence. Roll again and the same thing happens, just to different requests. The code is fine. What is broken is the handoff between Kubernetes tearing a pod down and your FastAPI server noticing.
This post is for engineers running FastAPI (or any ASGI app) on Kubernetes who lose in-flight requests on every rolling deploy and want to know exactly why. The short version: a pod termination is not one event but several, and they fire in an order most people guess wrong. Get the order right and the drop rate goes to zero. I’ll walk through that sequence, what uvicorn does on shutdown, the preStop hook that closes the race, where cleanup belongs, and how to handle long-lived streams.
I hit this running the FastAPI services behind the CMS workflow operations stack at CERN. There, a rolling update should be invisible to an operator watching a dashboard, not a reason for the request they just submitted to vanish.
A pod termination is a sequence, not a moment
When you (or a rolling update) delete a pod, the kubelet (the agent on each node that runs pods) runs a specific sequence, and two of the steps happen at the same time:
- The pod is marked
Terminating, and the EndpointSlice controller starts removing it from the Service endpoints (the list of pod addresses the Service routes traffic to). - Concurrently, the kubelet runs your
preStophook if you have one, then sendsSIGTERMto the container’s main process. - The
terminationGracePeriodSecondsclock (default 30) counts down. That clock started at step 1, so it includes the time spent inpreStop. - If the process has not exited when the clock hits zero, the kubelet sends
SIGKILL. That is theexit 137you never want to see for a web server.
The trap is the word concurrently in step 2. Endpoint removal is not instant, and it is not synchronized with SIGTERM. The removal has to propagate to every node’s kube-proxy (or your ingress, or your service mesh) before those data planes, the components that actually forward traffic, stop routing to the dying pod. Meanwhile, SIGTERM has already told your app to shut down. So there is a window where the load balancer is still sending new requests to a server that has already decided to stop accepting them.
That window is the whole problem. Everything below is about closing it.
What uvicorn does when it gets SIGTERM
The good news is that uvicorn already handles SIGTERM gracefully. On the signal, it stops accepting new connections, lets the requests already in flight finish, runs the ASGI lifespan shutdown, and exits. You do not have to write signal handling yourself for the request-draining part.
Two conditions have to hold for that to actually happen, and both are easy to break.
Uvicorn has to be the process that receives the signal
If your container starts the server through a shell, the shell is PID 1 (the container’s main process), and a shell does not forward SIGTERM to its child by default. For example, CMD python -m uvicorn ... is fine, but CMD sh -c "uvicorn app:app" is not. The signal lands on the shell and uvicorn never hears it. Nothing drains, and the pod gets SIGKILLed at the end of the grace period every single time. Use the exec form (the JSON-array syntax) so uvicorn is PID 1:
# PID 1 is uvicorn, so it receives SIGTERM directly
CMD ["uvicorn", "app:app", "--host", "0.0.0.0", "--port", "8000"]If you genuinely need a wrapper, run the server with exec (exec uvicorn ...) so it replaces the shell rather than running underneath it, or add an init like tini to forward signals.
The drain has to fit inside the grace period
Uvicorn will wait for in-flight requests, but the kubelet will not wait for uvicorn past terminationGracePeriodSeconds. If a request takes 40 seconds and your grace period is 30, that request gets SIGKILLed no matter how polite uvicorn is being. Cap uvicorn’s own wait so it does not sit forever on a stuck connection:
uvicorn app:app --host 0.0.0.0 --port 8000 \
--timeout-graceful-shutdown 25Set that value a little below your grace period. Uvicorn then force-closes stragglers and exits cleanly before the kubelet reaches for the bigger hammer.
Closing the endpoint-removal race with preStop
Uvicorn’s drain handles the requests already in flight. It does nothing about the new requests still arriving during the propagation window, because from the app’s point of view they are legitimate new connections. The fix is not in the app at all. You delay SIGTERM long enough for endpoint removal to finish first, using a preStop hook that just sleeps:
containers:
- name: api
lifecycle:
preStop:
exec:
command: ["sleep", "15"]
terminationGracePeriodSeconds: 45 # must exceed preStop + drainThe kubelet runs preStop and sends SIGTERM only after it returns. During those 15 seconds the app keeps serving normally (it has not been told to stop), while the endpoint removal propagates out to every kube-proxy and ingress. By the time SIGTERM arrives, nothing is routing new traffic to this pod; new traffic already goes to the other replicas. Uvicorn’s drain only has to finish the requests that were genuinely in flight.
The number is not magic. It means “longer than your data plane takes to converge.” Five seconds is usually enough for kube-proxy in a small cluster, but a large cluster, an external load balancer, or a mesh like Istio can need more. Measure it: watch the reset rate during a rollout and raise the sleep until the rate hits zero.
And mind the arithmetic in step 3 above: the sleep is inside the grace period. A 15-second preStop under a 30-second grace period leaves uvicorn only 15 seconds to drain. I set the grace period to preStop sleep + worst-case drain + a little slack, which is why the example uses 45.
The sleep binary has to exist
One footgun deserves its own section: preStop: exec: command: ["sleep", "15"] needs a sleep binary in the image. Distroless and scratch-based images often do not have one. The hook then fails silently: the kubelet logs a FailedPreStopHook event and proceeds straight to SIGTERM, which puts you right back in the race you were trying to close.
If your base image is minimal, either add a static sleep, or move the delay into the app by handling SIGTERM yourself and sleeping before you let uvicorn shut down. Check for the FailedPreStopHook event after your first deploy; do not assume the hook ran.
Readiness probes are not the drain mechanism
A common piece of advice is “flip your readiness probe to failing on shutdown so Kubernetes stops sending traffic.” It sounds right, and it is mostly a distraction.
Failing readiness is one of the signals that triggers endpoint removal. But during termination, the pod is already being removed from endpoints for a more direct reason: it is Terminating. Racing your own readiness probe against that does not reliably beat the propagation delay, and it adds a moving part. The preStop sleep addresses the real problem, data-plane propagation lag, head on, which is why it is the pattern the Kubernetes docs themselves reach for.
Keep readiness for its actual job, which I wrote about in Kubernetes liveness, readiness, and startup probes: telling Kubernetes when a starting pod is ready, not choreographing a shutdown.
Cleanup belongs in the lifespan shutdown
Once the drain works, use the graceful path for the cleanup that a SIGKILL would have skipped: closing database pools, flushing a metrics buffer, releasing a lease. In FastAPI that goes in the lifespan context. The code after yield runs during uvicorn’s graceful shutdown:
from contextlib import asynccontextmanager
from fastapi import FastAPI
@asynccontextmanager
async def lifespan(app: FastAPI):
pool = await create_pool()
app.state.pool = pool
yield
# runs on graceful shutdown, after in-flight requests drain
await pool.close()
app = FastAPI(lifespan=lifespan)Be honest about the guarantee here: this runs on a graceful shutdown, not on SIGKILL. If the pod blows past its grace period, or the node dies, this code does not execute. So lifespan cleanup is for being tidy, not for correctness.
Anything that must not be lost (a half-finished write, an unacknowledged job) needs to be safe against a hard kill anyway. Use a transaction, an idempotent retry, or a queue ack you send only after the work is durable. Treat graceful shutdown as an optimization on top of a design that already survives a kill, not as the thing that makes the kill safe.
Long-lived streams: SSE, WebSockets, and agent runs
Here is where the tidy “drain in-flight requests” model stops being enough, and it is the case I care about most. A Server-Sent Events (SSE) stream or a WebSocket does not finish on its own; it stays open as long as the client is listening. If you are streaming LLM tokens to a browser, a “request” can last minutes. Uvicorn’s graceful drain will politely wait for it, hit --timeout-graceful-shutdown, and cut it anyway. So a pure drain either stalls your whole rollout on the longest-lived connection or chops it mid-stream. Neither is what you want.
The fix is to design the stream to survive being interrupted, not to hold the deploy hostage to it. On SIGTERM:
- Stop starting new streams. The
preStopwindow plus endpoint removal already handles that. - For the streams still open, send an explicit end-of-stream event and close, rather than letting them get guillotined at the timeout.
In this generator, a set shutting_down event (or a client disconnect) makes the stream send an interrupted event and return:
import asyncio, contextlib
shutting_down = asyncio.Event()
async def token_stream(request, prompt):
async for token in generate(prompt):
if shutting_down.is_set() or await request.is_disconnected():
# tell the client this stream ended on our terms
yield {"event": "interrupted", "data": "reconnect"}
return
yield {"event": "token", "data": token}The client watches for that interrupted event and reopens the stream. The new stream lands on a healthy replica, because the dying pod is already out of the endpoints. I leaned on the same reconnect-and-resume design for the streaming backends in CloudCanvasAI and Archi. Once a stream can recover from any disconnect, a deploy is just one more disconnect it already knows how to handle.
If SSE mechanics are new to you, I covered the server side in FastAPI SSE streaming for LLMs and the browser side in cancelling an LLM stream in React.
Failure modes that waste an afternoon
The shell-wrapper signal trap. This is the single most common cause, so it is worth repeating. If uvicorn is not PID 1 and nothing forwards the signal, every pod gets SIGKILLed at the grace-period deadline and no drain ever runs. The symptom is that every rollout drops the same number of connections and shutdown always takes exactly terminationGracePeriodSeconds; that fixed duration is the tell.
A blocking call that starves the shutdown. If a request handler is doing synchronous work on the event loop, uvicorn cannot process the shutdown promptly because the loop is busy. The drain stalls, the grace period expires, and SIGKILL follows. This is the same problem I dug into in FastAPI event loop blocking, and it shows up during shutdown too, not just under load.
Grace period shorter than the work. If your slowest legitimate request takes 20 seconds and your grace period is 30 with a 15-second preStop, the request has 15 seconds and gets killed. The preStop sleep, the uvicorn timeout, and the grace period are three numbers that have to be set together, not independently.
SIGKILL mistaken for OOM. A graceful-shutdown timeout and an out-of-memory kill both surface as exit 137, because both are 128 + 9. Read the pod’s lastState.terminated.reason: OOMKilled is memory, and a plain timeout kill is not. I pulled that thread apart in Kubernetes OOMKilled: requests vs limits; do not spend an afternoon tuning memory limits for what is actually a drain that ran out of time.
What I would do differently
Early on, I treated graceful shutdown as an app problem and went looking for the right signal handler to write. Most of the fix actually lived in the pod spec: a preStop sleep and a grace period that added up. The app-side work that mattered was smaller than I expected: make uvicorn PID 1, cap its shutdown timeout, and put cleanup in lifespan.
The larger lesson was to stop treating streams as requests that happen to be long. A request wants to finish; a stream wants to resume. Once I designed the streaming endpoints to reconnect cleanly from any interruption, deploys stopped being a special case, and so did flaky networks and closed laptops. The deploy was never really the hard part. A stream that could not recover was.
Where this runs
This is the shutdown discipline behind the FastAPI operator console for CMS workflow operations, deployed on Kubernetes through ArgoCD. There, a rolling update lands during the working day, and an operator should never notice one happened. Getting the termination sequence right makes that possible: endpoint removal before SIGTERM, a drain that fits the grace period, and streams that resume instead of snapping. It is the difference between a deploy you can run at 2pm without thinking and one you schedule for a quiet window and watch nervously.
It pairs with the rollout ordering I wrote about in ArgoCD sync waves. Sync waves decide what comes up in what order; graceful shutdown makes sure that what goes down does so without dropping anyone mid-request.
Diagram by M. Hassan Ahmed, made for this post.
Image credit: Diagram created by M. Hassan Ahmed for this post, released under CC0 (public domain).