{"componentChunkName":"component---src-templates-blog-post-js","path":"/blog/2026-09-12-fastapi-websockets-streaming-llm-chat/","result":{"data":{"site":{"siteMetadata":{"title":"M.Hassan Ahmed","author":"Hassan11196"}},"markdownRemark":{"id":"b31d0de0-b10e-58dc-8998-d7a497a25e1f","excerpt":"Most LLM chat backends start with Server-Sent Events (SSE), and for good reason. You open a  response, push tokens as they come off the model, and the browser…","html":"<p>Most LLM chat backends start with Server-Sent Events (SSE), and for good reason. You open a <code class=\"language-text\">text/event-stream</code> response, push tokens as they come off the model, and the browser paints them. I covered that setup in an earlier post on <a href=\"/blog/2026-06-30-fastapi-sse-streaming-llm/\">streaming from FastAPI with SSE</a>. For a plain question-and-answer box, it is the right amount of machinery.</p>\n<p>The trouble starts when the client needs to say something back <em>while the model is still talking</em>. A user hits stop halfway through a 900-token answer, or starts typing the next message before the current one finishes. SSE is a one-way pipe by design (<a href=\"https://developer.mozilla.org/en-US/docs/Web/API/Server-sent_events/Using_server-sent_events\">MDN’s own docs</a> call this out), so the usual workaround is a second HTTP request to a <code class=\"language-text\">/cancel</code> endpoint that finds the right in-flight generation and tears it down. That works, but now you are correlating two connections and hoping they agree on which request is being cancelled.</p>\n<p>A WebSocket collapses that into one connection: tokens go down and control messages come up over the same socket. This post is for engineers who already stream LLM output from FastAPI and have hit the point where one direction isn’t enough. I’ll cover the endpoint itself and how to run the send and receive loops concurrently, so a cancel actually interrupts generation. Then come the operational parts: keeping dead connections from piling up, authenticating the connection, and what breaks once you run more than one worker.</p>\n<h2>SSE or WebSocket: pick the smaller tool that fits</h2>\n<p>Before reaching for a WebSocket, be honest about whether you need one. The <a href=\"https://datatracker.ietf.org/doc/html/rfc6455\">WebSocket protocol</a> buys you full-duplex framing (both sides can send at any time) over a single TCP connection. You need that only if messages genuinely flow <em>both</em> ways during a single exchange:</p>\n<ul>\n<li>One request in, a stream of tokens out, nothing from the client until it’s done: SSE is simpler, survives proxies better, and reconnects for free with <code class=\"language-text\">Last-Event-ID</code>.</li>\n<li>The client sends things mid-stream (cancel, edits, tool approvals, live cursors): a WebSocket earns its keep.</li>\n</ul>\n<p>An LLM chat with a stop button, live typing indicators, or interactive tool confirmation falls in the second case. That is the scenario this post builds.</p>\n<h2>The basic endpoint</h2>\n<p>FastAPI exposes WebSockets through <a href=\"https://www.starlette.io/websockets/\">Starlette</a>, the toolkit FastAPI is built on. The handler is an <code class=\"language-text\">async def</code> that accepts the connection, then reads and writes messages until someone hangs up. This minimal version just acknowledges each message it receives:</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\"><span class=\"token keyword\">from</span> fastapi <span class=\"token keyword\">import</span> FastAPI<span class=\"token punctuation\">,</span> WebSocket<span class=\"token punctuation\">,</span> WebSocketDisconnect\n\napp <span class=\"token operator\">=</span> FastAPI<span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span>\n\n<span class=\"token decorator annotation punctuation\">@app<span class=\"token punctuation\">.</span>websocket</span><span class=\"token punctuation\">(</span><span class=\"token string\">\"/ws/chat\"</span><span class=\"token punctuation\">)</span>\n<span class=\"token keyword\">async</span> <span class=\"token keyword\">def</span> <span class=\"token function\">chat</span><span class=\"token punctuation\">(</span>ws<span class=\"token punctuation\">:</span> WebSocket<span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n    <span class=\"token keyword\">await</span> ws<span class=\"token punctuation\">.</span>accept<span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span>\n    <span class=\"token keyword\">try</span><span class=\"token punctuation\">:</span>\n        <span class=\"token keyword\">while</span> <span class=\"token boolean\">True</span><span class=\"token punctuation\">:</span>\n            msg <span class=\"token operator\">=</span> <span class=\"token keyword\">await</span> ws<span class=\"token punctuation\">.</span>receive_json<span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span>\n            <span class=\"token keyword\">await</span> ws<span class=\"token punctuation\">.</span>send_json<span class=\"token punctuation\">(</span><span class=\"token punctuation\">{</span><span class=\"token string\">\"type\"</span><span class=\"token punctuation\">:</span> <span class=\"token string\">\"ack\"</span><span class=\"token punctuation\">,</span> <span class=\"token string\">\"id\"</span><span class=\"token punctuation\">:</span> msg<span class=\"token punctuation\">[</span><span class=\"token string\">\"id\"</span><span class=\"token punctuation\">]</span><span class=\"token punctuation\">}</span><span class=\"token punctuation\">)</span>\n    <span class=\"token keyword\">except</span> WebSocketDisconnect<span class=\"token punctuation\">:</span>\n        <span class=\"token comment\"># client closed the tab or lost the network; nothing to clean up here</span>\n        <span class=\"token keyword\">pass</span></code></pre></div>\n<p><code class=\"language-text\">receive_json</code> waits until a frame arrives, and <code class=\"language-text\">WebSocketDisconnect</code> is raised when the peer goes away. That much is standard <a href=\"https://fastapi.tiangolo.com/advanced/websockets/\">FastAPI WebSocket usage</a>. The interesting part is what happens when one message kicks off a long-running token stream and you still want to hear the client during it.</p>\n<h2>Sending and receiving at the same time</h2>\n<p>Here is the mistake I made the first time. I read a message, ran an <code class=\"language-text\">async for</code> over the model’s tokens to send each one, then looped back to read the next message. While that <code class=\"language-text\">async for</code> runs, nothing is calling <code class=\"language-text\">receive</code>. A cancel frame from the browser just sits in the buffer until generation finishes on its own, which is exactly the moment you no longer care about it.</p>\n<p>The fix is to run two coroutines against the same socket. One produces tokens; the other listens for control frames. When a cancel arrives, the listener cancels the producer.</p>\n<p><img src=\"/de5e28890f224a6acc794c5cc7005cbe/duplex-tasks.svg\" alt=\"One WebSocket connection between a React client and a FastAPI worker. Tokens stream down while a cancel message travels up. On the server, two asyncio tasks share the socket: a producer that pushes LLM tokens, and a receiver that watches for a cancel frame and calls cancel on the producer.\"></p>\n<p>In the code, each incoming chat message starts a producer task and a watcher task, and the handler waits for whichever finishes first:</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\"><span class=\"token keyword\">import</span> asyncio\n<span class=\"token keyword\">from</span> fastapi <span class=\"token keyword\">import</span> WebSocket<span class=\"token punctuation\">,</span> WebSocketDisconnect\n\n<span class=\"token keyword\">async</span> <span class=\"token keyword\">def</span> <span class=\"token function\">chat</span><span class=\"token punctuation\">(</span>ws<span class=\"token punctuation\">:</span> WebSocket<span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n    <span class=\"token keyword\">await</span> ws<span class=\"token punctuation\">.</span>accept<span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span>\n    <span class=\"token keyword\">try</span><span class=\"token punctuation\">:</span>\n        <span class=\"token keyword\">while</span> <span class=\"token boolean\">True</span><span class=\"token punctuation\">:</span>\n            req <span class=\"token operator\">=</span> <span class=\"token keyword\">await</span> ws<span class=\"token punctuation\">.</span>receive_json<span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span>\n            <span class=\"token keyword\">if</span> req<span class=\"token punctuation\">.</span>get<span class=\"token punctuation\">(</span><span class=\"token string\">\"type\"</span><span class=\"token punctuation\">)</span> <span class=\"token operator\">!=</span> <span class=\"token string\">\"message\"</span><span class=\"token punctuation\">:</span>\n                <span class=\"token keyword\">continue</span>\n\n            producer <span class=\"token operator\">=</span> asyncio<span class=\"token punctuation\">.</span>create_task<span class=\"token punctuation\">(</span>stream_answer<span class=\"token punctuation\">(</span>ws<span class=\"token punctuation\">,</span> req<span class=\"token punctuation\">)</span><span class=\"token punctuation\">)</span>\n            watcher <span class=\"token operator\">=</span> asyncio<span class=\"token punctuation\">.</span>create_task<span class=\"token punctuation\">(</span>watch_for_cancel<span class=\"token punctuation\">(</span>ws<span class=\"token punctuation\">,</span> producer<span class=\"token punctuation\">)</span><span class=\"token punctuation\">)</span>\n\n            <span class=\"token comment\"># whichever finishes first, tear down the other</span>\n            done<span class=\"token punctuation\">,</span> pending <span class=\"token operator\">=</span> <span class=\"token keyword\">await</span> asyncio<span class=\"token punctuation\">.</span>wait<span class=\"token punctuation\">(</span>\n                <span class=\"token punctuation\">{</span>producer<span class=\"token punctuation\">,</span> watcher<span class=\"token punctuation\">}</span><span class=\"token punctuation\">,</span>\n                return_when<span class=\"token operator\">=</span>asyncio<span class=\"token punctuation\">.</span>FIRST_COMPLETED<span class=\"token punctuation\">,</span>\n            <span class=\"token punctuation\">)</span>\n            <span class=\"token keyword\">for</span> task <span class=\"token keyword\">in</span> pending<span class=\"token punctuation\">:</span>\n                task<span class=\"token punctuation\">.</span>cancel<span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span>\n            <span class=\"token keyword\">await</span> asyncio<span class=\"token punctuation\">.</span>gather<span class=\"token punctuation\">(</span><span class=\"token operator\">*</span>pending<span class=\"token punctuation\">,</span> return_exceptions<span class=\"token operator\">=</span><span class=\"token boolean\">True</span><span class=\"token punctuation\">)</span>\n    <span class=\"token keyword\">except</span> WebSocketDisconnect<span class=\"token punctuation\">:</span>\n        <span class=\"token keyword\">pass</span>\n\n\n<span class=\"token keyword\">async</span> <span class=\"token keyword\">def</span> <span class=\"token function\">stream_answer</span><span class=\"token punctuation\">(</span>ws<span class=\"token punctuation\">:</span> WebSocket<span class=\"token punctuation\">,</span> req<span class=\"token punctuation\">:</span> <span class=\"token builtin\">dict</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n    <span class=\"token keyword\">async</span> <span class=\"token keyword\">for</span> token <span class=\"token keyword\">in</span> llm_stream<span class=\"token punctuation\">(</span>req<span class=\"token punctuation\">[</span><span class=\"token string\">\"text\"</span><span class=\"token punctuation\">]</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>   <span class=\"token comment\"># your model client</span>\n        <span class=\"token keyword\">await</span> ws<span class=\"token punctuation\">.</span>send_json<span class=\"token punctuation\">(</span><span class=\"token punctuation\">{</span><span class=\"token string\">\"type\"</span><span class=\"token punctuation\">:</span> <span class=\"token string\">\"token\"</span><span class=\"token punctuation\">,</span> <span class=\"token string\">\"text\"</span><span class=\"token punctuation\">:</span> token<span class=\"token punctuation\">}</span><span class=\"token punctuation\">)</span>\n    <span class=\"token keyword\">await</span> ws<span class=\"token punctuation\">.</span>send_json<span class=\"token punctuation\">(</span><span class=\"token punctuation\">{</span><span class=\"token string\">\"type\"</span><span class=\"token punctuation\">:</span> <span class=\"token string\">\"done\"</span><span class=\"token punctuation\">}</span><span class=\"token punctuation\">)</span>\n\n\n<span class=\"token keyword\">async</span> <span class=\"token keyword\">def</span> <span class=\"token function\">watch_for_cancel</span><span class=\"token punctuation\">(</span>ws<span class=\"token punctuation\">:</span> WebSocket<span class=\"token punctuation\">,</span> producer<span class=\"token punctuation\">:</span> asyncio<span class=\"token punctuation\">.</span>Task<span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n    <span class=\"token keyword\">while</span> <span class=\"token boolean\">True</span><span class=\"token punctuation\">:</span>\n        msg <span class=\"token operator\">=</span> <span class=\"token keyword\">await</span> ws<span class=\"token punctuation\">.</span>receive_json<span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span>\n        <span class=\"token keyword\">if</span> msg<span class=\"token punctuation\">.</span>get<span class=\"token punctuation\">(</span><span class=\"token string\">\"type\"</span><span class=\"token punctuation\">)</span> <span class=\"token operator\">==</span> <span class=\"token string\">\"cancel\"</span><span class=\"token punctuation\">:</span>\n            producer<span class=\"token punctuation\">.</span>cancel<span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span>\n            <span class=\"token keyword\">return</span></code></pre></div>\n<p>Two pieces are load-bearing here:</p>\n<ul>\n<li><strong><code class=\"language-text\">asyncio.wait</code> with <code class=\"language-text\">FIRST_COMPLETED</code>.</strong> Both the normal case (the answer finishes) and the interrupt case (a cancel comes in) end the round, and the loser gets cancelled either way.</li>\n<li><strong><code class=\"language-text\">return_exceptions=True</code> on the <code class=\"language-text\">gather</code>.</strong> It swallows the <code class=\"language-text\">CancelledError</code> that the cancelled task raises, so it doesn’t crash the handler.</li>\n</ul>\n<p>This is the server-side mirror of what I did on the frontend with an <a href=\"/blog/2026-07-16-cancel-llm-stream-react-abortcontroller/\">AbortController to cancel a stream</a>. The difference is that the cancel signal now rides the same connection as the tokens instead of firing off a separate request.</p>\n<p>One subtlety: cancelling the producer task stops your loop from sending more tokens, but your model client may still hold an open HTTP request to the provider. Make sure the client’s context manager sits inside <code class=\"language-text\">stream_answer</code>, so the <code class=\"language-text\">CancelledError</code> propagates into it and closes the request. Otherwise you stop paying attention to the response while still paying for it.</p>\n<h2>Heartbeats, because TCP won’t tell you a connection is dead</h2>\n<p>A WebSocket sits on a TCP connection, and TCP will happily believe a connection is alive long after the client’s laptop went into a tunnel. If you never write to a dead socket, you may never learn it’s dead, and the coroutine holding it leaks. On top of that, most reverse proxies close idle connections on a timer: nginx defaults <code class=\"language-text\">proxy_read_timeout</code> to 60 seconds, and AWS load balancers idle out around the same mark.</p>\n<p>The protocol has ping/pong frames for exactly this. Send a ping on an interval, and if no pong comes back within a timeout, treat the connection as gone and close it. A small keepalive task running alongside the others does the job:</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\"><span class=\"token keyword\">async</span> <span class=\"token keyword\">def</span> <span class=\"token function\">keepalive</span><span class=\"token punctuation\">(</span>ws<span class=\"token punctuation\">:</span> WebSocket<span class=\"token punctuation\">,</span> interval<span class=\"token punctuation\">:</span> <span class=\"token builtin\">float</span> <span class=\"token operator\">=</span> <span class=\"token number\">20.0</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n    <span class=\"token keyword\">while</span> <span class=\"token boolean\">True</span><span class=\"token punctuation\">:</span>\n        <span class=\"token keyword\">await</span> asyncio<span class=\"token punctuation\">.</span>sleep<span class=\"token punctuation\">(</span>interval<span class=\"token punctuation\">)</span>\n        <span class=\"token comment\"># a lightweight app-level ping; the client answers with a pong frame</span>\n        <span class=\"token keyword\">await</span> ws<span class=\"token punctuation\">.</span>send_json<span class=\"token punctuation\">(</span><span class=\"token punctuation\">{</span><span class=\"token string\">\"type\"</span><span class=\"token punctuation\">:</span> <span class=\"token string\">\"ping\"</span><span class=\"token punctuation\">}</span><span class=\"token punctuation\">)</span></code></pre></div>\n<p>The point is to keep the interval under the proxy’s idle timeout. Twenty seconds against a 60-second proxy timeout leaves margin.</p>\n<p>If you terminate the socket at the proxy and only speak HTTP to the app, confirm the proxy is configured for WebSocket upgrades at all. A proxy that doesn’t pass through the <code class=\"language-text\">Upgrade</code>/<code class=\"language-text\">Connection</code> headers is the classic reason a connection works locally and dies behind nginx.</p>\n<h2>Authenticating the handshake</h2>\n<p>Browsers won’t let you set arbitrary headers on a <code class=\"language-text\">WebSocket</code>, so the <code class=\"language-text\">Authorization: Bearer ...</code> pattern you use for REST doesn’t carry over. You have two realistic options:</p>\n<ul>\n<li>Pass a short-lived token as a query parameter.</li>\n<li>Use the <code class=\"language-text\">Sec-WebSocket-Protocol</code> subprotocol header, which the browser <em>does</em> let you set.</li>\n</ul>\n<p>Either way, validate before you call <code class=\"language-text\">accept()</code>, and close with a policy-violation code if validation fails. This version reads the token from the query string:</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\"><span class=\"token keyword\">from</span> fastapi <span class=\"token keyword\">import</span> WebSocket<span class=\"token punctuation\">,</span> status\n\n<span class=\"token decorator annotation punctuation\">@app<span class=\"token punctuation\">.</span>websocket</span><span class=\"token punctuation\">(</span><span class=\"token string\">\"/ws/chat\"</span><span class=\"token punctuation\">)</span>\n<span class=\"token keyword\">async</span> <span class=\"token keyword\">def</span> <span class=\"token function\">chat</span><span class=\"token punctuation\">(</span>ws<span class=\"token punctuation\">:</span> WebSocket<span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n    token <span class=\"token operator\">=</span> ws<span class=\"token punctuation\">.</span>query_params<span class=\"token punctuation\">.</span>get<span class=\"token punctuation\">(</span><span class=\"token string\">\"token\"</span><span class=\"token punctuation\">)</span>\n    user <span class=\"token operator\">=</span> <span class=\"token keyword\">await</span> verify_token<span class=\"token punctuation\">(</span>token<span class=\"token punctuation\">)</span>         <span class=\"token comment\"># your own check</span>\n    <span class=\"token keyword\">if</span> user <span class=\"token keyword\">is</span> <span class=\"token boolean\">None</span><span class=\"token punctuation\">:</span>\n        <span class=\"token keyword\">await</span> ws<span class=\"token punctuation\">.</span>close<span class=\"token punctuation\">(</span>code<span class=\"token operator\">=</span>status<span class=\"token punctuation\">.</span>WS_1008_POLICY_VIOLATION<span class=\"token punctuation\">)</span>\n        <span class=\"token keyword\">return</span>\n    <span class=\"token keyword\">await</span> ws<span class=\"token punctuation\">.</span>accept<span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span>\n    <span class=\"token punctuation\">.</span><span class=\"token punctuation\">.</span><span class=\"token punctuation\">.</span></code></pre></div>\n<p>Prefer a short-lived token minted for this purpose over your main session token. Query strings land in access logs and proxy logs, so a token that expires in a minute is a smaller thing to leak than one good for a week.</p>\n<h2>What breaks when you add a second worker</h2>\n<p>This is the part that surprises people. Everything above works on one Uvicorn worker. The moment you scale to several workers, or several pods, the design falls apart, because a WebSocket’s state lives in the process that accepted it. Two consequences follow:</p>\n<ul>\n<li><strong>Routing.</strong> If you load-balance without sticky routing, a client’s frames can land on a worker that never saw its connection. You need the balancer to pin a connection to a backend, or you accept that any given socket is owned by exactly one process and route accordingly.</li>\n<li><strong>Reaching other users.</strong> If one user’s action needs to reach another user (a shared room, a broadcast, a “someone else is typing” signal), the worker holding <em>your</em> socket has no way to reach the worker holding <em>theirs</em>. The in-memory <code class=\"language-text\">set()</code> of connections that every tutorial shows only sees the sockets on the local process.</li>\n</ul>\n<p>The standard fix is a message broker that every worker subscribes to. Redis pub/sub is the common choice: a worker publishes an event, every worker receives it on its subscription, and each one forwards it to the local sockets that care. The <a href=\"https://redis.io/docs/latest/develop/interact/pubsub/\">Redis pub/sub docs</a> cover the primitive. In FastAPI, the shape is a background subscriber task per process that fans messages out to that process’s own connections.</p>\n<p>For plain one-user token streaming, you may never need this. For anything multi-user, design for it up front, because retrofitting a broadcast layer after you have assumed a single process is a rewrite.</p>\n<h2>Backpressure and slow clients</h2>\n<p>When <code class=\"language-text\">send_json</code> returns, the frame is only queued; the client hasn’t necessarily received it. If the model generates faster than a phone on hotel wifi can drain the socket, that queue grows in your process memory. This is backpressure: the consumer can’t keep up, so something on the sending side has to give.</p>\n<p>For token-by-token text this is rarely fatal. If you stream large chunks, cap it. Keep a bounded buffer, and if it fills, either coalesce tokens into bigger, less frequent frames, or drop the connection rather than let one slow client bloat a worker. If you run this in production, send-queue depth is worth a metric.</p>\n<h2>Tradeoffs, and what I’d reach for first</h2>\n<p>I don’t default to WebSockets. For the common case (one prompt, one streamed answer, nothing from the client until it finishes), SSE is less to get wrong. It also degrades more gracefully through corporate proxies that mangle upgrades, and the reconnect-with-<code class=\"language-text\">Last-Event-ID</code> story alone saves real code.</p>\n<p>WebSockets win when the interaction is genuinely two-way during a single turn: a stop button that must land immediately, tool-use confirmations the user approves inline, collaborative sessions, or a UI that streams state in both directions. The cost is everything above: heartbeats you have to run yourself, auth that can’t use headers, and a broker the day you outgrow one process. Pay it when the interaction demands it, not before.</p>\n<p>I lean on this pattern in the interactive AI work on my portfolio. In the split-panel chat in <a href=\"/project/cloud-canvas-ai/\">CloudCanvasAI</a>, a document renders as the model writes it. In the operator copilot in <a href=\"/project/archi/\">Archi</a>, a running answer sometimes needs to be cut short. In both, the deciding factor was the same: the client had something to say before the server was done talking, and a single full-duplex connection was the honest way to carry it.</p>\n<hr>\n<p><em>Further reading: the <a href=\"https://fastapi.tiangolo.com/advanced/websockets/\">FastAPI WebSockets guide</a>, <a href=\"https://www.starlette.io/websockets/\">Starlette’s WebSocket API</a>, and <a href=\"https://datatracker.ietf.org/doc/html/rfc6455\">RFC 6455</a> for the protocol itself.</em></p>","frontmatter":{"title":"FastAPI WebSockets for Streaming LLM Chat","date":"2026-09-12T00:00:00.000Z","description":"SSE streams one direction only. When an LLM chat needs the browser to interrupt a running response, FastAPI WebSockets give you a full-duplex channel.","thumbnail":{"childImageSharp":{"fluid":{"base64":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAABQAAAALCAIAAADwazoUAAAACXBIWXMAAAsSAAALEgHS3X78AAABtUlEQVQoz03Oa3OiMBQGYD5VRQlBJORiEiCRq9KlRVtvrd3tzn7Y//9/9ljcjjPPnAnhfZM4jTagkumubn+9HD62r5f+5eduD/Ptafv5ejx3/fvzDha/D6e+XNcqGyrAGfkKPCBJ0x95ewamOazas92czfpUdh9ZvTebo92cqu7C7TMkr/mvljMLFPDmGsdFyOto2fgkd4Nk4i9dLCdIzMJsIRoiNzBxnENyqABniuVVoKCjVn39+KZXWyLqkGR+qK8iq2xvm4PMd4isbvkvjgs33Iixx3CoGbdcFJznghnJQQaLVFjJ0gU1KNTfFSiLwcwX8FTDlKUiISKJhY5FQiXsSKb/dE/7skLEBlH6XXFcxEGn2DFXWUxbW3ZFUyizTm2TWMuVoUtNxKVZ10I6Y+J6dKgAZ4IYnLHKqtKUUuWQLpZJlRVlWhRpUWZlbSuTln9P721mH5DAC+X6HFrAmXgUjBEfefDNg4CHAfM94iMCc47pYs5BYjtlHpluFsxMfTa0oBzfmyHqY0aJojHQMVFRJDFm8NopFihU15tmZAg7sLo3mhIX0Zhq+l9EFOyMbn+j+/A/YwZGwBp7r5QAAAAASUVORK5CYII=","aspectRatio":1.899441340782123,"src":"/static/f592754569a19a404a61ee0fa9a91ca8/40a76/hero.png","srcSet":"/static/f592754569a19a404a61ee0fa9a91ca8/c972b/hero.png 340w,\n/static/f592754569a19a404a61ee0fa9a91ca8/27625/hero.png 680w,\n/static/f592754569a19a404a61ee0fa9a91ca8/40a76/hero.png 1360w,\n/static/f592754569a19a404a61ee0fa9a91ca8/ed396/hero.png 2000w","sizes":"(max-width: 1360px) 100vw, 1360px"}}}}}},"pageContext":{"slug":"/2026-09-12-fastapi-websockets-streaming-llm-chat/","previous":"blog/2026-09-13-agentic-rag-when-to-retrieve/","next":"blog/2026-09-16-fastapi-pydantic-v2-request-validation/"}},"staticQueryHashes":["32046230"]}