{"componentChunkName":"component---src-templates-blog-post-js","path":"/blog/2026-06-30-fastapi-sse-streaming-llm/","result":{"data":{"site":{"siteMetadata":{"title":"M.Hassan Ahmed","author":"Hassan11196"}},"markdownRemark":{"id":"81828bec-9bbf-5ba7-a2b6-7951f5b7175c","excerpt":"A language model takes a few seconds to write a paragraph. If your API waits for the full completion and returns it in one shot, the user stares at a spinner…","html":"<p>A language model takes a few seconds to write a paragraph. If your API waits for the full completion and returns it in one shot, the user stares at a spinner the whole time, and then the answer appears all at once. The model was producing words the entire time; you just hid them. Streaming gives that time back: the first token shows up in a few hundred milliseconds, and the rest fill in as the model writes.</p>\n<p>This post is for engineers who have a <a href=\"https://fastapi.tiangolo.com/\">FastAPI</a> backend in front of an LLM, plus a JavaScript frontend, and want the response to land word by word. It covers the server endpoint, the adapter that pulls text out of the model, the browser client, and the problems that appear after the demo works: proxy buffering, idle timeouts, errors mid-stream, and the limits of <code class=\"language-text\">EventSource</code>.</p>\n<p>The transport is Server-Sent Events (SSE). I reached for it when I built the streaming chat panel in <a href=\"/project/cloud-canvas-ai/\">CloudCanvasAI</a>, where you talk to Claude on the left and a document renders live on the right. The same shape works for any provider; <a href=\"/project/gemini-alchemy/\">Gemini Alchemy</a> is another FastAPI backend of mine that streams the same way.</p>\n<h2>Why SSE instead of WebSockets</h2>\n<p>LLM streaming is one-directional. The client sends one prompt, and the server pushes a long sequence of tokens back. Nothing flows the other way until the next turn. That is exactly the shape <a href=\"https://developer.mozilla.org/en-US/docs/Web/API/Server-sent_events\">Server-Sent Events</a> was designed for: a one-way channel from server to browser over a single long-lived HTTP response.</p>\n<p>WebSockets give you a full-duplex socket, where both sides can send at any time. That is more than this job needs, and it costs more to operate. A WebSocket leaves HTTP behind, so your auth middleware, logging, and load balancer rules all need a second path. SSE is plain HTTP with a particular content type, so everything that already understands a request keeps working. The browser side is a few lines, and reconnection is built into the <a href=\"https://html.spec.whatwg.org/multipage/server-sent-events.html#the-eventsource-interface\"><code class=\"language-text\">EventSource</code></a> interface instead of being something you write yourself.</p>\n<p>The diagram below shows the whole system. A request goes out, and tokens come back on the same connection until the server signals that it is done.</p>\n<p><img src=\"/9be7e772d476f8f289103902ca00b5b4/sse-flow.svg\" alt=\"SSE streaming pipeline: the LLM emits token deltas, FastAPI relays them as a text/event-stream, and the browser EventSource appends each one to the React UI\"></p>\n<p>The bottom half of the diagram is the part people skip and then end up debugging. SSE is a concrete wire format. Each event is one or more lines of text, and a blank line ends it. Get the blank line wrong and the browser buffers your whole stream, waiting for an event that never closes.</p>\n<h2>The server: an async generator behind StreamingResponse</h2>\n<p>FastAPI can stream whatever an async generator yields. (An async generator is an <code class=\"language-text\">async def</code> function that uses <code class=\"language-text\">yield</code> to produce values one at a time.) Wrap the generator in a <a href=\"https://fastapi.tiangolo.com/advanced/custom-response/#streamingresponse\"><code class=\"language-text\">StreamingResponse</code></a> with the <code class=\"language-text\">text/event-stream</code> media type, and you have a working SSE endpoint. The handler below uses a small <code class=\"language-text\">sse</code> helper to format each message, sends model output as <code class=\"language-text\">token</code> events, and ends with a <code class=\"language-text\">done</code> event.</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\"><span class=\"token keyword\">import</span> json\n<span class=\"token keyword\">from</span> fastapi <span class=\"token keyword\">import</span> FastAPI\n<span class=\"token keyword\">from</span> fastapi<span class=\"token punctuation\">.</span>responses <span class=\"token keyword\">import</span> StreamingResponse\n\napp <span class=\"token operator\">=</span> FastAPI<span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span>\n\n<span class=\"token keyword\">def</span> <span class=\"token function\">sse</span><span class=\"token punctuation\">(</span>event<span class=\"token punctuation\">:</span> <span class=\"token builtin\">str</span><span class=\"token punctuation\">,</span> data<span class=\"token punctuation\">:</span> <span class=\"token builtin\">dict</span><span class=\"token punctuation\">)</span> <span class=\"token operator\">-</span><span class=\"token operator\">></span> <span class=\"token builtin\">str</span><span class=\"token punctuation\">:</span>\n    <span class=\"token triple-quoted-string string\">\"\"\"Format one SSE message. The trailing blank line ends the event.\"\"\"</span>\n    <span class=\"token keyword\">return</span> <span class=\"token string-interpolation\"><span class=\"token string\">f\"event: </span><span class=\"token interpolation\"><span class=\"token punctuation\">{</span>event<span class=\"token punctuation\">}</span></span><span class=\"token string\">\\ndata: </span><span class=\"token interpolation\"><span class=\"token punctuation\">{</span>json<span class=\"token punctuation\">.</span>dumps<span class=\"token punctuation\">(</span>data<span class=\"token punctuation\">)</span><span class=\"token punctuation\">}</span></span><span class=\"token string\">\\n\\n\"</span></span>\n\n<span class=\"token keyword\">async</span> <span class=\"token keyword\">def</span> <span class=\"token function\">token_stream</span><span class=\"token punctuation\">(</span>prompt<span class=\"token punctuation\">:</span> <span class=\"token builtin\">str</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n    <span class=\"token keyword\">async</span> <span class=\"token keyword\">for</span> delta <span class=\"token keyword\">in</span> call_model<span class=\"token punctuation\">(</span>prompt<span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>   <span class=\"token comment\"># your provider's streaming call</span>\n        <span class=\"token keyword\">yield</span> sse<span class=\"token punctuation\">(</span><span class=\"token string\">\"token\"</span><span class=\"token punctuation\">,</span> <span class=\"token punctuation\">{</span><span class=\"token string\">\"text\"</span><span class=\"token punctuation\">:</span> delta<span class=\"token punctuation\">}</span><span class=\"token punctuation\">)</span>\n    <span class=\"token keyword\">yield</span> sse<span class=\"token punctuation\">(</span><span class=\"token string\">\"done\"</span><span class=\"token punctuation\">,</span> <span class=\"token punctuation\">{</span><span class=\"token punctuation\">}</span><span class=\"token punctuation\">)</span>\n\n<span class=\"token decorator annotation punctuation\">@app<span class=\"token punctuation\">.</span>get</span><span class=\"token punctuation\">(</span><span class=\"token string\">\"/chat\"</span><span class=\"token punctuation\">)</span>\n<span class=\"token keyword\">async</span> <span class=\"token keyword\">def</span> <span class=\"token function\">chat</span><span class=\"token punctuation\">(</span>prompt<span class=\"token punctuation\">:</span> <span class=\"token builtin\">str</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n    <span class=\"token keyword\">return</span> StreamingResponse<span class=\"token punctuation\">(</span>\n        token_stream<span class=\"token punctuation\">(</span>prompt<span class=\"token punctuation\">)</span><span class=\"token punctuation\">,</span>\n        media_type<span class=\"token operator\">=</span><span class=\"token string\">\"text/event-stream\"</span><span class=\"token punctuation\">,</span>\n        headers<span class=\"token operator\">=</span><span class=\"token punctuation\">{</span>\n            <span class=\"token string\">\"Cache-Control\"</span><span class=\"token punctuation\">:</span> <span class=\"token string\">\"no-cache\"</span><span class=\"token punctuation\">,</span>\n            <span class=\"token string\">\"X-Accel-Buffering\"</span><span class=\"token punctuation\">:</span> <span class=\"token string\">\"no\"</span><span class=\"token punctuation\">,</span>   <span class=\"token comment\"># tell nginx not to buffer</span>\n        <span class=\"token punctuation\">}</span><span class=\"token punctuation\">,</span>\n    <span class=\"token punctuation\">)</span></code></pre></div>\n<p>Two details in that handler do real work:</p>\n<ul>\n<li><strong>The <code class=\"language-text\">\\n\\n</code> at the end of every message</strong> is the event boundary. Without it, the client never sees a complete event.</li>\n<li><strong>The <code class=\"language-text\">X-Accel-Buffering: no</code> header</strong> is meant for the reverse proxy. Leaving it out causes the single most common “streaming doesn’t stream” bug, which the proxy section below explains.</li>\n</ul>\n<p>If you would rather not format the wire bytes by hand, <a href=\"https://github.com/sysid/sse_starlette\"><code class=\"language-text\">sse-starlette</code></a> provides an <code class=\"language-text\">EventSourceResponse</code> that takes the same async generator. It handles the framing and sends periodic pings. On a small service, the manual version above is fine and keeps the dependency list short.</p>\n<h2>Getting text deltas out of the model</h2>\n<p>Every major SDK exposes its streaming output as an async iterator of deltas, meaning the small chunks of new text the model produced since the previous chunk. With Anthropic’s SDK, the <a href=\"https://platform.claude.com/docs/en/api/messages-streaming\">streaming API</a> looks like this:</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\"><span class=\"token keyword\">from</span> anthropic <span class=\"token keyword\">import</span> AsyncAnthropic\n\nclient <span class=\"token operator\">=</span> AsyncAnthropic<span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span>\n\n<span class=\"token keyword\">async</span> <span class=\"token keyword\">def</span> <span class=\"token function\">call_model</span><span class=\"token punctuation\">(</span>prompt<span class=\"token punctuation\">:</span> <span class=\"token builtin\">str</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n    <span class=\"token keyword\">async</span> <span class=\"token keyword\">with</span> client<span class=\"token punctuation\">.</span>messages<span class=\"token punctuation\">.</span>stream<span class=\"token punctuation\">(</span>\n        model<span class=\"token operator\">=</span><span class=\"token string\">\"claude-sonnet-4-6\"</span><span class=\"token punctuation\">,</span>\n        max_tokens<span class=\"token operator\">=</span><span class=\"token number\">1024</span><span class=\"token punctuation\">,</span>\n        messages<span class=\"token operator\">=</span><span class=\"token punctuation\">[</span><span class=\"token punctuation\">{</span><span class=\"token string\">\"role\"</span><span class=\"token punctuation\">:</span> <span class=\"token string\">\"user\"</span><span class=\"token punctuation\">,</span> <span class=\"token string\">\"content\"</span><span class=\"token punctuation\">:</span> prompt<span class=\"token punctuation\">}</span><span class=\"token punctuation\">]</span><span class=\"token punctuation\">,</span>\n    <span class=\"token punctuation\">)</span> <span class=\"token keyword\">as</span> stream<span class=\"token punctuation\">:</span>\n        <span class=\"token keyword\">async</span> <span class=\"token keyword\">for</span> text <span class=\"token keyword\">in</span> stream<span class=\"token punctuation\">.</span>text_stream<span class=\"token punctuation\">:</span>\n            <span class=\"token keyword\">yield</span> text</code></pre></div>\n<p><code class=\"language-text\">call_model</code> opens the stream and passes each piece of text from <code class=\"language-text\">text_stream</code> straight through. Other providers follow the same pattern; only the method names change.</p>\n<p>Keep this adapter separate from the SSE formatting. The endpoint should not care which model is behind it, and you will want to swap models without touching the transport. That separation is what let CloudCanvasAI route different document tasks through different skills while the streaming endpoint stayed the same.</p>\n<h2>The browser client: EventSource</h2>\n<p>In the browser, the consumer is small. Open an <code class=\"language-text\">EventSource</code>, listen for your named events (<code class=\"language-text\">token</code> and <code class=\"language-text\">done</code>), and append each token as it arrives.</p>\n<div class=\"gatsby-highlight\" data-language=\"javascript\"><pre class=\"language-javascript\"><code class=\"language-javascript\"><span class=\"token keyword\">const</span> source <span class=\"token operator\">=</span> <span class=\"token keyword\">new</span> <span class=\"token class-name\">EventSource</span><span class=\"token punctuation\">(</span><span class=\"token template-string\"><span class=\"token template-punctuation string\">`</span><span class=\"token string\">/chat?prompt=</span><span class=\"token interpolation\"><span class=\"token interpolation-punctuation punctuation\">${</span><span class=\"token function\">encodeURIComponent</span><span class=\"token punctuation\">(</span>prompt<span class=\"token punctuation\">)</span><span class=\"token interpolation-punctuation punctuation\">}</span></span><span class=\"token template-punctuation string\">`</span></span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">;</span>\n\nsource<span class=\"token punctuation\">.</span><span class=\"token function\">addEventListener</span><span class=\"token punctuation\">(</span><span class=\"token string\">\"token\"</span><span class=\"token punctuation\">,</span> <span class=\"token punctuation\">(</span><span class=\"token parameter\">e</span><span class=\"token punctuation\">)</span> <span class=\"token operator\">=></span> <span class=\"token punctuation\">{</span>\n  <span class=\"token keyword\">const</span> <span class=\"token punctuation\">{</span> text <span class=\"token punctuation\">}</span> <span class=\"token operator\">=</span> <span class=\"token constant\">JSON</span><span class=\"token punctuation\">.</span><span class=\"token function\">parse</span><span class=\"token punctuation\">(</span>e<span class=\"token punctuation\">.</span>data<span class=\"token punctuation\">)</span><span class=\"token punctuation\">;</span>\n  <span class=\"token function\">appendToMessage</span><span class=\"token punctuation\">(</span>text<span class=\"token punctuation\">)</span><span class=\"token punctuation\">;</span>          <span class=\"token comment\">// your state update</span>\n<span class=\"token punctuation\">}</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">;</span>\n\nsource<span class=\"token punctuation\">.</span><span class=\"token function\">addEventListener</span><span class=\"token punctuation\">(</span><span class=\"token string\">\"done\"</span><span class=\"token punctuation\">,</span> <span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span> <span class=\"token operator\">=></span> <span class=\"token punctuation\">{</span>\n  source<span class=\"token punctuation\">.</span><span class=\"token function\">close</span><span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">;</span>                 <span class=\"token comment\">// stop, or the browser will reconnect</span>\n<span class=\"token punctuation\">}</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">;</span>\n\nsource<span class=\"token punctuation\">.</span><span class=\"token function-variable function\">onerror</span> <span class=\"token operator\">=</span> <span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span> <span class=\"token operator\">=></span> <span class=\"token punctuation\">{</span>\n  source<span class=\"token punctuation\">.</span><span class=\"token function\">close</span><span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">;</span>                 <span class=\"token comment\">// handle the failure, then decide to retry</span>\n<span class=\"token punctuation\">}</span><span class=\"token punctuation\">;</span></code></pre></div>\n<p>Always call <code class=\"language-text\">close()</code> on <code class=\"language-text\">done</code>. If you do not, <code class=\"language-text\">EventSource</code> treats the closed connection as a dropped one and reconnects automatically, which sends your whole request again. Auto-reconnect is useful for genuine network blips. It is annoying when it re-runs a paid completion because you forgot to close a finished stream.</p>\n<h2>Why streaming breaks behind nginx: proxy buffering</h2>\n<p>You write all of the above, it streams perfectly against <code class=\"language-text\">localhost</code>, you deploy it behind nginx, and the tokens arrive in one lump at the end. Nothing in your code changed. The reverse proxy (the server in front of your app that forwards requests to it) is buffering the response. It collects the whole body before passing it on, which defeats the point.</p>\n<p>The fix has two sides:</p>\n<ul>\n<li><strong>In the app:</strong> the <code class=\"language-text\">X-Accel-Buffering: no</code> response header from the handler tells nginx to pass this one response straight through.</li>\n<li><strong>In nginx:</strong> the streaming location also needs <code class=\"language-text\">proxy_buffering off;</code>.</li>\n</ul>\n<p>Both settings exist because buffering is the right default for normal responses and the wrong one for a stream, so you have to opt out explicitly. The behavior is documented in the <a href=\"https://nginx.org/en/docs/http/ngx_http_proxy_module.html#proxy_buffering\">nginx proxy module</a>. Check it first whenever streaming works locally but not in production.</p>\n<h2>Heartbeats for quiet connections</h2>\n<p>A stream that goes quiet is ambiguous. The model might be thinking, or the connection might be dead, and the client cannot tell which. Proxies and load balancers also close idle connections after a timeout. A heartbeat handles both problems. Every fifteen seconds or so, send an SSE comment line: a line that begins with a colon, which the client ignores. The generator below sends one as soon as the stream opens, before the first token:</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\"><span class=\"token keyword\">async</span> <span class=\"token keyword\">def</span> <span class=\"token function\">token_stream</span><span class=\"token punctuation\">(</span>prompt<span class=\"token punctuation\">:</span> <span class=\"token builtin\">str</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n    <span class=\"token keyword\">yield</span> <span class=\"token string\">\": keep-alive\\n\\n\"</span>          <span class=\"token comment\"># comment line, ignored by EventSource</span>\n    <span class=\"token keyword\">async</span> <span class=\"token keyword\">for</span> delta <span class=\"token keyword\">in</span> call_model<span class=\"token punctuation\">(</span>prompt<span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n        <span class=\"token keyword\">yield</span> sse<span class=\"token punctuation\">(</span><span class=\"token string\">\"token\"</span><span class=\"token punctuation\">,</span> <span class=\"token punctuation\">{</span><span class=\"token string\">\"text\"</span><span class=\"token punctuation\">:</span> delta<span class=\"token punctuation\">}</span><span class=\"token punctuation\">)</span>\n    <span class=\"token keyword\">yield</span> sse<span class=\"token punctuation\">(</span><span class=\"token string\">\"done\"</span><span class=\"token punctuation\">,</span> <span class=\"token punctuation\">{</span><span class=\"token punctuation\">}</span><span class=\"token punctuation\">)</span></code></pre></div>\n<p>Once tokens are flowing, the data itself keeps the connection warm, so the heartbeat mostly matters in the gap before the first token, while the model is still reading a long prompt. If you use <code class=\"language-text\">sse-starlette</code>, its ping mechanism does this for you.</p>\n<h2>Handling errors in the middle of a stream</h2>\n<p>Once you have sent a <code class=\"language-text\">200 OK</code> and started streaming, you cannot take it back. If the model call throws on token five hundred, the HTTP status went out long ago. The honest option is to send the error as its own event and let the client react. The generator below catches the failure, emits an <code class=\"language-text\">error</code> event, and always finishes with <code class=\"language-text\">done</code>:</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\"><span class=\"token keyword\">async</span> <span class=\"token keyword\">def</span> <span class=\"token function\">token_stream</span><span class=\"token punctuation\">(</span>prompt<span class=\"token punctuation\">:</span> <span class=\"token builtin\">str</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n    <span class=\"token keyword\">try</span><span class=\"token punctuation\">:</span>\n        <span class=\"token keyword\">async</span> <span class=\"token keyword\">for</span> delta <span class=\"token keyword\">in</span> call_model<span class=\"token punctuation\">(</span>prompt<span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n            <span class=\"token keyword\">yield</span> sse<span class=\"token punctuation\">(</span><span class=\"token string\">\"token\"</span><span class=\"token punctuation\">,</span> <span class=\"token punctuation\">{</span><span class=\"token string\">\"text\"</span><span class=\"token punctuation\">:</span> delta<span class=\"token punctuation\">}</span><span class=\"token punctuation\">)</span>\n    <span class=\"token keyword\">except</span> Exception <span class=\"token keyword\">as</span> exc<span class=\"token punctuation\">:</span>\n        <span class=\"token keyword\">yield</span> sse<span class=\"token punctuation\">(</span><span class=\"token string\">\"error\"</span><span class=\"token punctuation\">,</span> <span class=\"token punctuation\">{</span><span class=\"token string\">\"message\"</span><span class=\"token punctuation\">:</span> <span class=\"token string\">\"generation failed\"</span><span class=\"token punctuation\">}</span><span class=\"token punctuation\">)</span>\n        log<span class=\"token punctuation\">.</span>exception<span class=\"token punctuation\">(</span><span class=\"token string\">\"stream failed: %s\"</span><span class=\"token punctuation\">,</span> exc<span class=\"token punctuation\">)</span>\n    <span class=\"token keyword\">finally</span><span class=\"token punctuation\">:</span>\n        <span class=\"token keyword\">yield</span> sse<span class=\"token punctuation\">(</span><span class=\"token string\">\"done\"</span><span class=\"token punctuation\">,</span> <span class=\"token punctuation\">{</span><span class=\"token punctuation\">}</span><span class=\"token punctuation\">)</span></code></pre></div>\n<p>The client listens for <code class=\"language-text\">error</code> the same way it listens for <code class=\"language-text\">token</code>. It then marks the partial message as failed, instead of leaving a half-written reply hanging. Log the real exception on the server, and send the user a generic message so a stack trace never goes out on the wire.</p>\n<h2>EventSource only does GET</h2>\n<p>The real limitation of <code class=\"language-text\">EventSource</code> is that it issues a GET and cannot set a request body or custom headers. For a short prompt, a query string is fine. For a long chat history, GET stops being workable. A bearer token does not belong in the URL either, because a token in a query string ends up in access logs.</p>\n<p>There are two ways out:</p>\n<ul>\n<li><strong>Keep the payload on the server.</strong> Pass an opaque session id in the URL and keep the real payload server-side.</li>\n<li><strong>Replace <code class=\"language-text\">EventSource</code> with <code class=\"language-text\">fetch</code>.</strong> Read the same <code class=\"language-text\">text/event-stream</code> response with <code class=\"language-text\">fetch</code> and a <code class=\"language-text\">ReadableStream</code>. This lets you POST a JSON body and set headers, while you parse the exact same event format yourself.</li>\n</ul>\n<p>I used the second option in CloudCanvasAI, because each turn carries context and auth that have no business sitting in a URL. You give up the built-in reconnect and parse the stream by hand. For an authenticated chat endpoint, that is the right trade.</p>\n<h2>Tradeoffs: what SSE can’t do, and what it costs to run</h2>\n<p>SSE is the smaller, calmer choice for this job, and the price is its narrowness. By design, it carries text in one direction. Streaming raw binary means base64 encoding, which bloats the payload. Anything genuinely bidirectional and low-latency, like a voice loop, is a WebSocket job. SSE earns its place when the interaction is request in, long text out, which covers most LLM chat and most agent output.</p>\n<p>The operational cost is that you now hold a connection open for the length of a generation. Long-lived requests change how you think about worker counts, timeouts, and graceful shutdown. Run an async server such as <code class=\"language-text\">uvicorn</code> so a streaming request does not pin a whole worker. And make sure your deploy drains in-flight streams instead of cutting them off mid-token.</p>\n<h2>What I would do differently</h2>\n<p>Early on, I let token formatting and model calls bleed into the same function, so swapping a model meant editing the streaming path. Splitting the provider adapter from the SSE layer, as the code above does, is worth doing from the first commit.</p>\n<p>I would also send a stable message id in the first event from the start. When a stream drops and the user retries, that id lets you reconcile the partial message instead of leaving a duplicate. It is far cheaper to add on day one than to backfill once the UI already assumes there is no id.</p>\n<p>If you want to see this pattern inside a full product, <a href=\"/project/cloud-canvas-ai/\">CloudCanvasAI</a> pairs a streaming FastAPI backend with a live document preview. <a href=\"/project/gemini-alchemy/\">Gemini Alchemy</a> streams structured output from a FastAPI service into an interactive UI. In both, the transport is the handful of lines above.</p>\n<hr>\n<p><em>Image credit: SSE streaming diagrams by M. Hassan Ahmed, created for this post, released under CC0 (public domain).</em></p>","frontmatter":{"title":"Streaming LLM Responses from FastAPI with SSE","date":"2026-06-30T00:00:00.000Z","description":"Stream LLM output token by token from a FastAPI backend with Server-Sent Events: working code, the EventSource client, proxy buffering, and failure modes.","thumbnail":{"childImageSharp":{"fluid":{"base64":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAABQAAAALCAIAAADwazoUAAAACXBIWXMAAAsSAAALEgHS3X78AAABk0lEQVQoz32P63KbMBCF+QnCuiJsg7gYHG4yJNgZ1zXGafwASZr2/d+lKxon7p/OfHNmdfastLJ4NUh91s+/q/NbcXqtn36W49vd8AJHKKrRmPr5V3l6bX68T4F3rxl5dRL1yUIknFFV1PuuPzftcfMwNt1Q6QMUuhvqzfdqc2h7Y97vnur2uK4eIQ9TLlGWK1IkEptGiCeAw+K/BWLJZwGmQ2OTYYlNlMsTg0gtRJWDA3MTi1wWT0RX4g/lMfcz4mfYA12DYrEC3yKhlvmjSHvXy+AiBI/A4zdAyCFREGuZaC9qRdhgvyB+YYZltq8PL2l3kastz3ue7+Z6YPlWNkdRfuPlXj5cvKRbZn2zvRTdyIIazwss17CUxVQrVMfje7wozE+8FQ1KurijyxL7uSuBDNaWSof5DpSolqoWNnepshwSOCScCODzDl4iEiAaouvRQAKXpwaRzrwVAAWY1kfbEHzW9pf5P2B4cWXuspDKFIuYeAn1UtAZj0CJSEChC5mb/MKyZ/MvJsu+ad86/yQn/gCzpkI6b86R9QAAAABJRU5ErkJggg==","aspectRatio":1.899441340782123,"src":"/static/aa9d918b5ab39f1c4481960a11c4f9a9/40a76/hero.png","srcSet":"/static/aa9d918b5ab39f1c4481960a11c4f9a9/c972b/hero.png 340w,\n/static/aa9d918b5ab39f1c4481960a11c4f9a9/27625/hero.png 680w,\n/static/aa9d918b5ab39f1c4481960a11c4f9a9/40a76/hero.png 1360w,\n/static/aa9d918b5ab39f1c4481960a11c4f9a9/ed396/hero.png 2000w","sizes":"(max-width: 1360px) 100vw, 1360px"}}}}}},"pageContext":{"slug":"/2026-06-30-fastapi-sse-streaming-llm/","previous":"blog/2026-06-30-llm-agent-tool-loop/","next":"blog/2026-06-29-incremental-rag-indexing/"}},"staticQueryHashes":["32046230"]}