{"componentChunkName":"component---src-templates-blog-post-js","path":"/blog/2026-07-13-streaming-markdown-react-llm/","result":{"data":{"site":{"siteMetadata":{"title":"M.Hassan Ahmed","author":"Hassan11196"}},"markdownRemark":{"id":"b5978d2e-ba20-59fe-b99c-c408050cd62e","excerpt":"When an LLM streams its answer into a chat UI, the browser receives the text a few tokens at a time, and the model is usually writing Markdown. Rendering a…","html":"<p>When an LLM streams its answer into a chat UI, the browser receives the text a few tokens at a time, and the model is usually writing Markdown. Rendering a finished Markdown document is a solved problem. Rendering a string that grows by a few characters every few milliseconds is not, because for most of the response that string is not valid Markdown yet:</p>\n<ul>\n<li>A code fence is open but not closed.</li>\n<li>A bold marker has one <code class=\"language-text\">*</code> but not its partner.</li>\n<li>A link has <code class=\"language-text\">[text](</code> and no URL.</li>\n</ul>\n<p>Render that naively and the message area flickers between raw syntax and formatted output. On a long answer, the whole thing also starts to lag.</p>\n<p>This post is for frontend engineers building an LLM chat UI in React. I hit every one of these problems building the split-panel chat in <a href=\"/project/cloud-canvas-ai/\">CloudCanvasAI</a>, where you watch a document take shape on the right as Claude writes it on the left, and again in the <a href=\"/project/archi/\">Archi</a> copilot at CERN. I will show why the naive version is both slow and ugly. Then I’ll walk through the two fixes that matter (parse less often, and parse smarter), how to stop half-written syntax from flickering, and why you should sanitize what the model writes.</p>\n<p>Most of my earlier posts covered the backend: how to <a href=\"/blog/2026-06-30-fastapi-sse-streaming-llm/\">stream tokens out of FastAPI over SSE</a> (Server-Sent Events), how the <a href=\"/blog/2026-06-30-llm-agent-tool-loop/\">agent tool loop</a> decides when to stop, and how one <a href=\"/blog/2026-07-10-fastapi-blocking-event-loop/\">blocking call stalls the event loop</a>. This post is the other end of that wire. The tokens have arrived in the browser, and now you have to render them.</p>\n<h2>The naive version, and why it breaks</h2>\n<p>You have an SSE stream (the <a href=\"/blog/2026-06-30-fastapi-sse-streaming-llm/\">backend post</a> shows where it comes from). Each message appends a token to a string, and you pass that string to <a href=\"https://github.com/remarkjs/react-markdown\">react-markdown</a>:</p>\n<div class=\"gatsby-highlight\" data-language=\"jsx\"><pre class=\"language-jsx\"><code class=\"language-jsx\"><span class=\"token keyword\">import</span> ReactMarkdown <span class=\"token keyword\">from</span> <span class=\"token string\">'react-markdown'</span><span class=\"token punctuation\">;</span>\n<span class=\"token keyword\">import</span> remarkGfm <span class=\"token keyword\">from</span> <span class=\"token string\">'remark-gfm'</span><span class=\"token punctuation\">;</span>\n\n<span class=\"token keyword\">function</span> <span class=\"token function\">Message</span><span class=\"token punctuation\">(</span><span class=\"token parameter\"><span class=\"token punctuation\">{</span> content <span class=\"token punctuation\">}</span></span><span class=\"token punctuation\">)</span> <span class=\"token punctuation\">{</span>\n  <span class=\"token keyword\">return</span> <span class=\"token punctuation\">(</span>\n    <span class=\"token tag\"><span class=\"token tag\"><span class=\"token punctuation\">&lt;</span><span class=\"token class-name\">ReactMarkdown</span></span> <span class=\"token attr-name\">remarkPlugins</span><span class=\"token script language-javascript\"><span class=\"token script-punctuation punctuation\">=</span><span class=\"token punctuation\">{</span><span class=\"token punctuation\">[</span>remarkGfm<span class=\"token punctuation\">]</span><span class=\"token punctuation\">}</span></span><span class=\"token punctuation\">></span></span><span class=\"token punctuation\">{</span>content<span class=\"token punctuation\">}</span><span class=\"token tag\"><span class=\"token tag\"><span class=\"token punctuation\">&lt;/</span><span class=\"token class-name\">ReactMarkdown</span></span><span class=\"token punctuation\">></span></span>\n  <span class=\"token punctuation\">)</span><span class=\"token punctuation\">;</span>\n<span class=\"token punctuation\">}</span></code></pre></div>\n<p><code class=\"language-text\">content</code> updates on every token. This works, and on a two-line answer you will never notice a problem. As the answer gets longer, two things go wrong.</p>\n<p>The first is <strong>cost</strong>. react-markdown does not diff: it never compares the new string with the old one to update only what changed. Every time <code class=\"language-text\">content</code> changes, <a href=\"https://github.com/remarkjs/remark\">remark</a> (the parser react-markdown uses) parses the <em>entire</em> accumulated string into an AST, or abstract syntax tree. react-markdown then rebuilds the element tree from scratch. Token 1 parses 4 characters. Token 500 parses the whole 2,000-character answer. Because you repeat that on every token, the total work over a full response grows with the square of its length. A 4,000-token answer is not twice the work of a 2,000-token one; it is four times. You feel it as the cursor stuttering near the end of a long reply.</p>\n<p>The second is <strong>flicker</strong>. A <a href=\"https://commonmark.org/\">CommonMark</a> parser (CommonMark is the standardized Markdown spec) is deterministic, but its output for a <em>prefix</em> of a document is not its output for the whole document. Mid-stream, the string <code class=\"language-text\">Here is the code: ```</code> is a paragraph containing two backticks, so react-markdown renders literal backticks. A few tokens later, the third backtick and a newline arrive. The same leading text is now the start of a fenced code block, and the DOM snaps to a grey code box. The user sees the text jump from one layout to another, and on a response full of code and lists it jumps the whole way down.</p>\n<p><img src=\"/fe1edd9d672dfaecd7e1556b2edc3928/render-pipeline.svg\" alt=\"A stream of tokens feeds a growing buffer. The naive path re-parses the entire buffer on every token, so cost climbs with the square of the length and the layout flickers as partial syntax resolves. The fixed path parses only the changed block and throttles updates, so cost stays flat and the rendered output is stable.\"></p>\n<p>Neither problem is react-markdown’s fault. It renders the string you gave it, as often as you changed it. Both fixes come down to asking it for less.</p>\n<h2>Fix one: parse less often by throttling updates</h2>\n<p>Tokens arrive faster than anyone can read, so there is no reason to re-render on every one. A person reads somewhere around 4 to 8 words a second; the model might emit 50 tokens a second. Batch the tokens and flush them on a timer instead of on arrival.</p>\n<p>Keep the live text in a ref, which holds a value without triggering a render when it changes, and copy it into state on an interval:</p>\n<div class=\"gatsby-highlight\" data-language=\"jsx\"><pre class=\"language-jsx\"><code class=\"language-jsx\"><span class=\"token keyword\">import</span> <span class=\"token punctuation\">{</span> useEffect<span class=\"token punctuation\">,</span> useRef<span class=\"token punctuation\">,</span> useState <span class=\"token punctuation\">}</span> <span class=\"token keyword\">from</span> <span class=\"token string\">'react'</span><span class=\"token punctuation\">;</span>\n\n<span class=\"token comment\">// `pending` accumulates every token; `text` is what we actually render.</span>\n<span class=\"token keyword\">function</span> <span class=\"token function\">useThrottledStream</span><span class=\"token punctuation\">(</span><span class=\"token parameter\">interval <span class=\"token operator\">=</span> <span class=\"token number\">60</span></span><span class=\"token punctuation\">)</span> <span class=\"token punctuation\">{</span>\n  <span class=\"token keyword\">const</span> pending <span class=\"token operator\">=</span> <span class=\"token function\">useRef</span><span class=\"token punctuation\">(</span><span class=\"token string\">''</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">;</span>\n  <span class=\"token keyword\">const</span> <span class=\"token punctuation\">[</span>text<span class=\"token punctuation\">,</span> setText<span class=\"token punctuation\">]</span> <span class=\"token operator\">=</span> <span class=\"token function\">useState</span><span class=\"token punctuation\">(</span><span class=\"token string\">''</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">;</span>\n\n  <span class=\"token keyword\">const</span> <span class=\"token function-variable function\">push</span> <span class=\"token operator\">=</span> <span class=\"token punctuation\">(</span><span class=\"token parameter\">chunk</span><span class=\"token punctuation\">)</span> <span class=\"token operator\">=></span> <span class=\"token punctuation\">{</span>\n    pending<span class=\"token punctuation\">.</span>current <span class=\"token operator\">+=</span> chunk<span class=\"token punctuation\">;</span>\n  <span class=\"token punctuation\">}</span><span class=\"token punctuation\">;</span>\n\n  <span class=\"token function\">useEffect</span><span class=\"token punctuation\">(</span><span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span> <span class=\"token operator\">=></span> <span class=\"token punctuation\">{</span>\n    <span class=\"token keyword\">const</span> id <span class=\"token operator\">=</span> <span class=\"token function\">setInterval</span><span class=\"token punctuation\">(</span><span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span> <span class=\"token operator\">=></span> <span class=\"token punctuation\">{</span>\n      <span class=\"token comment\">// Only setState when something changed, so we don't</span>\n      <span class=\"token comment\">// re-render on idle ticks after the stream ends.</span>\n      <span class=\"token function\">setText</span><span class=\"token punctuation\">(</span><span class=\"token punctuation\">(</span><span class=\"token parameter\">prev</span><span class=\"token punctuation\">)</span> <span class=\"token operator\">=></span> <span class=\"token punctuation\">(</span>prev <span class=\"token operator\">===</span> pending<span class=\"token punctuation\">.</span>current <span class=\"token operator\">?</span> prev <span class=\"token operator\">:</span> pending<span class=\"token punctuation\">.</span>current<span class=\"token punctuation\">)</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">;</span>\n    <span class=\"token punctuation\">}</span><span class=\"token punctuation\">,</span> interval<span class=\"token punctuation\">)</span><span class=\"token punctuation\">;</span>\n    <span class=\"token keyword\">return</span> <span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span> <span class=\"token operator\">=></span> <span class=\"token function\">clearInterval</span><span class=\"token punctuation\">(</span>id<span class=\"token punctuation\">)</span><span class=\"token punctuation\">;</span>\n  <span class=\"token punctuation\">}</span><span class=\"token punctuation\">,</span> <span class=\"token punctuation\">[</span>interval<span class=\"token punctuation\">]</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">;</span>\n\n  <span class=\"token keyword\">return</span> <span class=\"token punctuation\">{</span> text<span class=\"token punctuation\">,</span> push <span class=\"token punctuation\">}</span><span class=\"token punctuation\">;</span>\n<span class=\"token punctuation\">}</span></code></pre></div>\n<p>Notice that <code class=\"language-text\">push</code> only writes to the ref. The interval is the only thing that calls <code class=\"language-text\">setText</code>, so it alone decides when React renders.</p>\n<p>The SSE handler now calls <code class=\"language-text\">push(event.data)</code> on every token, but React re-renders only about 16 times a second instead of 50. That is one line of behavior change, and it removes most of the felt lag because it cuts the number of full re-parses by two-thirds or more. The <a href=\"https://developer.chrome.com/docs/ai/render-llm-responses\">Chrome team’s guidance on rendering streamed LLM responses</a> makes the same point: 16ms-per-token rendering is wasted work when a human reads far slower than the model writes.</p>\n<p>Sixty milliseconds is a starting point, not a law. Set it too high and the text arrives in visible chunks, which feels stuttery in a different way. Set it too low and you are back to paying for renders no one can see. I have landed between 50 and 100ms on every chat UI I have shipped. Tune it against a real long answer, not a one-liner.</p>\n<h2>Fix two: parse smarter by memoizing finished blocks</h2>\n<p>Throttling reduces how <em>often</em> you parse. It does nothing about the fact that each parse still chews through the whole document. To fix that, stop treating the answer as one string and split it into blocks.</p>\n<p>A Markdown document is a sequence of top-level blocks: paragraphs, headings, fenced code, lists. While the model streams, every block <em>above</em> the one it is currently writing is finished and will never change again. Only the last block is still growing. If you parse and memoize each block separately, React can skip the finished ones entirely.</p>\n<p>Use <a href=\"https://github.com/markedjs/marked\">marked</a> as a lexer to cut the string into block-level chunks, then render each chunk with a memoized component. <code class=\"language-text\">marked.lexer</code> returns one token per top-level block, and each token’s <code class=\"language-text\">raw</code> field holds that block’s original source text:</p>\n<div class=\"gatsby-highlight\" data-language=\"jsx\"><pre class=\"language-jsx\"><code class=\"language-jsx\"><span class=\"token keyword\">import</span> <span class=\"token punctuation\">{</span> memo<span class=\"token punctuation\">,</span> useMemo <span class=\"token punctuation\">}</span> <span class=\"token keyword\">from</span> <span class=\"token string\">'react'</span><span class=\"token punctuation\">;</span>\n<span class=\"token keyword\">import</span> <span class=\"token punctuation\">{</span> marked <span class=\"token punctuation\">}</span> <span class=\"token keyword\">from</span> <span class=\"token string\">'marked'</span><span class=\"token punctuation\">;</span>\n<span class=\"token keyword\">import</span> ReactMarkdown <span class=\"token keyword\">from</span> <span class=\"token string\">'react-markdown'</span><span class=\"token punctuation\">;</span>\n<span class=\"token keyword\">import</span> remarkGfm <span class=\"token keyword\">from</span> <span class=\"token string\">'remark-gfm'</span><span class=\"token punctuation\">;</span>\n\n<span class=\"token comment\">// Lexing is far cheaper than a full parse-to-AST-to-DOM.</span>\n<span class=\"token keyword\">function</span> <span class=\"token function\">splitIntoBlocks</span><span class=\"token punctuation\">(</span><span class=\"token parameter\">markdown</span><span class=\"token punctuation\">)</span> <span class=\"token punctuation\">{</span>\n  <span class=\"token keyword\">return</span> marked<span class=\"token punctuation\">.</span><span class=\"token function\">lexer</span><span class=\"token punctuation\">(</span>markdown<span class=\"token punctuation\">)</span><span class=\"token punctuation\">.</span><span class=\"token function\">map</span><span class=\"token punctuation\">(</span><span class=\"token punctuation\">(</span><span class=\"token parameter\">token</span><span class=\"token punctuation\">)</span> <span class=\"token operator\">=></span> token<span class=\"token punctuation\">.</span>raw<span class=\"token punctuation\">)</span><span class=\"token punctuation\">;</span>\n<span class=\"token punctuation\">}</span>\n\n<span class=\"token keyword\">const</span> Block <span class=\"token operator\">=</span> <span class=\"token function\">memo</span><span class=\"token punctuation\">(</span>\n  <span class=\"token punctuation\">(</span><span class=\"token parameter\"><span class=\"token punctuation\">{</span> content <span class=\"token punctuation\">}</span></span><span class=\"token punctuation\">)</span> <span class=\"token operator\">=></span> <span class=\"token punctuation\">(</span>\n    <span class=\"token tag\"><span class=\"token tag\"><span class=\"token punctuation\">&lt;</span><span class=\"token class-name\">ReactMarkdown</span></span> <span class=\"token attr-name\">remarkPlugins</span><span class=\"token script language-javascript\"><span class=\"token script-punctuation punctuation\">=</span><span class=\"token punctuation\">{</span><span class=\"token punctuation\">[</span>remarkGfm<span class=\"token punctuation\">]</span><span class=\"token punctuation\">}</span></span><span class=\"token punctuation\">></span></span><span class=\"token punctuation\">{</span>content<span class=\"token punctuation\">}</span><span class=\"token tag\"><span class=\"token tag\"><span class=\"token punctuation\">&lt;/</span><span class=\"token class-name\">ReactMarkdown</span></span><span class=\"token punctuation\">></span></span>\n  <span class=\"token punctuation\">)</span><span class=\"token punctuation\">,</span>\n  <span class=\"token punctuation\">(</span><span class=\"token parameter\">prev<span class=\"token punctuation\">,</span> next</span><span class=\"token punctuation\">)</span> <span class=\"token operator\">=></span> prev<span class=\"token punctuation\">.</span>content <span class=\"token operator\">===</span> next<span class=\"token punctuation\">.</span>content\n<span class=\"token punctuation\">)</span><span class=\"token punctuation\">;</span>\n\n<span class=\"token keyword\">function</span> <span class=\"token function\">StreamedMessage</span><span class=\"token punctuation\">(</span><span class=\"token parameter\"><span class=\"token punctuation\">{</span> content <span class=\"token punctuation\">}</span></span><span class=\"token punctuation\">)</span> <span class=\"token punctuation\">{</span>\n  <span class=\"token keyword\">const</span> blocks <span class=\"token operator\">=</span> <span class=\"token function\">useMemo</span><span class=\"token punctuation\">(</span><span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span> <span class=\"token operator\">=></span> <span class=\"token function\">splitIntoBlocks</span><span class=\"token punctuation\">(</span>content<span class=\"token punctuation\">)</span><span class=\"token punctuation\">,</span> <span class=\"token punctuation\">[</span>content<span class=\"token punctuation\">]</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">;</span>\n  <span class=\"token keyword\">return</span> blocks<span class=\"token punctuation\">.</span><span class=\"token function\">map</span><span class=\"token punctuation\">(</span><span class=\"token punctuation\">(</span><span class=\"token parameter\">block<span class=\"token punctuation\">,</span> i</span><span class=\"token punctuation\">)</span> <span class=\"token operator\">=></span> <span class=\"token tag\"><span class=\"token tag\"><span class=\"token punctuation\">&lt;</span><span class=\"token class-name\">Block</span></span> <span class=\"token attr-name\">key</span><span class=\"token script language-javascript\"><span class=\"token script-punctuation punctuation\">=</span><span class=\"token punctuation\">{</span>i<span class=\"token punctuation\">}</span></span> <span class=\"token attr-name\">content</span><span class=\"token script language-javascript\"><span class=\"token script-punctuation punctuation\">=</span><span class=\"token punctuation\">{</span>block<span class=\"token punctuation\">}</span></span> <span class=\"token punctuation\">/></span></span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">;</span>\n<span class=\"token punctuation\">}</span></code></pre></div>\n<p>The <code class=\"language-text\">memo</code> comparator is the whole point. When a new token lands, blocks <code class=\"language-text\">0</code> through <code class=\"language-text\">N-1</code> have the same <code class=\"language-text\">content</code> as on the last render, so <code class=\"language-text\">memo</code> returns the cached tree and does zero parsing. Only block <code class=\"language-text\">N</code>, the one being written, re-parses. The per-token cost stops growing with the length of the answer and stays flat at the size of one block. This is the pattern <a href=\"https://ai-sdk.dev/cookbook/next/markdown-chatbot-with-memoization\">Vercel’s AI SDK ships in its memoized-Markdown cookbook</a>, and it is the single biggest win for long responses.</p>\n<p>One caveat on <code class=\"language-text\">key</code>. The array index works as a key here only because blocks are append-only during a stream: block 3 stays block 3, and a new block 4 appears at the end. If you ever reorder or splice blocks, index keys make React reuse the wrong DOM, and you need a stable id per block instead.</p>\n<h2>Stopping the flicker: close what the model left open</h2>\n<p>Parsing smarter does not fix the code-fence flash. The problem is not how often you parse; a half-written fence genuinely is not a fence yet. The cleanest fix is to lie to the parser: before rendering, find syntax the model has opened but not closed, and close it temporarily.</p>\n<p>An open code fence is the case that matters most, because the miss is so visible: raw source text with backticks instead of a code box. Count the fences, and if the count is odd, append a closing fence for this render only:</p>\n<div class=\"gatsby-highlight\" data-language=\"js\"><pre class=\"language-js\"><code class=\"language-js\"><span class=\"token keyword\">function</span> <span class=\"token function\">closeOpenFence</span><span class=\"token punctuation\">(</span><span class=\"token parameter\">markdown</span><span class=\"token punctuation\">)</span> <span class=\"token punctuation\">{</span>\n  <span class=\"token keyword\">const</span> fenceCount <span class=\"token operator\">=</span> <span class=\"token punctuation\">(</span>markdown<span class=\"token punctuation\">.</span><span class=\"token function\">match</span><span class=\"token punctuation\">(</span><span class=\"token regex\">/```/g</span><span class=\"token punctuation\">)</span> <span class=\"token operator\">||</span> <span class=\"token punctuation\">[</span><span class=\"token punctuation\">]</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">.</span>length<span class=\"token punctuation\">;</span>\n  <span class=\"token comment\">// Odd number of ``` means a fence is still open.</span>\n  <span class=\"token keyword\">return</span> fenceCount <span class=\"token operator\">%</span> <span class=\"token number\">2</span> <span class=\"token operator\">===</span> <span class=\"token number\">1</span> <span class=\"token operator\">?</span> <span class=\"token template-string\"><span class=\"token template-punctuation string\">`</span><span class=\"token interpolation\"><span class=\"token interpolation-punctuation punctuation\">${</span>markdown<span class=\"token interpolation-punctuation punctuation\">}</span></span><span class=\"token string\">\\n\\`\\`\\`</span><span class=\"token template-punctuation string\">`</span></span> <span class=\"token operator\">:</span> markdown<span class=\"token punctuation\">;</span>\n<span class=\"token punctuation\">}</span></code></pre></div>\n<p>Run the raw buffer through <code class=\"language-text\">closeOpenFence</code> before you split it into blocks. As soon as the model writes the opening fence, the content renders as a code block and <em>stays</em> a code block as the code fills in, instead of flashing from raw backticks to a box. When the real closing fence arrives, the count is even, so your temporary fence drops out on the next parse. The same trick works for unclosed bold and inline code, but the code fence is the one worth doing first.</p>\n<p><img src=\"/a031c590da6b5bc97bc8b3f883d41929/incomplete-markdown.svg\" alt=\"Mid-stream, an odd number of triple-backticks means the model has opened a code block it has not yet closed. Rendered as-is, the parser shows literal backticks and raw source. After appending a temporary closing fence, the same text renders as a stable code box that fills in as tokens arrive.\"></p>\n<p>If you would rather not maintain this yourself, purpose-built renderers now handle it. <a href=\"https://github.com/vercel/streamdown\">Streamdown</a>, a drop-in react-markdown replacement from the Vercel team, renders an open code block as soon as the model starts one and is built for exactly these partial-syntax states. Once your own preprocessing grows past the code fence, switching to it is a reasonable call.</p>\n<h2>Do not render model output as trusted HTML</h2>\n<p>By default, react-markdown does not render raw HTML. That is the safe default, and one reason to prefer it over hand-rolling with <code class=\"language-text\">dangerouslySetInnerHTML</code>. The moment you add <code class=\"language-text\">rehype-raw</code> to support inline HTML in answers, model output becomes a script-injection path. A retrieved document or a prompt-injected instruction can carry an <code class=\"language-text\">&lt;img onerror=...&gt;</code> or a <code class=\"language-text\">javascript:</code> link straight into your DOM. If you allow HTML, sanitize it:</p>\n<div class=\"gatsby-highlight\" data-language=\"jsx\"><pre class=\"language-jsx\"><code class=\"language-jsx\"><span class=\"token keyword\">import</span> rehypeSanitize <span class=\"token keyword\">from</span> <span class=\"token string\">'rehype-sanitize'</span><span class=\"token punctuation\">;</span>\n\n<span class=\"token tag\"><span class=\"token tag\"><span class=\"token punctuation\">&lt;</span><span class=\"token class-name\">ReactMarkdown</span></span>\n  <span class=\"token attr-name\">remarkPlugins</span><span class=\"token script language-javascript\"><span class=\"token script-punctuation punctuation\">=</span><span class=\"token punctuation\">{</span><span class=\"token punctuation\">[</span>remarkGfm<span class=\"token punctuation\">]</span><span class=\"token punctuation\">}</span></span>\n  <span class=\"token attr-name\">rehypePlugins</span><span class=\"token script language-javascript\"><span class=\"token script-punctuation punctuation\">=</span><span class=\"token punctuation\">{</span><span class=\"token punctuation\">[</span>rehypeSanitize<span class=\"token punctuation\">]</span><span class=\"token punctuation\">}</span></span>\n<span class=\"token punctuation\">></span></span><span class=\"token plain-text\">\n  </span><span class=\"token punctuation\">{</span>content<span class=\"token punctuation\">}</span><span class=\"token plain-text\">\n</span><span class=\"token tag\"><span class=\"token tag\"><span class=\"token punctuation\">&lt;/</span><span class=\"token class-name\">ReactMarkdown</span></span><span class=\"token punctuation\">></span></span><span class=\"token punctuation\">;</span></code></pre></div>\n<p><a href=\"https://github.com/rehypejs/rehype-sanitize\">rehype-sanitize</a> runs inside the same AST, so you are not shipping a second parser. It strips event handlers and dangerous URLs while keeping the formatting. This matters more with RAG (retrieval-augmented generation) in the loop, because the text you are rendering was partly written by documents you crawled, not just by the model.</p>\n<h2>Failure modes</h2>\n<p><strong>Auto-scroll fighting the user.</strong> Scrolling to the bottom on every update keeps the newest text visible, until the user scrolls up to re-read something and your code yanks them back down on the next token. Check whether they are near the bottom before you scroll, and auto-scroll only if they already were. Missing this small check is one of the most common broken-chat-UI complaints.</p>\n<p><strong>Throttling that outlives the stream.</strong> The interval in <code class=\"language-text\">useThrottledStream</code> runs forever unless you clear it. The guard <code class=\"language-text\">prev === pending.current</code> keeps it from re-rendering on idle ticks, but still stop the timer when the stream ends, so you are not spinning a <code class=\"language-text\">setInterval</code> for every finished message on screen.</p>\n<p><strong>Index keys after the stream.</strong> Append-only index keys are fine while streaming. If the same component later supports editing or deleting a block, switch to stable ids first, or React will paste new content into an old block’s DOM node.</p>\n<p><strong>Trusting the fence count with inline backticks.</strong> Counting <code class=\"language-text\">```</code> is a heuristic, not a parser, and a message that legitimately discusses triple-backticks in prose can throw the count off. It is right the vast majority of the time, and when it is wrong the damage is cosmetic.</p>\n<p><strong>Testing only on short answers.</strong> Every one of these problems is invisible on a two-line reply and obvious on a two-page one. Test against a long, code-heavy response streamed at real speed, or you will ship the naive version and hear about it from users on the answers that matter most.</p>\n<h2>Tradeoffs: what to add first</h2>\n<p>The throttle and the block memoization are close to free, and I would put them in from the first commit: the value is real and the cost is a few lines. The fence-closing preprocessing needs judgment. It is a heuristic that accepts a small chance of a cosmetic miss in exchange for killing the most jarring flicker. That trade is almost always worth it, but it is a trade. Once you find yourself special-casing bold, then inline code, then tables, then math, you have started writing a streaming Markdown renderer. That is the point to adopt one, like Streamdown or the <a href=\"https://ai-sdk.dev/cookbook/next/markdown-chatbot-with-memoization\">Vercel AI SDK</a>, instead of maintaining your own.</p>\n<h2>Putting it together</h2>\n<p>The rendering side of an LLM chat UI looks trivial until the stream is live. Then it becomes a small pile of specific problems: quadratic re-parsing, flicker because a prefix is not a document, an auto-scroll that fights the reader, and model text you should not trust as HTML. None of them needs a framework:</p>\n<ol>\n<li>Throttle the updates, so you parse a few times a second instead of fifty.</li>\n<li>Split the answer into blocks and memoize them, so a new token re-parses only the block being written.</li>\n<li>Close the fences the model has not closed yet.</li>\n<li>Sanitize what you render.</li>\n</ol>\n<p>That is the same chat surface behind <a href=\"/project/cloud-canvas-ai/\">CloudCanvasAI</a> and the <a href=\"/project/archi/\">Archi copilot</a>. The whole point there is watching the answer build in real time, so the rendering has to be as clean as the <a href=\"/blog/2026-06-30-fastapi-sse-streaming-llm/\">stream feeding it</a>.</p>\n<hr>\n<p><em>Diagrams by M. Hassan Ahmed, released under CC0. Image credit: original work by the author.</em></p>","frontmatter":{"title":"Streaming LLM Markdown in React Without Flicker","date":"2026-07-13T00:00:00.000Z","description":"LLM tokens arrive one at a time, and re-parsing Markdown on every token flickers and drags. How to render streaming Markdown in React without the jank.","thumbnail":{"childImageSharp":{"fluid":{"base64":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAABQAAAALCAIAAADwazoUAAAACXBIWXMAAAsSAAALEgHS3X78AAAB4UlEQVQoz02R6W7iMBRG8zN4N3EScByHbA4hLAGGpYWyqHSqvv8L9TKVRpWOPl1dfedKTjzi1rRZE9fjejV5eS/Pf/PXR3H6qN4+i9PD7K/15RM2ze3LHu6oWlLXU7d+Wm7tYWYQ0VxmQezq6X6+epstTt3yDNnOX2HzM/w5fMBSxU4EOaIJWICHh5aogutWxA0OChJWNKwgiSqB56wqPMx8bpC0LMxJMAHlH5mHecJGTdKdTbWLJ/1Qd3I8+80wmfOoJjLFwiBhsEj/4w1YwlRl3aHoznn7GiTdUM/gxA9PWXcidlCFpk+1zzTi5nmIG5A1C6s43yf1OalPdnq17dW407jc6/IY2l7qGQ9rGkGnz9ze1jtlFqAgnngDOoZJmS6yczlyNCyfqJyonKoCko8aHlZCt5PFZbp5NOt3257EuEV07Pk0BlmOpirpAt3C83hUPTOuQYYr8CGhQKTFIkM8xdLS4QQLO2Agk4ipUuWbqDrE9XHkjuPmJSy3waQXpgNkuoACkgmJMxJlvtTAAKAjb4BDaqbp9T273O3lri831m/ZagPwfsv7jdgciJ2x5Sr/+rCPm9ptwpcdnc99Ens+DpE0zDYsdQBNHTH1b6hx8IdQbHnR8MLJshVVi3QO4jdXbk0MRLGiNgAAAABJRU5ErkJggg==","aspectRatio":1.899441340782123,"src":"/static/7fde826b4ceb97a0223355de964fb9fe/40a76/hero.png","srcSet":"/static/7fde826b4ceb97a0223355de964fb9fe/c972b/hero.png 340w,\n/static/7fde826b4ceb97a0223355de964fb9fe/27625/hero.png 680w,\n/static/7fde826b4ceb97a0223355de964fb9fe/40a76/hero.png 1360w,\n/static/7fde826b4ceb97a0223355de964fb9fe/ed396/hero.png 2000w","sizes":"(max-width: 1360px) 100vw, 1360px"}}}}}},"pageContext":{"slug":"/2026-07-13-streaming-markdown-react-llm/","previous":"blog/2026-07-14-reliable-json-from-llms/","next":"blog/2026-07-17-build-mcp-server-llm-agent/"}},"staticQueryHashes":["32046230"]}