{"componentChunkName":"component---src-templates-blog-post-js","path":"/blog/2026-07-23-query-rewriting-rag-retrieval/","result":{"data":{"site":{"siteMetadata":{"title":"M.Hassan Ahmed","author":"Hassan11196"}},"markdownRemark":{"id":"fa22fec1-7895-5c1c-89f7-f1169b41089f","excerpt":"An operator on shift types “did it break again?” into the copilot. The retriever embeds that four-word string, searches the index, and hands back the nearest…","html":"<p>An operator on shift types “did it break again?” into the copilot. The retriever embeds that four-word string, searches the index, and hands back the nearest chunks, which could be about anything, because nothing in the query says what “it” is. A minute earlier the same person had asked about a <code class=\"language-text\">T1_US_FNAL</code> transfer error, so a human would know exactly what they meant. The index has no memory of that turn. It only saw four words with a pronoun in them.</p>\n<p>Most RAG (retrieval-augmented generation) debugging skips past this failure. When retrieval returns junk, the reflex is to blame the index: bad chunks, the wrong embedding model, a missing reranker. Sometimes that is right. But often the chunk you need is sitting in the index, embedded perfectly well. The problem is that the query you searched with looks nothing like the document that answers it. You cannot fix that downstream. You fix it by rewriting the query before it reaches the retriever.</p>\n<p>This post is for engineers who already have a working RAG pipeline and keep watching it miss on short questions, vague questions, and follow-ups. I will cover why the raw query is so often the weak link, then three techniques that sit in front of retrieval:</p>\n<ol>\n<li>contextualizing follow-ups against the chat history;</li>\n<li>multi-query expansion;</li>\n<li><a href=\"https://arxiv.org/abs/2212.10496\">HyDE</a> (Hypothetical Document Embeddings).</li>\n</ol>\n<p>Each one buys you something different, and each can quietly make retrieval worse if you use it in the wrong place.</p>\n<h2>Why the raw query is the weak link</h2>\n<p>Dense retrieval compares a query vector against document vectors, so it only works when the two land near each other in embedding space. The trouble is that a question and its answer are written in different registers (styles of language). The question is short, informal, and often assumes context the asker already has. The document is a full paragraph of prose that never repeats the question back.</p>\n<p>Take “why is FNAL slow” and a logbook entry titled “FTS staging backlog on the US Tier-1 causing transfer throughput degradation.” They are about the same thing. Their embeddings can still sit far apart, because almost none of the words overlap and the two texts have different shapes.</p>\n<p>That gap shows up in three recurring forms.</p>\n<p><strong>Too short to be specific.</strong> A three-word query carries almost no signal. There is not enough text for the embedding to land anywhere precise, so it lands in a vague neighborhood and pulls back a vague set of chunks.</p>\n<p><strong>Vocabulary mismatch.</strong> The user says “slow”; the document says “throughput degradation.” The user says “FNAL”; the document says <code class=\"language-text\">T1_US_FNAL</code>. Dense retrieval is supposed to bridge synonyms better than keyword search does, and it partly does. That is exactly why the cases it misses are easy to overlook. <a href=\"/blog/2026-07-02-hybrid-search-rag-bm25-vectors/\">Hybrid search</a> recovers some of these by adding exact-match BM25 keyword scoring back into the mix, but it cannot invent a term the user never typed.</p>\n<p><strong>Conversational follow-ups.</strong> In a chat interface, half the questions are not standalone: “did it break again?”, “what about the other site?”, “why?” They make sense only against the previous turns. Embed them literally and you retrieve for the wrong thing entirely.</p>\n<p>The fix for all three is a transformation stage between the user and the retriever. The index stays exactly as it is; you change what you ask it.</p>\n<p><img src=\"/858ec6ff87bdefd106d04a2e00245dc7/query-transform-pipeline.svg\" alt=\"A RAG pipeline with a query transformation stage inserted before retrieval. The raw query passes through a transform box offering three approaches, contextualize, multi-query, and HyDE, before reaching the retriever, index, and LLM. The index never changes; only the query does.\"></p>\n<h2>Contextualizing follow-up questions</h2>\n<p>If you run a chat copilot, build this one first. It fixes the most common failure for the least effort.</p>\n<p>The idea is to use the conversation history to rewrite a context-dependent question into a standalone one before retrieval. “did it break again?” becomes “Did the <code class=\"language-text\">T1_US_FNAL</code> transfer error recur after the fix?” The rewritten query stands on its own, so the retriever now has something to work with.</p>\n<p>It takes one LLM call with the history and the latest question. The code below builds that prompt, skips the call when there is no history, and retrieves with the rewritten query:</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\">CONDENSE_PROMPT <span class=\"token operator\">=</span> <span class=\"token triple-quoted-string string\">\"\"\"Given the conversation and a follow-up question,\nrewrite the follow-up as a standalone question that includes the\nentities and context it depends on. Do not answer it. If it is already\nstandalone, return it unchanged.\n\nConversation:\n{history}\n\nFollow-up: {question}\nStandalone question:\"\"\"</span>\n\n<span class=\"token keyword\">def</span> <span class=\"token function\">condense</span><span class=\"token punctuation\">(</span>history<span class=\"token punctuation\">:</span> <span class=\"token builtin\">list</span><span class=\"token punctuation\">[</span><span class=\"token builtin\">dict</span><span class=\"token punctuation\">]</span><span class=\"token punctuation\">,</span> question<span class=\"token punctuation\">:</span> <span class=\"token builtin\">str</span><span class=\"token punctuation\">)</span> <span class=\"token operator\">-</span><span class=\"token operator\">></span> <span class=\"token builtin\">str</span><span class=\"token punctuation\">:</span>\n    <span class=\"token keyword\">if</span> <span class=\"token keyword\">not</span> history<span class=\"token punctuation\">:</span>\n        <span class=\"token keyword\">return</span> question  <span class=\"token comment\"># nothing to fold in, skip the call</span>\n    convo <span class=\"token operator\">=</span> <span class=\"token string\">\"\\n\"</span><span class=\"token punctuation\">.</span>join<span class=\"token punctuation\">(</span><span class=\"token string-interpolation\"><span class=\"token string\">f\"</span><span class=\"token interpolation\"><span class=\"token punctuation\">{</span>m<span class=\"token punctuation\">[</span><span class=\"token string\">'role'</span><span class=\"token punctuation\">]</span><span class=\"token punctuation\">}</span></span><span class=\"token string\">: </span><span class=\"token interpolation\"><span class=\"token punctuation\">{</span>m<span class=\"token punctuation\">[</span><span class=\"token string\">'content'</span><span class=\"token punctuation\">]</span><span class=\"token punctuation\">}</span></span><span class=\"token string\">\"</span></span> <span class=\"token keyword\">for</span> m <span class=\"token keyword\">in</span> history<span class=\"token punctuation\">[</span><span class=\"token operator\">-</span><span class=\"token number\">6</span><span class=\"token punctuation\">:</span><span class=\"token punctuation\">]</span><span class=\"token punctuation\">)</span>\n    rewritten <span class=\"token operator\">=</span> llm<span class=\"token punctuation\">.</span>generate<span class=\"token punctuation\">(</span>\n        CONDENSE_PROMPT<span class=\"token punctuation\">.</span><span class=\"token builtin\">format</span><span class=\"token punctuation\">(</span>history<span class=\"token operator\">=</span>convo<span class=\"token punctuation\">,</span> question<span class=\"token operator\">=</span>question<span class=\"token punctuation\">)</span>\n    <span class=\"token punctuation\">)</span><span class=\"token punctuation\">.</span>strip<span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span>\n    <span class=\"token keyword\">return</span> rewritten\n\n<span class=\"token comment\"># then retrieve with the rewritten query, not the raw one</span>\nstandalone <span class=\"token operator\">=</span> condense<span class=\"token punctuation\">(</span>history<span class=\"token punctuation\">,</span> user_question<span class=\"token punctuation\">)</span>\nchunks <span class=\"token operator\">=</span> retriever<span class=\"token punctuation\">.</span>search<span class=\"token punctuation\">(</span>standalone<span class=\"token punctuation\">,</span> k<span class=\"token operator\">=</span><span class=\"token number\">8</span><span class=\"token punctuation\">)</span></code></pre></div>\n<p>Two details matter:</p>\n<ul>\n<li><strong>Only the last handful of turns go into the prompt.</strong> A full transcript costs more, and it drags in stale context that pulls the rewrite off target.</li>\n<li><strong>The early return saves a call.</strong> When there is no history, which is the first question of every conversation and a meaningful fraction of traffic, no LLM call is made.</li>\n</ul>\n<p>This is the “condense question” pattern that <a href=\"https://python.langchain.com/docs/how_to/qa_chat_history_how_to/\">LangChain’s history-aware retriever</a> formalizes. You can write it in a dozen lines without the framework.</p>\n<h2>Multi-query expansion</h2>\n<p>Contextualizing fixes follow-ups. It does nothing for a standalone question that is simply too vague or unluckily phrased. For that, generate several rewordings and search with all of them.</p>\n<p>The reasoning: any single phrasing is one sample from the many ways to ask the question, and it might be a bad sample. Ask an LLM for three or four alternate phrasings, retrieve for each, and merge the results. Where one phrasing misses the right chunk, another catches it, so recall goes up. <a href=\"https://python.langchain.com/docs/how_to/MultiQueryRetriever/\">LangChain’s <code class=\"language-text\">MultiQueryRetriever</code></a> packages exactly this, but the mechanics are worth seeing directly. The function below searches with the original question plus the generated variants, then fuses the ranked lists:</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\">EXPAND_PROMPT <span class=\"token operator\">=</span> <span class=\"token triple-quoted-string string\">\"\"\"Generate 3 alternative phrasings of this question,\neach from a different angle, one per line, no numbering:\n\n{question}\"\"\"</span>\n\n<span class=\"token keyword\">def</span> <span class=\"token function\">multi_query</span><span class=\"token punctuation\">(</span>question<span class=\"token punctuation\">:</span> <span class=\"token builtin\">str</span><span class=\"token punctuation\">,</span> k<span class=\"token punctuation\">:</span> <span class=\"token builtin\">int</span> <span class=\"token operator\">=</span> <span class=\"token number\">8</span><span class=\"token punctuation\">)</span> <span class=\"token operator\">-</span><span class=\"token operator\">></span> <span class=\"token builtin\">list</span><span class=\"token punctuation\">[</span><span class=\"token builtin\">dict</span><span class=\"token punctuation\">]</span><span class=\"token punctuation\">:</span>\n    variants <span class=\"token operator\">=</span> llm<span class=\"token punctuation\">.</span>generate<span class=\"token punctuation\">(</span>EXPAND_PROMPT<span class=\"token punctuation\">.</span><span class=\"token builtin\">format</span><span class=\"token punctuation\">(</span>question<span class=\"token operator\">=</span>question<span class=\"token punctuation\">)</span><span class=\"token punctuation\">)</span>\n    queries <span class=\"token operator\">=</span> <span class=\"token punctuation\">[</span>question<span class=\"token punctuation\">]</span> <span class=\"token operator\">+</span> <span class=\"token punctuation\">[</span>q<span class=\"token punctuation\">.</span>strip<span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span> <span class=\"token keyword\">for</span> q <span class=\"token keyword\">in</span> variants<span class=\"token punctuation\">.</span>splitlines<span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span> <span class=\"token keyword\">if</span> q<span class=\"token punctuation\">.</span>strip<span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">]</span>\n    ranked_lists <span class=\"token operator\">=</span> <span class=\"token punctuation\">[</span>retriever<span class=\"token punctuation\">.</span>search<span class=\"token punctuation\">(</span>q<span class=\"token punctuation\">,</span> k<span class=\"token operator\">=</span>k<span class=\"token punctuation\">)</span> <span class=\"token keyword\">for</span> q <span class=\"token keyword\">in</span> queries<span class=\"token punctuation\">]</span>\n    <span class=\"token keyword\">return</span> reciprocal_rank_fusion<span class=\"token punctuation\">(</span>ranked_lists<span class=\"token punctuation\">)</span></code></pre></div>\n<h3>Merging the results with Reciprocal Rank Fusion</h3>\n<p>The part people get wrong is the merge. You now have several ranked lists and need one. Do not just concatenate and dedupe, because that throws away the rank information.</p>\n<p>Use <a href=\"https://plg.uwaterloo.ca/~gvcormack/cormacksigir09-rrf.pdf\">Reciprocal Rank Fusion</a> (RRF) instead, the same fusion I used to combine keyword and vector results in the <a href=\"/blog/2026-07-02-hybrid-search-rag-bm25-vectors/\">hybrid search post</a>. RRF scores each document by summing <code class=\"language-text\">1 / (k + rank)</code> across every list it appears in. A chunk that shows up near the top of two different rewrites therefore beats one that ranks first in a single list and nowhere else:</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\"><span class=\"token keyword\">def</span> <span class=\"token function\">reciprocal_rank_fusion</span><span class=\"token punctuation\">(</span>ranked_lists<span class=\"token punctuation\">,</span> k<span class=\"token punctuation\">:</span> <span class=\"token builtin\">int</span> <span class=\"token operator\">=</span> <span class=\"token number\">60</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n    scores<span class=\"token punctuation\">,</span> seen <span class=\"token operator\">=</span> <span class=\"token punctuation\">{</span><span class=\"token punctuation\">}</span><span class=\"token punctuation\">,</span> <span class=\"token punctuation\">{</span><span class=\"token punctuation\">}</span>\n    <span class=\"token keyword\">for</span> lst <span class=\"token keyword\">in</span> ranked_lists<span class=\"token punctuation\">:</span>\n        <span class=\"token keyword\">for</span> rank<span class=\"token punctuation\">,</span> doc <span class=\"token keyword\">in</span> <span class=\"token builtin\">enumerate</span><span class=\"token punctuation\">(</span>lst<span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n            scores<span class=\"token punctuation\">[</span>doc<span class=\"token punctuation\">[</span><span class=\"token string\">\"id\"</span><span class=\"token punctuation\">]</span><span class=\"token punctuation\">]</span> <span class=\"token operator\">=</span> scores<span class=\"token punctuation\">.</span>get<span class=\"token punctuation\">(</span>doc<span class=\"token punctuation\">[</span><span class=\"token string\">\"id\"</span><span class=\"token punctuation\">]</span><span class=\"token punctuation\">,</span> <span class=\"token number\">0</span><span class=\"token punctuation\">)</span> <span class=\"token operator\">+</span> <span class=\"token number\">1</span> <span class=\"token operator\">/</span> <span class=\"token punctuation\">(</span>k <span class=\"token operator\">+</span> rank<span class=\"token punctuation\">)</span>\n            seen<span class=\"token punctuation\">[</span>doc<span class=\"token punctuation\">[</span><span class=\"token string\">\"id\"</span><span class=\"token punctuation\">]</span><span class=\"token punctuation\">]</span> <span class=\"token operator\">=</span> doc\n    ordered <span class=\"token operator\">=</span> <span class=\"token builtin\">sorted</span><span class=\"token punctuation\">(</span>scores<span class=\"token punctuation\">,</span> key<span class=\"token operator\">=</span>scores<span class=\"token punctuation\">.</span>get<span class=\"token punctuation\">,</span> reverse<span class=\"token operator\">=</span><span class=\"token boolean\">True</span><span class=\"token punctuation\">)</span>\n    <span class=\"token keyword\">return</span> <span class=\"token punctuation\">[</span>seen<span class=\"token punctuation\">[</span>i<span class=\"token punctuation\">]</span> <span class=\"token keyword\">for</span> i <span class=\"token keyword\">in</span> ordered<span class=\"token punctuation\">]</span></code></pre></div>\n<p>The cost is real. Three rewrites mean three retrieval round trips, plus the generation that produced them. On a vector index that is usually cheap, but it is not free, and it adds up if you also rerank afterward.</p>\n<p><img src=\"/e4413882c8dedaa09d17355d79d949ca/three-rewrites.svg\" alt=\"Three query-rewriting techniques compared side by side. Contextualize folds chat history into one standalone query. Multi-query generates three paraphrases, retrieves each, and fuses them with RRF. HyDE drafts a hypothetical answer and embeds that instead of the question. All three trade one extra generation for a better match without changing the index.\"></p>\n<h2>HyDE: search with a hypothetical answer</h2>\n<p>Multi-query still searches with questions. HyDE, from <a href=\"https://arxiv.org/abs/2212.10496\">Gao et al. at ACL 2023</a>, takes a stranger route. It asks the LLM to write a fake answer to the question, then embeds that fake answer and searches with it.</p>\n<p>The reasoning follows directly from the register problem. Your index is full of answer-shaped documents, so an answer-shaped query lands closer to the right neighborhood than the terse question ever could.</p>\n<p>It does not matter if the drafted answer is factually wrong. You never show it to the user; you only use its embedding as a better search key. A plausible-but-wrong paragraph about FTS staging backlogs still sits near the real logbook entries about FTS staging backlogs. In code, the only change from normal retrieval is what gets embedded:</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\">HYDE_PROMPT <span class=\"token operator\">=</span> <span class=\"token triple-quoted-string string\">\"\"\"Write a short, plausible paragraph that could answer\nthis question. It does not need to be correct, just realistic in style\nand terminology:\n\n{question}\"\"\"</span>\n\n<span class=\"token keyword\">def</span> <span class=\"token function\">hyde_search</span><span class=\"token punctuation\">(</span>question<span class=\"token punctuation\">:</span> <span class=\"token builtin\">str</span><span class=\"token punctuation\">,</span> k<span class=\"token punctuation\">:</span> <span class=\"token builtin\">int</span> <span class=\"token operator\">=</span> <span class=\"token number\">8</span><span class=\"token punctuation\">)</span> <span class=\"token operator\">-</span><span class=\"token operator\">></span> <span class=\"token builtin\">list</span><span class=\"token punctuation\">[</span><span class=\"token builtin\">dict</span><span class=\"token punctuation\">]</span><span class=\"token punctuation\">:</span>\n    draft <span class=\"token operator\">=</span> llm<span class=\"token punctuation\">.</span>generate<span class=\"token punctuation\">(</span>HYDE_PROMPT<span class=\"token punctuation\">.</span><span class=\"token builtin\">format</span><span class=\"token punctuation\">(</span>question<span class=\"token operator\">=</span>question<span class=\"token punctuation\">)</span><span class=\"token punctuation\">)</span>\n    <span class=\"token keyword\">return</span> retriever<span class=\"token punctuation\">.</span>search_by_text<span class=\"token punctuation\">(</span>draft<span class=\"token punctuation\">,</span> k<span class=\"token operator\">=</span>k<span class=\"token punctuation\">)</span>  <span class=\"token comment\"># embeds the draft, not the query</span></code></pre></div>\n<p>HyDE earns its keep when the query and the documents are written in genuinely different registers: terse operator shorthand against formal write-ups, or a domain where the user does not know the right vocabulary. In the paper, it was strong precisely where there were no training labels to fine-tune a retriever, which describes most internal corpora. Still, it is the technique I reach for last, because it is the most likely to backfire, as the next section covers.</p>\n<h2>Tradeoffs and failure modes</h2>\n<p>Every technique here spends an extra LLM generation to improve retrieval, and each has a specific way of going wrong.</p>\n<p><strong>The latency is not free, and it lands early.</strong> The rewrite happens before retrieval, which happens before generation. You have added a full model call to the front of the request, and the user sees nothing until it finishes. If you <a href=\"/blog/2026-06-30-fastapi-sse-streaming-llm/\">stream the answer over SSE</a>, the rewrite is pure dead time before the first token, so measure it as part of time-to-first-token, not as a background cost. Use a small, fast model for the rewrite; it does not need the model that writes the final answer.</p>\n<p><strong>Rewrites drift.</strong> A condense or expansion step can hallucinate an entity that was never in the conversation, or quietly change the question’s meaning. If “why is FNAL slow” gets expanded into “why is the CERN Tier-0 slow,” you now retrieve confidently for the wrong site. Log the rewritten query next to the original whenever a request looks off. The bug is often obvious the moment you can see what you actually searched for.</p>\n<p><strong>HyDE amplifies wrong assumptions.</strong> This failure mode mirrors its strength. If the drafted answer is wrong in the wrong direction, its embedding points at the wrong neighborhood, and you retrieve confidently irrelevant chunks. On a question with a specific factual answer that the user already phrased well, HyDE can lose to plain search with the original query. It helps most on vague questions and hurts most on precise ones.</p>\n<p><strong>Over-expansion dilutes results.</strong> More rewrites are not always better. Past three or four, the extra phrasings start repeating each other. Each additional list you fuse in also gives off-target chunks another chance to accumulate RRF score. Recall stops climbing, and precision slips.</p>\n<p><strong>It breaks naive caching.</strong> If you cache retrieval results keyed on the query string, a rewrite step changes the key on every request. A non-deterministic rewrite can even produce two different keys for the same question. Cache on the rewritten query, or downstream of it, not on the raw input.</p>\n<p>None of this is visible unless you measure. Build a small golden set of real queries with their correct chunks, and score recall with and without each technique, exactly as in the <a href=\"/blog/2026-07-04-measuring-rag-retrieval-quality/\">retrieval-quality post</a>. Query rewriting is a change to your retriever, and it deserves the same measurement as any other retriever change.</p>\n<h2>What I would do differently: add one technique at a time</h2>\n<p>If I were adding this to a pipeline from scratch, I would resist turning on all three at once. The order I would follow:</p>\n<ol>\n<li><strong>Start with contextualization.</strong> In a chat product it fixes the single most common miss, the unresolved follow-up, and it is cheap and hard to get wrong. Ship it and measure it before doing anything else.</li>\n<li><strong>Add multi-query only if standalone queries still miss.</strong> It is the safer of the two remaining options, because RRF is forgiving of a bad rewrite: one off-target list gets outvoted.</li>\n<li><strong>Use HyDE last, and only with evidence.</strong> Reach for it only when you have a measured recall gap on vague queries with little vocabulary overlap that multi-query did not close. A/B test it against plain retrieval before you trust it, because it is the one most likely to quietly make some queries worse while helping others.</li>\n</ol>\n<p>The order matters because these techniques stack multiplicatively in cost and additively in ways to fail. One well-placed rewrite beats three stacked ones you cannot debug.</p>\n<h2>Closing</h2>\n<p>Query rewriting decides what your retriever actually gets asked. It is where a copilot stops falling over on “did it break again?” and starts resolving it to the question the operator meant.</p>\n<p>It is the query-side complement to the document-side work in the retrieval series behind <a href=\"/project/archi/\">Archi</a>, the RAG copilot I worked on for CMS computing operations at CERN. <a href=\"/blog/2026-07-06-chunking-strategies-for-rag/\">Chunking</a> sets the ceiling, <a href=\"/blog/2026-07-02-hybrid-search-rag-bm25-vectors/\">hybrid search</a> widens the net, and <a href=\"/blog/2026-07-08-cross-encoder-reranking-rag/\">reranking</a> sharpens the order. Rewriting comes before all of them, because none of that machinery helps if you searched for the wrong thing. Fix the question first, then let the rest of the pipeline do its job.</p>\n<hr>\n<p><em>Diagrams by M. Hassan Ahmed, released under CC0. Image credit: original work by the author.</em></p>","frontmatter":{"title":"Query Rewriting for Better RAG Retrieval","date":"2026-07-23T00:00:00.000Z","description":"Short, vague, follow-up questions don't match how your docs are written. How query rewriting, multi-query expansion, and HyDE fix retrieval before it runs.","thumbnail":{"childImageSharp":{"fluid":{"base64":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAABQAAAALCAIAAADwazoUAAAACXBIWXMAAAsSAAALEgHS3X78AAAB+0lEQVQoz1WPa2+bMBSG+dCu2GAbjAHbmFsgkAuEpMqaNE2bLU2rrtKmTdq0af//h+ykUj9MenR0dPxeZMttVqQ94zZLp16wyXW1f8q2x/HDi7k5VPvncnfKb4+wFLtH0i6d8QBK0oB+ZTnUODTBrma8ELKN1KQbHtrZbb86NNPtYvVpvrif9XtYJt1OxE0oW8ZLTBIwWg7PSDhiUU2CkoYlTORl2M9sL0VeajPzNhPbMw4HQUVE4fIcXIDleMaTbZB0LGpIWHtRQ8MxRDCeeUHhi9L1Uz+sAt3xdOBm8OXM8TNwOV5qYZYI02fjrak+ZvVG5te+mlJRRqqRuo0lBBVZtRZmycWUyy5Il+AHF2bGQlQLNS1nu7Lfj/p9mC+ZqDBRiGmb6SsiHT/3VM91z5+++evPQdxBP5RjpqFZ21RdJDUu5ihtr9TI5gaLHO7nCFeSqBLtlnd3/o/ffHvy4zlPFi6YibIgW22Ow/e/sy8/h69/utdf2eGVrg9uM+C0wcUEZy2v1tAWxPNA9QF8W88RVYBlu3HQ3RR3z/nuNLp/Mdsj6za022AzRmqEdYUCg2niqbmn+zNyhliC3BgRaSESXznhJRKXWFwg8QEL+/wQo/dpu1Go62K8HE1ugLxe+CKHI7xCc2STN90bsJ8v/4OppEHKRAYTQO+af1+0SGyi03LaAAAAAElFTkSuQmCC","aspectRatio":1.899441340782123,"src":"/static/16e70e892c537421cf56a7d2f5b2a2c6/40a76/hero.png","srcSet":"/static/16e70e892c537421cf56a7d2f5b2a2c6/c972b/hero.png 340w,\n/static/16e70e892c537421cf56a7d2f5b2a2c6/27625/hero.png 680w,\n/static/16e70e892c537421cf56a7d2f5b2a2c6/40a76/hero.png 1360w,\n/static/16e70e892c537421cf56a7d2f5b2a2c6/ed396/hero.png 2000w","sizes":"(max-width: 1360px) 100vw, 1360px"}}}}}},"pageContext":{"slug":"/2026-07-23-query-rewriting-rag-retrieval/","previous":"blog/2026-07-20-llm-as-a-judge-evaluating-outputs/","next":"blog/2026-07-22-tracing-llm-agent-opentelemetry/"}},"staticQueryHashes":["32046230"]}