Query Rewriting for Better RAG Retrieval
Short, vague, follow-up questions don't match how your docs are written. How query rewriting, multi-query expansion, and HyDE fix retrieval before it runs.
An operator on shift types “did it break again?” into the copilot. The retriever embeds that four-word string, searches the index, and hands back the nearest chunks, which could be about anything, because nothing in the query says what “it” is. A minute earlier the same person had asked about a T1_US_FNAL transfer error, so a human would know exactly what they meant. The index has no memory of that turn. It only saw four words with a pronoun in them.
Most RAG (retrieval-augmented generation) debugging skips past this failure. When retrieval returns junk, the reflex is to blame the index: bad chunks, the wrong embedding model, a missing reranker. Sometimes that is right. But often the chunk you need is sitting in the index, embedded perfectly well. The problem is that the query you searched with looks nothing like the document that answers it. You cannot fix that downstream. You fix it by rewriting the query before it reaches the retriever.
This post is for engineers who already have a working RAG pipeline and keep watching it miss on short questions, vague questions, and follow-ups. I will cover why the raw query is so often the weak link, then three techniques that sit in front of retrieval:
- contextualizing follow-ups against the chat history;
- multi-query expansion;
- HyDE (Hypothetical Document Embeddings).
Each one buys you something different, and each can quietly make retrieval worse if you use it in the wrong place.
Why the raw query is the weak link
Dense retrieval compares a query vector against document vectors, so it only works when the two land near each other in embedding space. The trouble is that a question and its answer are written in different registers (styles of language). The question is short, informal, and often assumes context the asker already has. The document is a full paragraph of prose that never repeats the question back.
Take “why is FNAL slow” and a logbook entry titled “FTS staging backlog on the US Tier-1 causing transfer throughput degradation.” They are about the same thing. Their embeddings can still sit far apart, because almost none of the words overlap and the two texts have different shapes.
That gap shows up in three recurring forms.
Too short to be specific. A three-word query carries almost no signal. There is not enough text for the embedding to land anywhere precise, so it lands in a vague neighborhood and pulls back a vague set of chunks.
Vocabulary mismatch. The user says “slow”; the document says “throughput degradation.” The user says “FNAL”; the document says T1_US_FNAL. Dense retrieval is supposed to bridge synonyms better than keyword search does, and it partly does. That is exactly why the cases it misses are easy to overlook. Hybrid search recovers some of these by adding exact-match BM25 keyword scoring back into the mix, but it cannot invent a term the user never typed.
Conversational follow-ups. In a chat interface, half the questions are not standalone: “did it break again?”, “what about the other site?”, “why?” They make sense only against the previous turns. Embed them literally and you retrieve for the wrong thing entirely.
The fix for all three is a transformation stage between the user and the retriever. The index stays exactly as it is; you change what you ask it.
Contextualizing follow-up questions
If you run a chat copilot, build this one first. It fixes the most common failure for the least effort.
The idea is to use the conversation history to rewrite a context-dependent question into a standalone one before retrieval. “did it break again?” becomes “Did the T1_US_FNAL transfer error recur after the fix?” The rewritten query stands on its own, so the retriever now has something to work with.
It takes one LLM call with the history and the latest question. The code below builds that prompt, skips the call when there is no history, and retrieves with the rewritten query:
CONDENSE_PROMPT = """Given the conversation and a follow-up question,
rewrite the follow-up as a standalone question that includes the
entities and context it depends on. Do not answer it. If it is already
standalone, return it unchanged.
Conversation:
{history}
Follow-up: {question}
Standalone question:"""
def condense(history: list[dict], question: str) -> str:
if not history:
return question # nothing to fold in, skip the call
convo = "\n".join(f"{m['role']}: {m['content']}" for m in history[-6:])
rewritten = llm.generate(
CONDENSE_PROMPT.format(history=convo, question=question)
).strip()
return rewritten
# then retrieve with the rewritten query, not the raw one
standalone = condense(history, user_question)
chunks = retriever.search(standalone, k=8)Two details matter:
- Only the last handful of turns go into the prompt. A full transcript costs more, and it drags in stale context that pulls the rewrite off target.
- The early return saves a call. When there is no history, which is the first question of every conversation and a meaningful fraction of traffic, no LLM call is made.
This is the “condense question” pattern that LangChain’s history-aware retriever formalizes. You can write it in a dozen lines without the framework.
Multi-query expansion
Contextualizing fixes follow-ups. It does nothing for a standalone question that is simply too vague or unluckily phrased. For that, generate several rewordings and search with all of them.
The reasoning: any single phrasing is one sample from the many ways to ask the question, and it might be a bad sample. Ask an LLM for three or four alternate phrasings, retrieve for each, and merge the results. Where one phrasing misses the right chunk, another catches it, so recall goes up. LangChain’s MultiQueryRetriever packages exactly this, but the mechanics are worth seeing directly. The function below searches with the original question plus the generated variants, then fuses the ranked lists:
EXPAND_PROMPT = """Generate 3 alternative phrasings of this question,
each from a different angle, one per line, no numbering:
{question}"""
def multi_query(question: str, k: int = 8) -> list[dict]:
variants = llm.generate(EXPAND_PROMPT.format(question=question))
queries = [question] + [q.strip() for q in variants.splitlines() if q.strip()]
ranked_lists = [retriever.search(q, k=k) for q in queries]
return reciprocal_rank_fusion(ranked_lists)Merging the results with Reciprocal Rank Fusion
The part people get wrong is the merge. You now have several ranked lists and need one. Do not just concatenate and dedupe, because that throws away the rank information.
Use Reciprocal Rank Fusion (RRF) instead, the same fusion I used to combine keyword and vector results in the hybrid search post. RRF scores each document by summing 1 / (k + rank) across every list it appears in. A chunk that shows up near the top of two different rewrites therefore beats one that ranks first in a single list and nowhere else:
def reciprocal_rank_fusion(ranked_lists, k: int = 60):
scores, seen = {}, {}
for lst in ranked_lists:
for rank, doc in enumerate(lst):
scores[doc["id"]] = scores.get(doc["id"], 0) + 1 / (k + rank)
seen[doc["id"]] = doc
ordered = sorted(scores, key=scores.get, reverse=True)
return [seen[i] for i in ordered]The cost is real. Three rewrites mean three retrieval round trips, plus the generation that produced them. On a vector index that is usually cheap, but it is not free, and it adds up if you also rerank afterward.
HyDE: search with a hypothetical answer
Multi-query still searches with questions. HyDE, from Gao et al. at ACL 2023, takes a stranger route. It asks the LLM to write a fake answer to the question, then embeds that fake answer and searches with it.
The reasoning follows directly from the register problem. Your index is full of answer-shaped documents, so an answer-shaped query lands closer to the right neighborhood than the terse question ever could.
It does not matter if the drafted answer is factually wrong. You never show it to the user; you only use its embedding as a better search key. A plausible-but-wrong paragraph about FTS staging backlogs still sits near the real logbook entries about FTS staging backlogs. In code, the only change from normal retrieval is what gets embedded:
HYDE_PROMPT = """Write a short, plausible paragraph that could answer
this question. It does not need to be correct, just realistic in style
and terminology:
{question}"""
def hyde_search(question: str, k: int = 8) -> list[dict]:
draft = llm.generate(HYDE_PROMPT.format(question=question))
return retriever.search_by_text(draft, k=k) # embeds the draft, not the queryHyDE earns its keep when the query and the documents are written in genuinely different registers: terse operator shorthand against formal write-ups, or a domain where the user does not know the right vocabulary. In the paper, it was strong precisely where there were no training labels to fine-tune a retriever, which describes most internal corpora. Still, it is the technique I reach for last, because it is the most likely to backfire, as the next section covers.
Tradeoffs and failure modes
Every technique here spends an extra LLM generation to improve retrieval, and each has a specific way of going wrong.
The latency is not free, and it lands early. The rewrite happens before retrieval, which happens before generation. You have added a full model call to the front of the request, and the user sees nothing until it finishes. If you stream the answer over SSE, the rewrite is pure dead time before the first token, so measure it as part of time-to-first-token, not as a background cost. Use a small, fast model for the rewrite; it does not need the model that writes the final answer.
Rewrites drift. A condense or expansion step can hallucinate an entity that was never in the conversation, or quietly change the question’s meaning. If “why is FNAL slow” gets expanded into “why is the CERN Tier-0 slow,” you now retrieve confidently for the wrong site. Log the rewritten query next to the original whenever a request looks off. The bug is often obvious the moment you can see what you actually searched for.
HyDE amplifies wrong assumptions. This failure mode mirrors its strength. If the drafted answer is wrong in the wrong direction, its embedding points at the wrong neighborhood, and you retrieve confidently irrelevant chunks. On a question with a specific factual answer that the user already phrased well, HyDE can lose to plain search with the original query. It helps most on vague questions and hurts most on precise ones.
Over-expansion dilutes results. More rewrites are not always better. Past three or four, the extra phrasings start repeating each other. Each additional list you fuse in also gives off-target chunks another chance to accumulate RRF score. Recall stops climbing, and precision slips.
It breaks naive caching. If you cache retrieval results keyed on the query string, a rewrite step changes the key on every request. A non-deterministic rewrite can even produce two different keys for the same question. Cache on the rewritten query, or downstream of it, not on the raw input.
None of this is visible unless you measure. Build a small golden set of real queries with their correct chunks, and score recall with and without each technique, exactly as in the retrieval-quality post. Query rewriting is a change to your retriever, and it deserves the same measurement as any other retriever change.
What I would do differently: add one technique at a time
If I were adding this to a pipeline from scratch, I would resist turning on all three at once. The order I would follow:
- Start with contextualization. In a chat product it fixes the single most common miss, the unresolved follow-up, and it is cheap and hard to get wrong. Ship it and measure it before doing anything else.
- Add multi-query only if standalone queries still miss. It is the safer of the two remaining options, because RRF is forgiving of a bad rewrite: one off-target list gets outvoted.
- Use HyDE last, and only with evidence. Reach for it only when you have a measured recall gap on vague queries with little vocabulary overlap that multi-query did not close. A/B test it against plain retrieval before you trust it, because it is the one most likely to quietly make some queries worse while helping others.
The order matters because these techniques stack multiplicatively in cost and additively in ways to fail. One well-placed rewrite beats three stacked ones you cannot debug.
Closing
Query rewriting decides what your retriever actually gets asked. It is where a copilot stops falling over on “did it break again?” and starts resolving it to the question the operator meant.
It is the query-side complement to the document-side work in the retrieval series behind Archi, the RAG copilot I worked on for CMS computing operations at CERN. Chunking sets the ceiling, hybrid search widens the net, and reranking sharpens the order. Rewriting comes before all of them, because none of that machinery helps if you searched for the wrong thing. Fix the question first, then let the rest of the pipeline do its job.
Diagrams by M. Hassan Ahmed, released under CC0. Image credit: original work by the author.