HyDE: Hypothetical Document Embeddings for RAG

Vector search fails when a short question looks nothing like its answer. HyDE has an LLM draft a fake answer, embeds that, and retrieves against it instead.

An operator types “why is the spring transfer error back” into the copilot. The retriever hands back three chunks about transfers in general and none about the incident they mean. Yet the corpus has the answer: a postmortem in the index describes exactly this incident, in full sentences, with the root cause and the fix. The retriever walked right past it.

The reason is a mismatch that pure vector search never really solves: a question and its answer do not look alike. The question is five words and a hedge. The answer is two paragraphs of declarative prose. You embed both into the same vector space and expect them to land near each other. But they were written for different purposes, and the embedding model puts them in different neighborhoods.

This post is about a small, slightly strange trick that closes that gap, called HyDE (Hypothetical Document Embeddings). It is for engineers who already have a working vector RAG (retrieval-augmented generation) pipeline and keep watching it miss on short or vague queries. I will show what the mismatch is, how HyDE works, the code, and where it makes things worse instead of better.

Why questions and answers land in different places

Dense retrieval assumes that the vector for a query sits close to the vector for the passage that answers it. That assumption mostly holds for long, well-formed queries that share vocabulary with the target. It falls apart for the queries real users actually type.

A short keyword query gives the encoder (the model that turns text into a vector) almost nothing to work with. “spring transfer error” has no verb, no context, and no hint of what a good answer would contain. The passage that answers it is full of the opposite: specifics, a timeline, a resolution. The two strings share a couple of tokens and nothing structural.

The asymmetry between short queries and long documents is a known problem in information retrieval. The usual fix is to train a retriever on labeled query-document pairs so it learns to bridge the two. That is expensive, and it needs data you probably do not have for an internal corpus.

A short query embeds far from the passage that answers it; HyDE embeds a drafted answer instead, which lands near the real passages

The insight behind HyDE is that you already own a tool that is very good at turning a question into something answer-shaped: the LLM sitting at the end of your pipeline.

How HyDE works: search with a fake answer

HyDE comes from a 2022 paper by Luyu Gao and colleagues, Precise Zero-Shot Dense Retrieval without Relevance Labels. The method is almost uncomfortably simple:

  1. Ask the LLM to write a document that would answer the query.
  2. Embed that generated document instead of the query.
  3. Retrieve real passages with that embedding, and hand them to the generator as usual.

You never show the generated document to the user, and you do not care whether it is factually correct. Its job is to be shaped like a real answer, so that its embedding lands in the same region as the real answers in your corpus.

Why a wrong draft still works

The part that trips people up on first read: the drafted document is often wrong. If the model has never seen your internal incident, it will confidently invent a plausible-sounding cause.

That is fine, and the paper is explicit about why. The generated text captures the pattern of a relevant answer. The encoder compresses text into a single dense vector, and that bottleneck keeps the pattern while discarding the specific invented facts. You are not retrieving the hallucination. You are using the hallucination’s shape to find the real thing.

What the paper showed

The original paper generated the hypothetical document with an instruction-tuned model (InstructGPT at the time) and embedded it with Contriever, an unsupervised dense encoder. The reported result is the interesting bit. With no labeled training data at all, HyDE matched or beat fine-tuned retrievers across several tasks and languages. For a team that cannot label its own corpus, “as good as fine-tuned, with zero labels” is the whole pitch.

The HyDE step sits between the query and the vector search: draft an answer, embed it, then retrieve; the corpus embeddings never change

Notice what does not change. Your index, your chunks, the embedding model for your corpus, and your vector store all stay untouched. HyDE only changes what you search with. That makes it cheap to bolt onto an existing pipeline, and cheap to remove if it does not help.

Implementation

The whole technique is one extra LLM call before retrieval. Here is the core, using an OpenAI-style client and whatever embedding model you already have wired up. The three numbered comments map to the three steps above:

async def hyde_retrieve(query: str, k: int = 8):
    # 1. Draft a hypothetical answer. Keep it short; you want the
    #    shape of an answer, not an essay.
    prompt = (
        "Write a short passage that directly answers the question "
        "as if it came from internal documentation. Two or three "
        "sentences. Do not hedge, do not say you are unsure.\n\n"
        f"Question: {query}\nPassage:"
    )
    draft = await llm.complete(prompt, max_tokens=160, temperature=0.7)

    # 2. Embed the draft, not the query.
    vector = await embed(draft.text)

    # 3. Search the untouched corpus with the draft's embedding.
    return await vector_store.search(vector, k=k)

Two details matter more than they look.

Keep the draft short. A long generated document drifts off topic and pulls the embedding toward whatever the model rambled into. Two or three sentences give the encoder an answer-shaped input, and they keep the extra latency and token cost down.

Consider averaging a few drafts. The paper generates several hypothetical documents and averages their embeddings, together with the query’s own embedding, to smooth out any single bad generation. In practice, that means N generation calls instead of one, so it is a direct trade of latency for stability. I would start with a single draft and measure. Add more drafts only if retrieval quality swings on one unlucky generation. The averaged version looks like this:

async def hyde_retrieve_averaged(query: str, k: int = 8, n: int = 4):
    drafts = await asyncio.gather(*(draft_answer(query) for _ in range(n)))
    vectors = await asyncio.gather(*(embed(d) for d in drafts))
    vectors.append(await embed(query))          # include the raw query too
    mean_vector = sum(vectors) / len(vectors)   # numpy arrays
    return await vector_store.search(mean_vector, k=k)

The code runs the N generation calls concurrently rather than awaiting them one by one. That is the same asyncio.gather fan-out I lean on across these backends. The FastAPI event loop will happily keep all N in flight, as long as nothing in the path blocks it.

Where HyDE makes things worse

HyDE is not free, and it is not always an improvement. The failure modes are specific enough to name.

It adds a full LLM call to the front of every query. Retrieval used to be one embedding call and a vector lookup: tens of milliseconds. Now a generation step runs in front of it, and generation is the slowest thing in the pipeline. In a chat UI where the user is already waiting on a streamed answer, an extra second before retrieval even starts is a real regression. Cache aggressively, or reserve HyDE for the queries that need it instead of running it on everything.

It hurts on exact-identifier queries. This is the important one, and it is the mirror image of the problem HyDE solves. When the user searches for T1_US_FNAL or a specific error code, they do not want the shape of an answer. They want the document that literally contains that string. HyDE will happily draft a passage about transfer errors in general and embed away the one token that mattered.

That is the weakness I wrote about in hybrid search for RAG: dense retrieval blurs exact strings together, and the fix is to keep a BM25 keyword path alongside the vector path. HyDE lives entirely on the dense side, so it inherits that weakness and amplifies it. Run it in parallel with a keyword retriever, not instead of one.

It can amplify a confidently wrong model. The paper claims the encoder filters out invented facts, and that mostly holds. But on a niche corpus where the model genuinely has no idea, the draft can be wrong in its structure, not just its facts. Then the embedding points at the wrong neighborhood entirely. If your domain is far from anything in the model’s training data, test before you trust it.

It is another thing to measure. HyDE changes retrieval, so you cannot tell whether it helped by reading a few outputs and nodding. You need the same golden-set discipline as any other retrieval change: a set of query-to-expected-chunk pairs, scored with recall@k and MRR, run with HyDE on and off. I have been surprised in both directions. On vague natural-language questions, it lifted recall noticeably. On a corpus full of identifiers, it quietly made things worse until I gated it behind a keyword path.

What I would do differently

The mistake is treating HyDE as a global switch for the whole pipeline. It is a per-query decision:

  • Where it earns its extra LLM call: short, natural-language questions with no exact tokens in them.
  • Where it costs you: identifier lookups and error-code searches.

The version I would build now checks the query first, with a cheap classifier or even a simple heuristic. Does it contain a code-like token? Is it under three words? Does it read like a keyword string? The query takes the HyDE path only when it looks like the kind HyDE helps; everything else goes straight to hybrid retrieval. That keeps the latency budget honest and stops HyDE from sabotaging the exact-match queries it was never meant to handle.

Also, measure it against a plain query rewrite before you commit. Rewriting the query into a cleaner search string is a lighter-weight member of the same family, and sometimes it captures most of the gain for a fraction of the cost. HyDE wins when the gap between question-shape and answer-shape is the actual problem. Confirm that it is your problem before you pay for it.

Closing

HyDE sounds like a hack and turns out to be principled. You are not retrieving a made-up answer. You are using its shape to reach the real one, and letting the encoder throw away everything the model invented.

It fits neatly next to the rest of the retrieval stack I have been writing about:

  • chunking decides what a passage is;
  • hybrid search decides how you find it;
  • reranking decides the final order;
  • HyDE decides what you search with in the first place.

It is the kind of change I reach for on Archi, the RAG copilot I worked on for CMS computing operations at CERN, because operators ask short, messy questions and the answers are long, careful writeups. The gap between the two is exactly what HyDE is for. Just keep the keyword path next to it, gate it to the queries that need it, and score it against a golden set before you believe it helped.


Diagrams by M. Hassan Ahmed, released under CC0. Image credit: original work by the author.