Contextual Retrieval for RAG Pipelines

Chunking a document for RAG strips each piece of its context. Contextual retrieval adds an LLM-written note to every chunk before you index it.

Here is a chunk pulled straight out of an operations logbook. It is the kind of text I index for Archi, the retrieval copilot I worked on for CMS computing operations at CERN:

The figure rose to 99.2% after the fix was applied on the second node.

Now imagine an operator asks, “what was Tier-1 uptime in Q3?” That chunk is the answer, and it will almost certainly not come back. The words “Tier-1”, “uptime”, and “Q3” appear nowhere in it. They were in the section heading three paragraphs up, which got dropped when the document was split. The embedding and the keyword index both see a chunk about “a figure rising after a fix on a node.” The chunk is correct and useless at the same time.

Contextual retrieval is meant to fix this failure. It is a small change to the indexing side of a RAG (retrieval-augmented generation) pipeline. Before you embed a chunk, you have an LLM write a sentence or two that places the chunk in its parent document, and you prepend that sentence to the chunk. The idea is Anthropic’s, published in late 2024. It is worth a post because it is cheap, it stacks cleanly with things you may already run, and the numbers behind it are specific enough to reason about.

This post is for people who already have a RAG pipeline that retrieves and want fewer of these near-misses. I assume you know what an embedding and a chunk are. If you are still deciding how to split documents in the first place, start with the chunking post and come back, because contextual retrieval sits one step downstream of that decision. I cover why chunks lose context, how the fix works, what it does for keyword search, the measured gains, the cost, and where it falls short.

Why chunks lose their context

RAG works by cutting your documents into chunks and embedding each one. At query time, it pulls back the chunks whose embeddings sit closest to the question. You chunk at all for two reasons. Embedding models have a fixed input size, and, more practically, a whole 40-page report embedded as one vector is too blurry to match anything specific.

The tradeoff is that a chunk is read in isolation but was written in context. A sentence in the middle of a report leans on everything above it:

  • the document is about the Q3 site-integration review,
  • this section is the Tier-1 availability table,
  • the “figure” is uptime,
  • “the fix” is the patch named two pages back.

The human author never repeats any of that, because they assume you read top to bottom. The embedding model gets no such luxury. It sees one paragraph with the pronouns and references dangling, and it embeds exactly that.

So the chunk that should match “Tier-1 uptime Q3” instead matches “figures, fixes, and nodes” in the abstract. It then loses to any chunk that happens to repeat the query words, even an off-topic one. You can watch this happen if you log your retrievals: the right document is in the corpus, but it never cracks the top-k (the k best-scoring chunks the retriever hands back). That is a retrieval miss, and it is the single most common reason a RAG answer comes back wrong or empty. Measuring misses is a topic of its own: recall@k and MRR are how you put a number on “the right chunk didn’t show up.”

The fix: write the context back into each chunk

The move is almost embarrassingly direct. For each chunk, send the whole document plus that chunk to a cheap, fast model and ask it for a short situating note. Then glue the note to the front of the chunk and index that. Anthropic’s prompt is essentially this:

<document>
{{WHOLE_DOCUMENT}}
</document>

Here is the chunk we want to situate within the whole document:
<chunk>
{{CHUNK_CONTENT}}
</chunk>

Give a short, succinct context to situate this chunk within the overall
document for the purposes of improving search retrieval of the chunk.
Answer only with the succinct context and nothing else.

Notice that the prompt sends the full document first, then the single chunk, and asks for the context alone with nothing else around it. For the logbook chunk above, a model returns something like “From the Q3 2025 site-integration report, Tier-1 availability section; the figure discussed is Tier-1 uptime.” You prepend that, so the chunk you actually store reads:

From the Q3 2025 site-integration report, Tier-1 availability section; the figure discussed is Tier-1 uptime. The figure rose to 99.2% after the fix was applied on the second node.

The dates, the entities, and the subject are back inside the chunk, where the embedding model can see them. The diagram below traces one chunk through the process. The LLM reads the full document alongside the chunk and writes the note, and only the combined text reaches the indexes.

Contextual indexing pipeline. A whole document feeds an LLM together with one highlighted chunk. The LLM writes a one-to-two sentence context. That context is prepended to the original chunk text to form a contextualized chunk, which is then written to both an embedding vector index and a BM25 keyword index. Below, a comparison band shows the same sentence indexed as a plain chunk, where which figure, system, and quarter are unknown, versus the contextualized chunk, where the entities and dates now live inside the chunk so both index types can match it.

Two details matter and are easy to skip:

  • The note is generated per chunk but conditioned on the whole document. That is why it can resolve “the figure” to “Tier-1 uptime”, even though that information isn’t in the chunk.
  • You keep the original chunk text. You are adding a prefix, not rewriting the content, so nothing the author actually wrote is lost.

Feed the same prefix to the keyword index

Contextual retrieval is usually paired with hybrid search. Hybrid search runs a dense vector search and a sparse keyword search side by side, then fuses the two rankings. The keyword half is typically BM25, a ranking function that scores documents on exact term overlap with the query. Decades on, BM25 still refuses to be beaten on error codes, ticket IDs, and hostnames, the tokens an embedding model tends to smear together.

The prepended context helps both indexes, not just the vector one. Once “Tier-1”, “uptime”, and “Q3 2025” are literally present in the stored chunk, a BM25 query for those exact terms matches it too. This is why Anthropic reports the technique as “contextual embeddings” and “contextual BM25”: the same prefix, fed to both indexes. If you only run dense retrieval today, contextual retrieval is a natural moment to add the sparse half, because the same generated prefix pays off twice.

The measured gains, layer by layer

You don’t have to rely on intuition here. The reductions were measured, and the way they stack tells you where to spend effort. Anthropic used a top-20 retrieval task: did the correct chunk land in the top 20 results? Averaged across their evaluation datasets, they reported these failure rates:

Bar chart of top-20 retrieval failure rate as each layer is added. Embeddings-only baseline is 5.7 percent. Adding contextual embeddings drops it to 3.7 percent, a 35 percent reduction. Adding contextual BM25 on top drops it to 2.9 percent, a 49 percent reduction from baseline. Adding a reranking pass that trims the top 150 candidates down to 20 drops it to 1.9 percent, a 67 percent reduction from baseline. Lower is better.

Read the chart left to right:

  • Contextual embeddings alone take the miss rate from 5.7% to 3.7%, a 35% cut.
  • Adding contextual BM25 brings it to 2.9%, a 49% cut from baseline.
  • Adding a reranking pass takes it to 1.9%, a 67% cut. In a reranking pass you retrieve a wide net of ~150 candidates, then a cross-encoder (a model that reads the query and a candidate together and scores the pair) keeps the best 20.

Anthropic’s full write-up is worth reading for the methodology and the per-dataset spread, which is wider than the average suggests.

The takeaway I keep is the ordering. The single biggest jump is the first one, from doing nothing to contextualizing the embeddings. Each further layer helps less than the one before. That is the usual shape of retrieval work: the first fix buys the most, and you add reranking when 2.9% still isn’t good enough for what the answer is used for.

What it costs, and how prompt caching makes it cheap

The obvious objection is cost. You are now running an LLM call for every chunk in your corpus, at index time. A document that splits into 200 chunks means 200 generations. For a corpus that grows daily, like a logbook, that adds up.

Prompt caching is what makes it affordable. Every one of those 200 calls sends the same whole document and varies only the chunk. If you cache the document prefix, you pay full price to process it once and get a large discount on every later chunk from the same document. Anthropic put the one-time cost at roughly $1.02 per million document tokens with caching on. For most corpora, that makes the indexing bill a rounding error next to the embedding and storage costs you were already paying.

A few practical notes from doing this:

  • It is an ingest-time cost, not a query-time one. The context is generated once, when you index a chunk, and stored with it. Retrieval latency is unchanged, and that is the part your users actually feel. This is the opposite tradeoff from query rewriting or reranking, both of which add work on every request.
  • Use a small, fast model for the note. You are not asking for reasoning. You are asking for one factual sentence grounded in the document. A cheap model is fine and keeps throughput high, so save the expensive model for answering.
  • Regeneration follows the document. If a source document changes, its chunks and their generated context need regenerating. This is the same incremental-indexing bookkeeping you need anyway. Contextual retrieval just adds the per-chunk note to the set of things that go stale when a source updates.

Where it doesn’t help, and how it fails

Contextual retrieval has sharp edges, and knowing them up front saves you a confusing week.

Self-contained chunks gain little. An FAQ entry, a standalone function’s docstring, or a glossary definition already carries its own context. The generated note just restates the obvious, and you still pay to generate it. If your corpus is mostly short, independent items, measure before assuming a win. The technique earns its keep on long, flowing documents where chunks lean hard on their surroundings.

The note is model output, so it can be wrong. The model can misread the document and write a confident, incorrect context, such as the wrong date or the wrong system. That bad prefix goes into your index and can pull the chunk up for queries it shouldn’t answer. It is the same class of risk as rendering model output as trusted content: retrieved text is not trustworthy just because your own pipeline produced it. Keeping the note short and strictly grounded in the document limits the blast radius, but it does not eliminate it.

It fixes context loss inside a document, not across documents. Sometimes a chunk’s true context lives in a different document, such as a config value defined in one file and used in another. Feeding the LLM only the chunk’s own document won’t recover it. Cross-document context is a harder problem that needs a different tool, usually explicit metadata or a graph over your sources.

It does not replace good chunking. Suppose your chunks are badly sized: a paragraph cut mid-sentence, or a chunk spanning two unrelated sections. A situating note papers over the seams, but the underlying split is still bad. Get chunking roughly right first. Contextual retrieval is a multiplier on a reasonable pipeline, not a rescue for a broken one.

What I would reach for

If I were adding retrieval quality to a pipeline from scratch today, I would go in this order:

  1. Chunk sensibly.
  2. Add contextual embeddings.
  3. Add contextual BM25 as the sparse half of hybrid search.
  4. Only then reach for reranking.

That order mirrors the failure-rate chart: most of the gain comes early, and each later layer is more machinery for less return. Archi indexes exactly the dense, cross-referential technical documents where this matters: logbooks where “the agent” and “the node” mean something only if you read the entry above. For a corpus like that, the first two layers are the ones I would not skip. The whole point of that copilot is that an operator can ask a question in their own words and land on the one logbook entry from eight months ago that actually explains the failure. A chunk that quietly forgot which system it was about is the difference between finding that entry and staring at a blank answer.

I like contextual retrieval because it is honest about where it acts. It does not touch your query path, it does not add latency, and it does not ask you to swap your vector store. It changes what you write to the index, once, at ingest, and lets everything downstream stay the same. That is usually the cheapest kind of improvement to adopt and the easiest to undo if it doesn’t pan out. For a change to production retrieval, that is worth as much as the accuracy itself.


Diagrams by M. Hassan Ahmed, released under CC0. No external image was used for this post; the figures are original work by the author. Retrieval failure-rate figures are from Anthropic’s “Introducing Contextual Retrieval” (2024).