Contextual Retrieval for RAG, Step by Step
Vector search misses chunks that lost their document context. Contextual Retrieval prepends a short LLM-written summary to each chunk before you embed it.
Here is a chunk that once cost me an afternoon of debugging. It read, roughly, “The failure rate rose to 5.7% after the migration and stayed there for two days before the rollback.” A perfectly good sentence, sitting in a vector index, invisible to every query that mattered. Someone searched “PhEDEx transfer errors in Q3”, and the chunk never came back. The chunk never says PhEDEx, never says Q3, and never says transfer. All of that lived in the report’s title and the section heading, three thousand tokens up the page. The moment I split the document into chunks, that context was gone.
Contextual Retrieval is built to fix this failure mode, and it is the one that quietly caps the quality of most RAG (retrieval-augmented generation) systems. This post is for engineers who already run a retrieval pipeline and have watched it miss chunks they know are in the index. If you want a fix more principled than shrinking the chunk size and hoping, read on. I worked on Archi, a retrieval copilot for CMS computing operations at CERN, and the example above is the kind of question shifters actually ask it.
I walk through why chunks lose context, what Contextual Retrieval does, the published numbers, a working indexing loop, what prompt caching does to the cost, where it bites, and how to measure whether it helped you.
Why chunks lose their context
Retrieval starts by cutting documents into chunks, because you cannot embed a 40-page runbook as one vector and expect it to match a one-line question. I wrote about where to make those cuts in chunking strategies for RAG. But every cut throws away the surrounding text.
An embedding model (the model that turns text into a vector) only sees the tokens inside the chunk. It has no idea the paragraph came from the transfer-system report rather than the tape-storage one. If the disambiguating words are not in the chunk, they are not in the vector. No amount of tuning the nearest-neighbor search recovers information that was never encoded.
The same blind spot hits the keyword half of a hybrid search setup. BM25, the standard keyword-ranking function, scores on exact term overlap. A query for “PhEDEx” only matches a chunk that literally contains “PhEDEx.” Both retrieval paths, dense (vector) and sparse (keyword), fail for the same underlying reason. The chunk was written for a reader who could see the rest of the document, and at query time nobody can.
What Contextual Retrieval does
The idea is small and a little obvious in hindsight:
- Before you embed a chunk, generate a short piece of text that says where the chunk sits in its document. An LLM that can see the whole document writes it.
- Prepend that text to the chunk.
- Embed the context plus the chunk together, and index the same combined text for BM25.
Anthropic’s writeup names the two variants: Contextual Embeddings for the dense side and Contextual BM25 for the sparse side.
The prefix does not need to be long. Fifty to a hundred tokens is enough to carry the identifiers a query is likely to use: the system name, the document title, the section, the time period. The instruction to the model is deliberately plain. Anthropic’s version reads close to this:
<document>
{{WHOLE_DOCUMENT}}
</document>
Here is the chunk we want to situate within the whole document:
<chunk>
{{CHUNK_CONTENT}}
</chunk>
Give a short, succinct context to situate this chunk within the overall
document for the purposes of improving search retrieval of the chunk.
Answer only with the succinct context and nothing else.The last line, “nothing else”, matters. You want the model to return the context string, not a preamble about how it will now describe the chunk. The output gets glued to the front of the chunk, and anything chatty pollutes the vector.
The published numbers, and how far to trust them
Anthropic ran this on a set of retrieval benchmarks and reported the top-20 retrieval failure rate: the fraction of queries whose relevant chunk never made it into the top 20 results.
- Contextual Embeddings alone took it from 5.7% to 3.7%, about a 35% reduction.
- Adding Contextual BM25 brought it to 2.9%, a 49% reduction against the baseline.
- A reranker on top of both pushed it to 1.9%, a 67% reduction.
Those figures are from Anthropic’s own writeup. The point I would hold onto is not the exact percentages but the shape: the two contextual variants stack, and reranking complements them rather than replacing them.
I would not treat 49% as a promise for your corpus. It is a benchmark result on their data. How close you get depends on how much your chunks actually depend on out-of-chunk context, and that is a property of your documents. Measure it; I come back to how at the end.
Building the indexing loop
The indexing loop is straightforward. For each document, generate a context string per chunk, prepend it, and hand the combined text to both your embedding call and your keyword index. In the code below, contextualize makes the LLM call for one chunk, and index_document runs it over every chunk of a document:
import anthropic
client = anthropic.Anthropic()
CONTEXT_PROMPT = """Here is the chunk we want to situate within the whole document:
<chunk>
{chunk}
</chunk>
Give a short, succinct context to situate this chunk within the overall
document for the purposes of improving search retrieval of the chunk.
Answer only with the succinct context and nothing else."""
def contextualize(document: str, chunk: str) -> str:
resp = client.messages.create(
model="claude-haiku-4-5",
max_tokens=120,
messages=[
{
"role": "user",
"content": [
{
"type": "text",
"text": f"<document>\n{document}\n</document>",
"cache_control": {"type": "ephemeral"},
},
{
"type": "text",
"text": CONTEXT_PROMPT.format(chunk=chunk),
},
],
}
],
)
return resp.content[0].text.strip()
def index_document(document: str, chunks: list[str]):
for chunk in chunks:
context = contextualize(document, chunk)
combined = f"{context}\n\n{chunk}"
vector = embed(combined) # your embedding model
bm25_index.add(combined) # same text into the keyword index
store(chunk=chunk, context=context, vector=vector)Two details carry the whole thing:
- The
cache_controlblock marks the document for prompt caching. The document is processed once, and every later chunk in the same document reads it from cache instead of re-sending it. - The original
chunkis stored separately from thecombinedtext. The context prefix exists to be found, not to be read by the generation model. When a chunk is retrieved, you feed the clean chunk into the answer prompt. The synthesized context can be slightly wrong, and you do not want that noise in what the model reasons over.
A small model is the right call here. The context step runs once per chunk at index time, not per query, so its latency never touches the user. Its cost is real, though, once you multiply by every chunk in the corpus. I use a fast, cheap model for the context step and save the larger model for generation.
Why prompt caching makes it viable
Without caching, this looks expensive. For every chunk you would re-send the entire parent document. An 8,000-token document split into ten chunks means sending 80,000 document tokens to produce ten short contexts.
Caching collapses that. The document is written to the cache once. Each chunk call then pays the much lower cache-read rate for the document, plus the full rate for its own ~800 tokens and the ~100-token completion.
Anthropic put the one-time cost at roughly $1.02 per million document tokens under these assumptions: 800-token chunks, 8k-token documents, a 50-token instruction, and about 100 tokens of context per chunk. Your number will move with model choice and chunk size. The structural point holds anyway: caching turns “run an LLM over every chunk” from a line item you argue about into a rounding error.
One caveat worth writing down: cache entries expire, with a default time-to-live measured in minutes. Process all chunks of a document back to back rather than interleaving documents, or you pay to warm the cache more than once.
What it costs you, and where it bites
The honest tradeoff is that indexing becomes heavier and slower, and there are a few failure modes to plan for.
Indexing is now a bigger job. You have added an LLM call per chunk to a pipeline that used to be a tokenizer and an embedding call. For a corpus you reindex often, that is a real operational cost. A full rebuild goes from minutes to something you schedule.
Re-indexing on change is the sharper edge. When a document is edited, the context strings for its chunks can shift. A change to one section can invalidate the synthesized context of chunks elsewhere in the same document. In practice I regenerate context for the whole document on any edit rather than trying to be clever about which chunks were affected. Getting that dependency tracking wrong leaves you with stale context, which is worse than none.
The context can be wrong. The model can misread a long document and situate a chunk under the wrong heading. That is exactly why the clean chunk, not the combined text, goes to the generation step.
The gain is uneven. Documents where each chunk already names its own subject see almost no benefit. If you index a pile of standalone FAQ entries, you have paid the indexing cost for close to nothing. The technique earns its keep on long, structured documents where meaning accumulates down the page, which describes most internal engineering documentation.
How to know it actually helped
Do not ship this on the strength of a benchmark from someone else’s corpus. The earlier post on measuring RAG retrieval quality exists to make this measurable:
- Take fifty to a hundred real queries.
- Label which chunks are actually relevant to each.
- Compute recall@k (the share of relevant chunks that appear in the top k results) with and without the context prefix, on the same index.
If recall@20 barely moves, your chunks were already self-describing, and Contextual Retrieval is a cost you do not need. If it jumps, you have found the queries that were failing silently, and now you can see them.
Run the comparison before you commit the indexing budget, not after. It is a cheap experiment: contextualize one slice of the corpus, index it twice, and diff the recall. The answer is specific to your documents in a way no blog post, including this one, can tell you in advance.
Where it fits in the pipeline
My first mistake was reaching for it too early. Contextual Retrieval sharpens retrieval, but it cannot fix chunks that were cut badly in the first place. It also cannot rescue a query that fails because the answer simply is not in the corpus. The order that works is the same one across this whole series:
- Chunk so the answer stays whole.
- Retrieve across both dense and sparse so the answer makes the shortlist.
- Add context so the shortlist matches the words people actually search.
- Rerank so the best chunk reaches the prompt.
Contextual Retrieval slots in as the third step, and it is most useful once the first two are already solid.
Closing
The bug at the top of this post was never a retrieval bug in the search sense. The right chunk was in the index and the query was reasonable. The two never met because the chunk had been separated from the words that identified it. Contextual Retrieval fixes that at index time, cheaply, by asking a model to write down the context a chunk lost when it was cut out of its document.
I add it to the retrieval pipeline behind Archi when a document is long enough that its chunks stop making sense on their own, which at CERN is most of them. For implementation details straight from the source, I keep two references open: Anthropic’s Contextual Retrieval writeup and the accompanying Claude cookbook guide.
Diagrams by M. Hassan Ahmed, released under CC0. Image credit: original work by the author.