Parent Document Retrieval: Search Small, Answer Big
Small chunks search precisely but read like fragments. Parent document retrieval matches on child chunks, then feeds the whole parent to the model.
Chunk size is the setting people tune last and regret first. Make chunks big and retrieval gets vague. A 2,000-token chunk embeds down to one vector that averages five different subtopics, so it matches everything weakly and nothing sharply. Make chunks small and search gets precise, but now the model reads a 200-token fragment that starts mid-sentence and drops the definition it needed two paragraphs up. A single number trades answer quality against retrieval quality, and no value of that number wins both.
Parent document retrieval stops treating those as the same number: you search over small chunks and answer with big ones. This post is for engineers who already have a working RAG pipeline and see retrieval find the right document but hand the model a passage too narrow to answer from. I walk through the two-index setup, a framework-free implementation, how to size the chunks, and the specific ways it breaks in practice.
The two jobs a chunk has to do
A chunk in a RAG (retrieval-augmented generation) system does two unrelated jobs:
- It is the unit you search: the thing that gets embedded and compared against the query.
- It is the unit you read: the text that lands in the model’s context.
Chunking strategies usually try to satisfy both with one split, and that is the tension. Search wants a chunk small and focused enough that its embedding means one thing. Generation wants a chunk large enough to stand on its own, with the surrounding sentences, the table header the row belongs to, and the “it” that refers to something named earlier.
The fix is to stop forcing one chunk to do both. Split each document twice. Big parent chunks are what the model reads. Small child chunks, carved out of each parent, are what you embed and search. When a child matches, you do not return the child; you return the parent it came from.
Only the children go into the vector index. The parents go into a plain key-value store keyed by an id that every child carries as metadata. That store can be a dict in memory, a Postgres table, a Redis hash, or whatever you already run. So you have two stores for one document, and the link between them is that parent_id.
Building the two indexes
Here is the index side without a framework, so the moving parts are visible. It has two splitters, a docstore, and a vector index where the payload on each vector is the child’s parent_id. Watch the last two lines: only the children are embedded.
from dataclasses import dataclass
@dataclass
class Child:
text: str
parent_id: str
def split(text: str, size: int, overlap: int) -> list[str]:
# token- or character-based; use your tokenizer in production
step = size - overlap
return [text[i:i + size] for i in range(0, len(text), step)]
parents: dict[str, str] = {} # parent_id -> full parent text
children: list[Child] = []
for i, parent in enumerate(split(document, size=1800, overlap=200)):
pid = f"{doc_id}:p{i}"
parents[pid] = parent
for chunk in split(parent, size=350, overlap=50):
children.append(Child(text=chunk, parent_id=pid))
# embed ONLY the children
vectors = embed([c.text for c in children])
index.add(vectors, payloads=children) # your vector store's add()The parent text is never embedded, and that is the point people miss on the first read. The doc store is dumb storage, and the vector index only ever sees the small chunks.
If you are on LangChain, its ParentDocumentRetriever wires this up for you. LlamaIndex ships the same idea as its auto-merging retriever. The mechanics underneath are the same, and knowing them is what lets you debug the framework when it does something surprising.
At query time, children collapse back to parents
Query time is where the two indexes come together. You embed the query and search the child index for the top-k children. Then you map each hit back to its parent and dedupe. Several children from the same parent will often match one query, and you do not want to send that parent three times.
The function below walks the child hits in score order, skips any parent it has already added, and stops once it has max_parents parents:
def retrieve(query: str, k_children: int = 8, max_parents: int = 3) -> list[str]:
hits = index.search(embed([query])[0], top_k=k_children) # sorted by score
seen, out = set(), []
for hit in hits:
pid = hit.payload.parent_id
if pid in seen:
continue
seen.add(pid)
out.append(parents[pid]) # fetch the whole parent
if len(out) >= max_parents:
break
return outTwo knobs matter here, and they are not the same knob:
k_childrencontrols how wide you search. Pull more children than you need, because several will collapse into the same parent.max_parentscontrols how much text you actually send. Search over eight children, come away with three parents, and keep the context focused.
If you set k_children too close to max_parents, you will occasionally get two parents back when you asked for three. That happens when, for example, the eighth-ranked child shared a parent with the second.
Sizing parents and children
The numbers in the code, 1,800-token parents and 350-token children, are a reasonable starting point, not a law. I reason about the two sizes differently.
Child size is bounded by your embedding model. Most retrieval models were trained on short passages, and their quality degrades well before their advertised token limit. Children of 250 to 400 tokens stay inside the range the model actually handles well. This is the same concern behind picking an embedding model in the first place: the child is what gets embedded, so the child is what the model’s limits apply to.
Parent size is bounded by two things pulling against each other. A parent must be big enough that the passage is self-contained, and small enough that three of them do not blow your context budget or bury the answer.
There is a real reason not to just make parents enormous and send one. Models attend unevenly across a long context and reliably lose information stranded in the middle of it, an effect documented in Liu et al.’s “Lost in the Middle”. Three tight parents beat one sprawling one, and both beat twenty fragments.
If your documents have real structure (Markdown headings, sections, a ## every few paragraphs), split parents on those boundaries instead of a fixed token count. That way a parent is a section rather than an arbitrary window.
Where it breaks
The happy path is a hundred lines. The failures all live in the seams between the two stores.
The two stores drift apart. The vector index and the doc store are separate systems that have to agree on parent_id. Re-index the documents, change the id scheme, or partially reload one store and not the other, and you get children whose parent_id points at a parent that no longer exists. The retrieve loop then throws a KeyError mid-request, or worse, silently skips the hit.
The fix is to treat the two writes as one operation. Never publish new children pointing at parents you have not written yet, and when you delete a document, delete it from both stores. This is the same zero-downtime reindexing discipline, spread across two stores instead of one.
Dedup throws away a relevance signal. A parent that matched on three of its children is probably more relevant than one that matched on one. The simple loop above ignores that: it keeps the first parent it sees and moves on. For a lot of corpora that is fine. When it is not, rank the deduped parents by their best child score, or by how many children hit, before you apply max_parents. The strong signal is sitting in the hit counts, and the naive loop just does not read it.
Metadata filters have to live on the children. In any multi-user system you must filter retrieval by fields such as source, tenant, or date. Those filters run against the vector index, which only holds children, so copy the filterable fields onto every child at index time. A filter that only exists on the parent cannot touch a search that only sees children. This trips up anyone who added metadata filtering after the fact and put the fields in the wrong store.
Reranking gets awkward. A cross-encoder reranker scores a (query, passage) pair, and like any other model it has an input limit. Rerank the small children, before you expand to parents: they fit comfortably, and the scores reflect the precise text that matched. Reranking full parents means feeding long, multi-topic passages to a model whose input window may not even hold them. The score you get back is muddier for the same reason the big chunk was a bad search unit to begin with.
Related variations
Parent document retrieval is one point on a spectrum. Two neighbors are worth knowing:
- Sentence-window retrieval takes it to the extreme on the search side. You embed individual sentences, then return the sentence plus a window of neighbors around it.
- Auto-merging goes hierarchical. You chunk into a tree, and when enough sibling leaves match, the retriever merges up to their shared parent automatically. The granularity you return then adapts to how concentrated the hits are.
All three share one idea: the unit you search is smaller than the unit you send, and you keep a link between them. Which one fits depends on how structured your documents are and how much the answer depends on surrounding context.
When I reach for it
I lean on this pattern whenever documents are long and the answer depends on context around the match: technical docs, incident write-ups, anything with tables. On ARGRAG, the multimodal RAG system I built for the Argusa AI Challenge, the corpus had exactly that shape. It held enterprise documents where a matched row means nothing without the header and the section it sits under. Searching on small children kept retrieval sharp against specific phrasing. Answering with parents meant the model saw the row and the frame around it.
The mental shift is small, but it clears up a question that has no good answer otherwise. Stop asking “what is the right chunk size” as if one number has to serve search and generation at once. Size the child for the embedding model, size the parent for the language model, and keep a parent_id between them.
Then measure whether it worked, with a proper way of measuring retrieval quality rather than reading a few answers and nodding. The only way to know your two sizes are right is to watch recall and answer quality move together instead of trading off.
Diagrams by M. Hassan Ahmed, released under CC0. No external image was used for this post; the figures are original work by the author.