{"componentChunkName":"component---src-templates-blog-post-js","path":"/blog/2026-09-09-parent-document-retrieval-rag/","result":{"data":{"site":{"siteMetadata":{"title":"M.Hassan Ahmed","author":"Hassan11196"}},"markdownRemark":{"id":"9d76bfa5-6fd9-5fdd-9240-9069d7a08f7b","excerpt":"Chunk size is the setting people tune last and regret first. Make chunks big and retrieval gets vague. A 2,000-token chunk embeds down to one vector that…","html":"<p>Chunk size is the setting people tune last and regret first. Make chunks big and retrieval gets vague. A 2,000-token chunk embeds down to one vector that averages five different subtopics, so it matches everything weakly and nothing sharply. Make chunks small and search gets precise, but now the model reads a 200-token fragment that starts mid-sentence and drops the definition it needed two paragraphs up. A single number trades answer quality against retrieval quality, and no value of that number wins both.</p>\n<p>Parent document retrieval stops treating those as the same number: you search over small chunks and answer with big ones. This post is for engineers who already have a working <a href=\"/project/archi/\">RAG pipeline</a> and see retrieval find the right document but hand the model a passage too narrow to answer from. I walk through the two-index setup, a framework-free implementation, how to size the chunks, and the specific ways it breaks in practice.</p>\n<h2>The two jobs a chunk has to do</h2>\n<p>A chunk in a RAG (retrieval-augmented generation) system does two unrelated jobs:</p>\n<ul>\n<li><strong>It is the unit you <em>search</em>:</strong> the thing that gets embedded and compared against the query.</li>\n<li><strong>It is the unit you <em>read</em>:</strong> the text that lands in the model’s context.</li>\n</ul>\n<p><a href=\"/blog/2026-07-06-chunking-strategies-for-rag/\">Chunking strategies</a> usually try to satisfy both with one split, and that is the tension. Search wants a chunk small and focused enough that its embedding means one thing. Generation wants a chunk large enough to stand on its own, with the surrounding sentences, the table header the row belongs to, and the “it” that refers to something named earlier.</p>\n<p>The fix is to stop forcing one chunk to do both. Split each document twice. Big <strong>parent</strong> chunks are what the model reads. Small <strong>child</strong> chunks, carved out of each parent, are what you embed and search. When a child matches, you do not return the child; you return the parent it came from.</p>\n<p><img src=\"/303ef938c41457c773a557ab9dfd0d76/index-build.svg\" alt=\"Indexing builds two stores from one document. A source document is split into parent chunks of roughly 1,500 to 2,000 tokens; each parent is split again into children of 250 to 400 tokens, and every child carries its parent&#x27;s id. Only the child chunks are embedded into a vector index. The parent text is stored separately in a doc store keyed by parent id, not embedded.\"></p>\n<p>Only the children go into the vector index. The parents go into a plain key-value store keyed by an id that every child carries as metadata. That store can be a dict in memory, a Postgres table, a Redis hash, or whatever you already run. So you have two stores for one document, and the link between them is that <code class=\"language-text\">parent_id</code>.</p>\n<h2>Building the two indexes</h2>\n<p>Here is the index side without a framework, so the moving parts are visible. It has two splitters, a docstore, and a vector index where the payload on each vector is the child’s <code class=\"language-text\">parent_id</code>. Watch the last two lines: only the children are embedded.</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\"><span class=\"token keyword\">from</span> dataclasses <span class=\"token keyword\">import</span> dataclass\n\n<span class=\"token decorator annotation punctuation\">@dataclass</span>\n<span class=\"token keyword\">class</span> <span class=\"token class-name\">Child</span><span class=\"token punctuation\">:</span>\n    text<span class=\"token punctuation\">:</span> <span class=\"token builtin\">str</span>\n    parent_id<span class=\"token punctuation\">:</span> <span class=\"token builtin\">str</span>\n\n<span class=\"token keyword\">def</span> <span class=\"token function\">split</span><span class=\"token punctuation\">(</span>text<span class=\"token punctuation\">:</span> <span class=\"token builtin\">str</span><span class=\"token punctuation\">,</span> size<span class=\"token punctuation\">:</span> <span class=\"token builtin\">int</span><span class=\"token punctuation\">,</span> overlap<span class=\"token punctuation\">:</span> <span class=\"token builtin\">int</span><span class=\"token punctuation\">)</span> <span class=\"token operator\">-</span><span class=\"token operator\">></span> <span class=\"token builtin\">list</span><span class=\"token punctuation\">[</span><span class=\"token builtin\">str</span><span class=\"token punctuation\">]</span><span class=\"token punctuation\">:</span>\n    <span class=\"token comment\"># token- or character-based; use your tokenizer in production</span>\n    step <span class=\"token operator\">=</span> size <span class=\"token operator\">-</span> overlap\n    <span class=\"token keyword\">return</span> <span class=\"token punctuation\">[</span>text<span class=\"token punctuation\">[</span>i<span class=\"token punctuation\">:</span>i <span class=\"token operator\">+</span> size<span class=\"token punctuation\">]</span> <span class=\"token keyword\">for</span> i <span class=\"token keyword\">in</span> <span class=\"token builtin\">range</span><span class=\"token punctuation\">(</span><span class=\"token number\">0</span><span class=\"token punctuation\">,</span> <span class=\"token builtin\">len</span><span class=\"token punctuation\">(</span>text<span class=\"token punctuation\">)</span><span class=\"token punctuation\">,</span> step<span class=\"token punctuation\">)</span><span class=\"token punctuation\">]</span>\n\nparents<span class=\"token punctuation\">:</span> <span class=\"token builtin\">dict</span><span class=\"token punctuation\">[</span><span class=\"token builtin\">str</span><span class=\"token punctuation\">,</span> <span class=\"token builtin\">str</span><span class=\"token punctuation\">]</span> <span class=\"token operator\">=</span> <span class=\"token punctuation\">{</span><span class=\"token punctuation\">}</span>          <span class=\"token comment\"># parent_id -> full parent text</span>\nchildren<span class=\"token punctuation\">:</span> <span class=\"token builtin\">list</span><span class=\"token punctuation\">[</span>Child<span class=\"token punctuation\">]</span> <span class=\"token operator\">=</span> <span class=\"token punctuation\">[</span><span class=\"token punctuation\">]</span>\n\n<span class=\"token keyword\">for</span> i<span class=\"token punctuation\">,</span> parent <span class=\"token keyword\">in</span> <span class=\"token builtin\">enumerate</span><span class=\"token punctuation\">(</span>split<span class=\"token punctuation\">(</span>document<span class=\"token punctuation\">,</span> size<span class=\"token operator\">=</span><span class=\"token number\">1800</span><span class=\"token punctuation\">,</span> overlap<span class=\"token operator\">=</span><span class=\"token number\">200</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n    pid <span class=\"token operator\">=</span> <span class=\"token string-interpolation\"><span class=\"token string\">f\"</span><span class=\"token interpolation\"><span class=\"token punctuation\">{</span>doc_id<span class=\"token punctuation\">}</span></span><span class=\"token string\">:p</span><span class=\"token interpolation\"><span class=\"token punctuation\">{</span>i<span class=\"token punctuation\">}</span></span><span class=\"token string\">\"</span></span>\n    parents<span class=\"token punctuation\">[</span>pid<span class=\"token punctuation\">]</span> <span class=\"token operator\">=</span> parent\n    <span class=\"token keyword\">for</span> chunk <span class=\"token keyword\">in</span> split<span class=\"token punctuation\">(</span>parent<span class=\"token punctuation\">,</span> size<span class=\"token operator\">=</span><span class=\"token number\">350</span><span class=\"token punctuation\">,</span> overlap<span class=\"token operator\">=</span><span class=\"token number\">50</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n        children<span class=\"token punctuation\">.</span>append<span class=\"token punctuation\">(</span>Child<span class=\"token punctuation\">(</span>text<span class=\"token operator\">=</span>chunk<span class=\"token punctuation\">,</span> parent_id<span class=\"token operator\">=</span>pid<span class=\"token punctuation\">)</span><span class=\"token punctuation\">)</span>\n\n<span class=\"token comment\"># embed ONLY the children</span>\nvectors <span class=\"token operator\">=</span> embed<span class=\"token punctuation\">(</span><span class=\"token punctuation\">[</span>c<span class=\"token punctuation\">.</span>text <span class=\"token keyword\">for</span> c <span class=\"token keyword\">in</span> children<span class=\"token punctuation\">]</span><span class=\"token punctuation\">)</span>\nindex<span class=\"token punctuation\">.</span>add<span class=\"token punctuation\">(</span>vectors<span class=\"token punctuation\">,</span> payloads<span class=\"token operator\">=</span>children<span class=\"token punctuation\">)</span>   <span class=\"token comment\"># your vector store's add()</span></code></pre></div>\n<p>The parent text is never embedded, and that is the point people miss on the first read. The doc store is dumb storage, and the vector index only ever sees the small chunks.</p>\n<p>If you are on LangChain, its <a href=\"https://python.langchain.com/docs/how_to/parent_document_retriever/\"><code class=\"language-text\">ParentDocumentRetriever</code></a> wires this up for you. LlamaIndex ships the same idea as its <a href=\"https://docs.llamaindex.ai/en/stable/examples/retrievers/auto_merging_retriever/\">auto-merging retriever</a>. The mechanics underneath are the same, and knowing them is what lets you debug the framework when it does something surprising.</p>\n<h2>At query time, children collapse back to parents</h2>\n<p>Query time is where the two indexes come together. You embed the query and search the child index for the top-k children. Then you map each hit back to its parent and dedupe. Several children from the same parent will often match one query, and you do not want to send that parent three times.</p>\n<p><img src=\"/33ef387de6e119591212940a2a1b60e7/query-expand.svg\" alt=\"At query time, five child hits collapse back to three parents. The top-5 child chunks match, but two of them point to parent A and two to parent B, so grouping by parent id leaves three unique parents: A, B, and C. Those parents are fetched whole from the doc store and concatenated into a context of three coherent passages, which goes to the LLM.\"></p>\n<p>The function below walks the child hits in score order, skips any parent it has already added, and stops once it has <code class=\"language-text\">max_parents</code> parents:</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\"><span class=\"token keyword\">def</span> <span class=\"token function\">retrieve</span><span class=\"token punctuation\">(</span>query<span class=\"token punctuation\">:</span> <span class=\"token builtin\">str</span><span class=\"token punctuation\">,</span> k_children<span class=\"token punctuation\">:</span> <span class=\"token builtin\">int</span> <span class=\"token operator\">=</span> <span class=\"token number\">8</span><span class=\"token punctuation\">,</span> max_parents<span class=\"token punctuation\">:</span> <span class=\"token builtin\">int</span> <span class=\"token operator\">=</span> <span class=\"token number\">3</span><span class=\"token punctuation\">)</span> <span class=\"token operator\">-</span><span class=\"token operator\">></span> <span class=\"token builtin\">list</span><span class=\"token punctuation\">[</span><span class=\"token builtin\">str</span><span class=\"token punctuation\">]</span><span class=\"token punctuation\">:</span>\n    hits <span class=\"token operator\">=</span> index<span class=\"token punctuation\">.</span>search<span class=\"token punctuation\">(</span>embed<span class=\"token punctuation\">(</span><span class=\"token punctuation\">[</span>query<span class=\"token punctuation\">]</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">[</span><span class=\"token number\">0</span><span class=\"token punctuation\">]</span><span class=\"token punctuation\">,</span> top_k<span class=\"token operator\">=</span>k_children<span class=\"token punctuation\">)</span>  <span class=\"token comment\"># sorted by score</span>\n    seen<span class=\"token punctuation\">,</span> out <span class=\"token operator\">=</span> <span class=\"token builtin\">set</span><span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">,</span> <span class=\"token punctuation\">[</span><span class=\"token punctuation\">]</span>\n    <span class=\"token keyword\">for</span> hit <span class=\"token keyword\">in</span> hits<span class=\"token punctuation\">:</span>\n        pid <span class=\"token operator\">=</span> hit<span class=\"token punctuation\">.</span>payload<span class=\"token punctuation\">.</span>parent_id\n        <span class=\"token keyword\">if</span> pid <span class=\"token keyword\">in</span> seen<span class=\"token punctuation\">:</span>\n            <span class=\"token keyword\">continue</span>\n        seen<span class=\"token punctuation\">.</span>add<span class=\"token punctuation\">(</span>pid<span class=\"token punctuation\">)</span>\n        out<span class=\"token punctuation\">.</span>append<span class=\"token punctuation\">(</span>parents<span class=\"token punctuation\">[</span>pid<span class=\"token punctuation\">]</span><span class=\"token punctuation\">)</span>          <span class=\"token comment\"># fetch the whole parent</span>\n        <span class=\"token keyword\">if</span> <span class=\"token builtin\">len</span><span class=\"token punctuation\">(</span>out<span class=\"token punctuation\">)</span> <span class=\"token operator\">>=</span> max_parents<span class=\"token punctuation\">:</span>\n            <span class=\"token keyword\">break</span>\n    <span class=\"token keyword\">return</span> out</code></pre></div>\n<p>Two knobs matter here, and they are not the same knob:</p>\n<ul>\n<li><strong><code class=\"language-text\">k_children</code> controls how wide you search.</strong> Pull more children than you need, because several will collapse into the same parent.</li>\n<li><strong><code class=\"language-text\">max_parents</code> controls how much text you actually send.</strong> Search over eight children, come away with three parents, and keep the context focused.</li>\n</ul>\n<p>If you set <code class=\"language-text\">k_children</code> too close to <code class=\"language-text\">max_parents</code>, you will occasionally get two parents back when you asked for three. That happens when, for example, the eighth-ranked child shared a parent with the second.</p>\n<h2>Sizing parents and children</h2>\n<p>The numbers in the code, 1,800-token parents and 350-token children, are a reasonable starting point, not a law. I reason about the two sizes differently.</p>\n<p><strong>Child size is bounded by your embedding model.</strong> Most retrieval models were trained on short passages, and their quality degrades well before their advertised token limit. Children of 250 to 400 tokens stay inside the range the model actually handles well. This is the same concern behind <a href=\"/blog/2026-07-28-choosing-embedding-model-rag/\">picking an embedding model</a> in the first place: the child is what gets embedded, so the child is what the model’s limits apply to.</p>\n<p><strong>Parent size is bounded by two things pulling against each other.</strong> A parent must be big enough that the passage is self-contained, and small enough that three of them do not blow your context budget or bury the answer.</p>\n<p>There is a real reason not to just make parents enormous and send one. Models attend unevenly across a long context and reliably lose information stranded in the middle of it, an effect documented in <a href=\"https://arxiv.org/abs/2307.03172\">Liu et al.’s “Lost in the Middle”</a>. Three tight parents beat one sprawling one, and both beat twenty fragments.</p>\n<p>If your documents have real structure (Markdown headings, sections, a <code class=\"language-text\">##</code> every few paragraphs), split parents on those boundaries instead of a fixed token count. That way a parent is a section rather than an arbitrary window.</p>\n<h2>Where it breaks</h2>\n<p>The happy path is a hundred lines. The failures all live in the seams between the two stores.</p>\n<p><strong>The two stores drift apart.</strong> The vector index and the doc store are separate systems that have to agree on <code class=\"language-text\">parent_id</code>. Re-index the documents, change the id scheme, or partially reload one store and not the other, and you get children whose <code class=\"language-text\">parent_id</code> points at a parent that no longer exists. The retrieve loop then throws a <code class=\"language-text\">KeyError</code> mid-request, or worse, silently skips the hit.</p>\n<p>The fix is to treat the two writes as one operation. Never publish new children pointing at parents you have not written yet, and when you delete a document, delete it from both stores. This is the same <a href=\"/blog/2026-08-13-zero-downtime-reindexing-opensearch/\">zero-downtime reindexing</a> discipline, spread across two stores instead of one.</p>\n<p><strong>Dedup throws away a relevance signal.</strong> A parent that matched on three of its children is probably more relevant than one that matched on one. The simple loop above ignores that: it keeps the first parent it sees and moves on. For a lot of corpora that is fine. When it is not, rank the deduped parents by their best child score, or by how many children hit, before you apply <code class=\"language-text\">max_parents</code>. The strong signal is sitting in the hit counts, and the naive loop just does not read it.</p>\n<p><strong>Metadata filters have to live on the children.</strong> In any multi-user system you must filter retrieval by fields such as source, tenant, or date. Those filters run against the vector index, which only holds children, so copy the filterable fields onto every child at index time. A filter that only exists on the parent cannot touch a search that only sees children. This trips up anyone who added <a href=\"/blog/2026-09-01-metadata-filtering-rag-vector-search/\">metadata filtering</a> after the fact and put the fields in the wrong store.</p>\n<p><strong>Reranking gets awkward.</strong> A <a href=\"/blog/2026-07-08-cross-encoder-reranking-rag/\">cross-encoder reranker</a> scores a <code class=\"language-text\">(query, passage)</code> pair, and like any other model it has an input limit. Rerank the small children, before you expand to parents: they fit comfortably, and the scores reflect the precise text that matched. Reranking full parents means feeding long, multi-topic passages to a model whose input window may not even hold them. The score you get back is muddier for the same reason the big chunk was a bad search unit to begin with.</p>\n<h2>Related variations</h2>\n<p>Parent document retrieval is one point on a spectrum. Two neighbors are worth knowing:</p>\n<ul>\n<li><strong>Sentence-window retrieval</strong> takes it to the extreme on the search side. You embed individual sentences, then return the sentence plus a window of neighbors around it.</li>\n<li><strong>Auto-merging</strong> goes hierarchical. You chunk into a tree, and when enough sibling leaves match, the retriever merges up to their shared parent automatically. The granularity you return then adapts to how concentrated the hits are.</li>\n</ul>\n<p>All three share one idea: the unit you search is smaller than the unit you send, and you keep a link between them. Which one fits depends on how structured your documents are and how much the answer depends on surrounding context.</p>\n<h2>When I reach for it</h2>\n<p>I lean on this pattern whenever documents are long and the answer depends on context around the match: technical docs, incident write-ups, anything with tables. On <a href=\"/project/argusa-ai-challenge-2025/\">ARGRAG</a>, the multimodal RAG system I built for the Argusa AI Challenge, the corpus had exactly that shape. It held enterprise documents where a matched row means nothing without the header and the section it sits under. Searching on small children kept retrieval sharp against specific phrasing. Answering with parents meant the model saw the row <em>and</em> the frame around it.</p>\n<p>The mental shift is small, but it clears up a question that has no good answer otherwise. Stop asking “what is the right chunk size” as if one number has to serve search and generation at once. Size the child for the embedding model, size the parent for the language model, and keep a <code class=\"language-text\">parent_id</code> between them.</p>\n<p>Then measure whether it worked, with a proper way of <a href=\"/blog/2026-07-04-measuring-rag-retrieval-quality/\">measuring retrieval quality</a> rather than reading a few answers and nodding. The only way to know your two sizes are right is to watch recall and answer quality move together instead of trading off.</p>\n<hr>\n<p><em>Diagrams by M. Hassan Ahmed, released under CC0. No external image was used for this post; the figures are original work by the author.</em></p>","frontmatter":{"title":"Parent Document Retrieval: Search Small, Answer Big","date":"2026-09-09T00:00:00.000Z","description":"Small chunks search precisely but read like fragments. Parent document retrieval matches on child chunks, then feeds the whole parent to the model.","thumbnail":{"childImageSharp":{"fluid":{"base64":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAABQAAAALCAIAAADwazoUAAAACXBIWXMAAAsSAAALEgHS3X78AAACE0lEQVQozy3R2W6jMBQGYK7SLBgIYDYvmMVgwEmhkzYkhaTppKmq6eRmpF5W8/5PMW410qff25Fl+Wh1iZWqQMehez7cbTfV09Adh1blYy9Phx/7h+Z0uBv3a7W538oyD+tC1WOV2gTQiU4BLHI5FqsjEzvRnvjq8O3I109JNRTrJy6P9d2Zif3C4RMQTwC7MZhmwszyctPLdSdRDJjpdrxYkrmF1WTpp3aQLb0E2AjY2HSpBRmwKXCIblMN0jbKege3LulstLYiaXrlwsJzk5gOlu397WaU3Q4ltY8LD5dxmssmpyxTNZodVVHSobSL0hbi2kW15fEpiKYAGQ6Oiw0qRlQ8BtnOT3slyvsouw9IqVtIK3giKoFjQVjFsjLJCkwT20sMVz0VM7GlzSVuzkxemHyNm5+oOLjxvY+5bkXatqPXc/Z+ys8j/3hLP9+RrLNS7op6gwm9PtO/V/z5+7+PV/xxQX9eyK1kMxBqufrmVVPVZSmqbp1fn+O2IX6I/ABREg4P6tL858C/jHzYlmNfHvpClGxmBJqNVpBtzahdog4E6zmUps8DIkyYzg01Dqj5heo3JSwvkXgJ+MkivYv4wvA0KxCQdjDuVNpIKgbMbxbwZuGZdhBnFc5byr+gpAnjKoxFQEVEkhnwtLkZ6UuiGqt/9Rap5cwIp7o3Bb7KBXCNpe945Bt1/VglsLw5gOr0H3xmWlqUGuSYAAAAAElFTkSuQmCC","aspectRatio":1.899441340782123,"src":"/static/73688f50ebb030dd3cc9f68a43db401a/40a76/hero.png","srcSet":"/static/73688f50ebb030dd3cc9f68a43db401a/c972b/hero.png 340w,\n/static/73688f50ebb030dd3cc9f68a43db401a/27625/hero.png 680w,\n/static/73688f50ebb030dd3cc9f68a43db401a/40a76/hero.png 1360w,\n/static/73688f50ebb030dd3cc9f68a43db401a/ed396/hero.png 2000w","sizes":"(max-width: 1360px) 100vw, 1360px"}}}}}},"pageContext":{"slug":"/2026-09-09-parent-document-retrieval-rag/","previous":"blog/2026-09-10-maximal-marginal-relevance-rag/","next":"blog/2026-09-13-agentic-rag-when-to-retrieve/"}},"staticQueryHashes":["32046230"]}