{"componentChunkName":"component---src-templates-blog-post-js","path":"/blog/2026-09-07-parsing-pdfs-for-rag/","result":{"data":{"site":{"siteMetadata":{"title":"M.Hassan Ahmed","author":"Hassan11196"}},"markdownRemark":{"id":"0cff266b-27bb-5932-b32a-4cd29e512c76","excerpt":"Most RAG (retrieval-augmented generation) demos start with a folder of clean Markdown. Most real RAG projects start with a shared drive full of PDFs nobody has…","html":"<p>Most RAG (retrieval-augmented generation) demos start with a folder of clean Markdown. Most real RAG projects start with a shared drive full of PDFs nobody has looked at in three years. Between the two sits the parsing step, and that is where retrieval quality quietly leaks away. You can pick the best embedding model and the sharpest reranker (the model that re-scores retrieved chunks). But if the parser turned a financial table into a run-on sentence of numbers, the retriever will happily fetch that chunk and the model will happily answer with the wrong figure.</p>\n<p>I hit this on two fronts. <a href=\"/project/archi/\">Archi</a> is the RAG copilot I led for CMS computing operations. Its ingestion crawler pulls from internal portals, ticketing systems, and logbooks, and a good share of that source material is PDF. <a href=\"/project/argusa-ai-challenge-2025/\">ARGRAG</a> is the enterprise-document system that took second at the Argusa AI Challenge, and there the whole problem <em>was</em> heterogeneous documents: contracts, reports, and scanned forms, all in one corpus. Both taught me the same lesson. The parser is not plumbing you set up once and forget. It decides what your model can and cannot know.</p>\n<p>This post is for engineers building a RAG pipeline who have moved past the “hello world” of <code class=\"language-text\">.txt</code> files and now have to ingest real PDFs. It covers why PDFs are hard, how to tell a born-digital file from a scan, and how to extract tables without destroying them. It also covers when OCR is worth the cost, and the failure modes that will embarrass you in a demo.</p>\n<h2>Why a PDF is a drawing, not a document</h2>\n<p>The common mistake is to assume a PDF contains text the way an HTML page does. It does not. A <a href=\"https://en.wikipedia.org/wiki/PDF\">PDF</a> describes how to <em>paint</em> a page: place this glyph (one drawn character shape) at these coordinates, draw a line here, show an image there. Often there is no notion of a paragraph, a column, a table cell, or even reading order. A two-column academic paper stores its glyphs in whatever order the layout engine emitted them, and that is frequently not the order a human reads them in.</p>\n<p>So “extracting text from a PDF” really means “reconstructing logical structure from a pile of positioned glyphs.” Some PDFs make that easy. They were exported from Word or LaTeX and carry a clean text layer, meaning the file stores the actual characters alongside their positions. Others make it impossible. They are a photo of a page wrapped in a PDF container, with no text at all. Treating both kinds the same way is the first mistake.</p>\n<p>The rest of this post follows the pipeline below: triage each page, extract its text and tables directly or send it to OCR, then normalize, chunk, and embed.</p>\n<p><img src=\"/0b282edb43fc0f154ffc07f5bdbef0fb/pdf-ingestion-pipeline.svg\" alt=\"A PDF ingestion pipeline for RAG. A PDF flows into a triage diamond asking whether it has a text layer. If yes, it goes to text and table extraction with PyMuPDF and pdfplumber or TableFormer. If no, it goes to OCR on each page image with Tesseract or a vision-language model, marked slow and error-prone. Both branches feed a normalize step for reading order, markdown, and metadata, then chunking, then embedding and storage. A caption reads: route once, on evidence; a page with almost no extractable text is scanned and needs OCR, everything else has a text layer to read directly.\"></p>\n<h2>Step one: triage each page as born-digital or scanned</h2>\n<p>Before you extract anything, decide what you are holding. A born-digital PDF, one generated by software rather than scanned from paper, has a real text layer you can read directly and cheaply. A scanned PDF is just images. The only way to get text out of it is optical character recognition (OCR), which is slower and lossier. Getting this wrong hurts in both directions:</p>\n<ul>\n<li>Running OCR on a born-digital file wastes compute and often produces <em>worse</em> text than the layer that was already there.</li>\n<li>Running plain extraction on a scan returns an empty string and leaves a silent gap in your index.</li>\n</ul>\n<p>The cheap, reliable test is to extract text and count how much you got per page. <a href=\"https://pymupdf.readthedocs.io/\">PyMuPDF</a> (imported as <code class=\"language-text\">fitz</code>) is my default for this because it is fast and gives you page-level access. The function below counts the pages that yield almost no text and returns them as a fraction of the document.</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\"><span class=\"token keyword\">import</span> fitz  <span class=\"token comment\"># PyMuPDF</span>\n\n<span class=\"token keyword\">def</span> <span class=\"token function\">looks_scanned</span><span class=\"token punctuation\">(</span>pdf_path<span class=\"token punctuation\">,</span> min_chars_per_page<span class=\"token operator\">=</span><span class=\"token number\">50</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n    doc <span class=\"token operator\">=</span> fitz<span class=\"token punctuation\">.</span><span class=\"token builtin\">open</span><span class=\"token punctuation\">(</span>pdf_path<span class=\"token punctuation\">)</span>\n    scanned_pages <span class=\"token operator\">=</span> <span class=\"token number\">0</span>\n    <span class=\"token keyword\">for</span> page <span class=\"token keyword\">in</span> doc<span class=\"token punctuation\">:</span>\n        text <span class=\"token operator\">=</span> page<span class=\"token punctuation\">.</span>get_text<span class=\"token punctuation\">(</span><span class=\"token string\">\"text\"</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">.</span>strip<span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span>\n        <span class=\"token keyword\">if</span> <span class=\"token builtin\">len</span><span class=\"token punctuation\">(</span>text<span class=\"token punctuation\">)</span> <span class=\"token operator\">&lt;</span> min_chars_per_page<span class=\"token punctuation\">:</span>\n            scanned_pages <span class=\"token operator\">+=</span> <span class=\"token number\">1</span>\n    <span class=\"token keyword\">return</span> scanned_pages <span class=\"token operator\">/</span> <span class=\"token builtin\">max</span><span class=\"token punctuation\">(</span><span class=\"token builtin\">len</span><span class=\"token punctuation\">(</span>doc<span class=\"token punctuation\">)</span><span class=\"token punctuation\">,</span> <span class=\"token number\">1</span><span class=\"token punctuation\">)</span>  <span class=\"token comment\"># fraction that look scanned</span></code></pre></div>\n<p>If most pages come back nearly empty, the file is scanned and needs the OCR branch. Make the decision per page, though, not per document. Mixed files are common in the wild, such as a born-digital report with three scanned appendix pages stapled on. If you decide “scanned or not” once for the whole file, you either OCR everything needlessly or drop the scanned appendix. Deciding per page costs nothing and handles both cases.</p>\n<h2>Step two: extract text in reading order</h2>\n<p>For born-digital pages, <code class=\"language-text\">page.get_text(&quot;text&quot;)</code> gets you most of the way, and PyMuPDF’s default already does a decent job of following reading flow. It falls down on multi-column layouts and on anything with sidebars, footnotes, or callout boxes. There, the glyphs for the left and right columns can interleave, and you get sentences spliced together across the gutter (the blank strip between columns).</p>\n<p>When order matters, switch to the structured output. <code class=\"language-text\">page.get_text(&quot;dict&quot;)</code> returns blocks, lines, and spans, each with its bounding box (the rectangle it occupies on the page). You can then sort and group them by position yourself. The snippet below handles a simple two-column page: it splits blocks at the page’s horizontal midpoint and reads each column top to bottom.</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\">blocks <span class=\"token operator\">=</span> page<span class=\"token punctuation\">.</span>get_text<span class=\"token punctuation\">(</span><span class=\"token string\">\"dict\"</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">[</span><span class=\"token string\">\"blocks\"</span><span class=\"token punctuation\">]</span>\n<span class=\"token comment\"># Each block has a \"bbox\" (x0, y0, x1, y1). For a two-column page,</span>\n<span class=\"token comment\"># split on the page midpoint, then read each column top-to-bottom.</span>\nmid <span class=\"token operator\">=</span> page<span class=\"token punctuation\">.</span>rect<span class=\"token punctuation\">.</span>width <span class=\"token operator\">/</span> <span class=\"token number\">2</span>\nleft  <span class=\"token operator\">=</span> <span class=\"token builtin\">sorted</span><span class=\"token punctuation\">(</span><span class=\"token punctuation\">(</span>b <span class=\"token keyword\">for</span> b <span class=\"token keyword\">in</span> blocks <span class=\"token keyword\">if</span> b<span class=\"token punctuation\">[</span><span class=\"token string\">\"bbox\"</span><span class=\"token punctuation\">]</span><span class=\"token punctuation\">[</span><span class=\"token number\">0</span><span class=\"token punctuation\">]</span> <span class=\"token operator\">&lt;</span> mid<span class=\"token punctuation\">)</span><span class=\"token punctuation\">,</span> key<span class=\"token operator\">=</span><span class=\"token keyword\">lambda</span> b<span class=\"token punctuation\">:</span> b<span class=\"token punctuation\">[</span><span class=\"token string\">\"bbox\"</span><span class=\"token punctuation\">]</span><span class=\"token punctuation\">[</span><span class=\"token number\">1</span><span class=\"token punctuation\">]</span><span class=\"token punctuation\">)</span>\nright <span class=\"token operator\">=</span> <span class=\"token builtin\">sorted</span><span class=\"token punctuation\">(</span><span class=\"token punctuation\">(</span>b <span class=\"token keyword\">for</span> b <span class=\"token keyword\">in</span> blocks <span class=\"token keyword\">if</span> b<span class=\"token punctuation\">[</span><span class=\"token string\">\"bbox\"</span><span class=\"token punctuation\">]</span><span class=\"token punctuation\">[</span><span class=\"token number\">0</span><span class=\"token punctuation\">]</span> <span class=\"token operator\">>=</span> mid<span class=\"token punctuation\">)</span><span class=\"token punctuation\">,</span> key<span class=\"token operator\">=</span><span class=\"token keyword\">lambda</span> b<span class=\"token punctuation\">:</span> b<span class=\"token punctuation\">[</span><span class=\"token string\">\"bbox\"</span><span class=\"token punctuation\">]</span><span class=\"token punctuation\">[</span><span class=\"token number\">1</span><span class=\"token punctuation\">]</span><span class=\"token punctuation\">)</span>\nordered <span class=\"token operator\">=</span> left <span class=\"token operator\">+</span> right</code></pre></div>\n<p>This is crude, and it works for simple two-column layouts. For genuinely messy pages, hand-rolled geometry logic stops paying off. You want a layout-detection model instead, one trained to recognize columns, headers, and tables. That is where tools like <a href=\"https://github.com/docling-project/docling\">Docling</a> (IBM’s open-source document converter, Apache 2.0) or <a href=\"https://github.com/Unstructured-IO/unstructured\">unstructured</a> earn their keep. They run a layout model, recover reading order, and emit clean Markdown or a structured list of elements. The tradeoff is speed and a heavier dependency. For a large corpus I profile both, and I usually end up with a fast path for clean files and the heavier model only for the files that need it.</p>\n<h2>Step three: extract tables without flattening them</h2>\n<p>Tables are the single biggest reason a RAG answer comes out confidently wrong. A table’s meaning lives in the <em>relationships</em> between its rows and columns, and naive text extraction throws those away. Read a table left to right and top to bottom as a stream of glyphs, and you get a line like <code class=\"language-text\">Site Jobs Fail % CERN 4200 1.2 FNAL 3100 3.8</code>. Embed that line and it will still be retrieved for a question about failure rates. The model will then answer with a number pulled from the wrong cell.</p>\n<p><img src=\"/a0f674105511d1b2753c5274d62f6003/table-flattening.svg\" alt=\"Why tables break naive PDF extraction. On the left, a small table with columns Site, Jobs, and Fail percent, and rows for CERN, FNAL, and DESY. A naive extraction flattens it into one line of text where the numbers run together, and it becomes impossible to tell whether DESY&#x27;s failure rate is 0.5 or 900. A table-aware extractor instead produces a Markdown table that preserves the row and column structure. The caption notes the row and column relationship is the information, and that an embedding of the flattened line still retrieves but answers wrong, a failure that stays silent until a user checks the number.\"></p>\n<p>For born-digital tables with visible ruling lines (the drawn borders between cells), <a href=\"https://github.com/jsvine/pdfplumber\">pdfplumber</a> is precise, because it works from the actual line and rectangle objects in the PDF. The code below pulls every table from every page and rebuilds each one as a Markdown table, treating the first row as the header.</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\"><span class=\"token keyword\">import</span> pdfplumber\n\n<span class=\"token keyword\">with</span> pdfplumber<span class=\"token punctuation\">.</span><span class=\"token builtin\">open</span><span class=\"token punctuation\">(</span>pdf_path<span class=\"token punctuation\">)</span> <span class=\"token keyword\">as</span> pdf<span class=\"token punctuation\">:</span>\n    <span class=\"token keyword\">for</span> page <span class=\"token keyword\">in</span> pdf<span class=\"token punctuation\">.</span>pages<span class=\"token punctuation\">:</span>\n        <span class=\"token keyword\">for</span> table <span class=\"token keyword\">in</span> page<span class=\"token punctuation\">.</span>extract_tables<span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n            <span class=\"token comment\"># table is a list of rows, each a list of cell strings</span>\n            header<span class=\"token punctuation\">,</span> <span class=\"token operator\">*</span>rows <span class=\"token operator\">=</span> table\n            md <span class=\"token operator\">=</span> <span class=\"token string\">\"| \"</span> <span class=\"token operator\">+</span> <span class=\"token string\">\" | \"</span><span class=\"token punctuation\">.</span>join<span class=\"token punctuation\">(</span>h <span class=\"token keyword\">or</span> <span class=\"token string\">\"\"</span> <span class=\"token keyword\">for</span> h <span class=\"token keyword\">in</span> header<span class=\"token punctuation\">)</span> <span class=\"token operator\">+</span> <span class=\"token string\">\" |\\n\"</span>\n            md <span class=\"token operator\">+=</span> <span class=\"token string\">\"|\"</span> <span class=\"token operator\">+</span> <span class=\"token string\">\"|\"</span><span class=\"token punctuation\">.</span>join<span class=\"token punctuation\">(</span><span class=\"token string\">\"---\"</span> <span class=\"token keyword\">for</span> _ <span class=\"token keyword\">in</span> header<span class=\"token punctuation\">)</span> <span class=\"token operator\">+</span> <span class=\"token string\">\"|\\n\"</span>\n            <span class=\"token keyword\">for</span> row <span class=\"token keyword\">in</span> rows<span class=\"token punctuation\">:</span>\n                md <span class=\"token operator\">+=</span> <span class=\"token string\">\"| \"</span> <span class=\"token operator\">+</span> <span class=\"token string\">\" | \"</span><span class=\"token punctuation\">.</span>join<span class=\"token punctuation\">(</span>c <span class=\"token keyword\">or</span> <span class=\"token string\">\"\"</span> <span class=\"token keyword\">for</span> c <span class=\"token keyword\">in</span> row<span class=\"token punctuation\">)</span> <span class=\"token operator\">+</span> <span class=\"token string\">\" |\\n\"</span></code></pre></div>\n<p>PyMuPDF now ships table detection too (<code class=\"language-text\">page.find_tables()</code>). Its implementation is <a href=\"https://pymupdf.readthedocs.io/en/latest/faq/index.html\">ported from pdfplumber</a>, so the behavior will feel familiar.</p>\n<p>The hard case is a table with no ruling lines, where only whitespace implies the grid. Line-based detection finds nothing there, and you fall back to one of two options:</p>\n<ul>\n<li><strong>Whitespace clustering</strong>, which is fragile.</li>\n<li><strong>A model-based table recognizer</strong> such as Docling’s TableFormer, which predicts the row and column structure from the rendered image instead of relying on drawn lines.</li>\n</ul>\n<p>Whatever produces the table, store it in the chunk as Markdown (or HTML), not as a flattened line. The pipe-delimited form keeps the header attached to each value, which is exactly the context the model needs to read a cell correctly. It also survives chunking better, which matters once a big table gets split. That ties back to how you <a href=\"/blog/2026-07-06-chunking-strategies-for-rag/\">chunk documents for RAG</a> in the first place.</p>\n<h2>Step four: OCR, only when you have to</h2>\n<p>If triage says a page is scanned, it has no text layer, so you have to generate one. <a href=\"https://tesseract-ocr.github.io/\">Tesseract</a> is the standard open-source OCR engine, and you call it from Python through <code class=\"language-text\">pytesseract</code>. The function below renders one page to an image and runs OCR on that image.</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\"><span class=\"token keyword\">import</span> fitz<span class=\"token punctuation\">,</span> pytesseract\n<span class=\"token keyword\">from</span> PIL <span class=\"token keyword\">import</span> Image\n<span class=\"token keyword\">import</span> io\n\n<span class=\"token keyword\">def</span> <span class=\"token function\">ocr_page</span><span class=\"token punctuation\">(</span>page<span class=\"token punctuation\">,</span> dpi<span class=\"token operator\">=</span><span class=\"token number\">300</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n    pix <span class=\"token operator\">=</span> page<span class=\"token punctuation\">.</span>get_pixmap<span class=\"token punctuation\">(</span>dpi<span class=\"token operator\">=</span>dpi<span class=\"token punctuation\">)</span>            <span class=\"token comment\"># render page to raster</span>\n    img <span class=\"token operator\">=</span> Image<span class=\"token punctuation\">.</span><span class=\"token builtin\">open</span><span class=\"token punctuation\">(</span>io<span class=\"token punctuation\">.</span>BytesIO<span class=\"token punctuation\">(</span>pix<span class=\"token punctuation\">.</span>tobytes<span class=\"token punctuation\">(</span><span class=\"token string\">\"png\"</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">)</span>\n    <span class=\"token keyword\">return</span> pytesseract<span class=\"token punctuation\">.</span>image_to_string<span class=\"token punctuation\">(</span>img<span class=\"token punctuation\">)</span>   <span class=\"token comment\"># OCR to text</span></code></pre></div>\n<p>Two things decide OCR quality more than anything else:</p>\n<ol>\n<li><strong>Resolution.</strong> 300 DPI (dots per inch) is the usual floor. Dropping below it to save time is the fastest way to get garbled output.</li>\n<li><strong>Preprocessing.</strong> Deskewing (straightening a tilted scan), denoising, and thresholding the image before OCR often matter more than any engine setting. Tesseract was trained on clean scans and struggles with faint or rotated pages.</li>\n</ol>\n<p>Vision-language models (VLMs, models that read images as well as text) have recently changed the calculus for OCR. A capable multimodal model can read a scanned page, preserve its table structure, and emit Markdown in one shot. On messy layouts that is genuinely better than Tesseract. The catch is cost and latency: you pay for a model call per page instead of running a local binary. For a small corpus of gnarly scans, that is a fine trade. For millions of pages it is not, and you want the cheap engine with good preprocessing.</p>\n<p>On ARGRAG we mixed both. Fast local extraction handled the clean majority, and we held heavier tools in reserve for the documents that actually needed them. Paying VLM prices for every page in the corpus would have been the wrong default.</p>\n<h2>The failure modes that bite in production</h2>\n<p>The happy path is a few dozen lines. The interesting part, as always, is where it breaks.</p>\n<p><strong>Silent empty extractions.</strong> Plain extraction on a scanned page returns an empty string, not an error. Without the triage step, those pages vanish from your index, and nobody notices until a user asks about content that simply is not there. Log a per-page character count during ingestion and alert on pages that extracted almost nothing; it is the cheapest observability you will ever add.</p>\n<p><strong>Reading order scrambled across columns.</strong> Two-column PDFs are the classic case. The extracted text looks fine to a regex but reads as nonsense to a model, because sentences jump the gutter. If retrieval quality is oddly bad on a subset of documents, render a few and check whether they are multi-column before you blame the embedding model.</p>\n<p><strong>Headers, footers, and page numbers as noise.</strong> Running headers and footers repeat on every page, and once chunked they dilute your embeddings with boilerplate. A footer that repeats the document title on all 80 pages will match a query about the title and bury the page you actually wanted. Before chunking, strip lines that repeat in the same position across most pages.</p>\n<p><strong>Tables split mid-row by chunking.</strong> A fixed-size chunker can cut even a perfectly extracted Markdown table in half, orphaning the second half from its header. Keep each table intact as its own chunk when it fits. When a table is too big, repeat the header row in each piece.</p>\n<p><strong>Ligatures and hyphenation.</strong> Extracted text often contains ligatures, two letters stored as a single glyph (the “fi” in “efficient”), and words hyphenated across line breaks. Both quietly hurt exact-match retrieval. A small normalization pass cleans up most of it: Unicode NFKC normalization, which splits ligatures back into plain letters, plus rejoining words that were hyphenated at a line break.</p>\n<p><strong>Encrypted or malformed files.</strong> Real corpora contain password-protected PDFs, truncated files, and PDFs that claim one page count and deliver another. Wrap extraction for each file in error handling, record which files failed and why, and keep going. One bad file should never take down a batch ingest of ten thousand.</p>\n<h2>What I would set up on day one</h2>\n<p>If I were starting a PDF ingestion pipeline fresh, this order saves the most pain:</p>\n<ol>\n<li>Triage every page for a text layer first.</li>\n<li>Run cheap extraction on the born-digital majority.</li>\n<li>Route only the genuinely scanned pages to OCR.</li>\n<li>Treat tables as a first-class output, not an afterthought.</li>\n<li>Emit Markdown, and carry the source document and page number as metadata on every chunk so answers can cite where they came from.</li>\n<li>Log per-page character counts, so a silent parsing regression shows up as a metric instead of a support ticket.</li>\n</ol>\n<p>I keep coming back to this because parsing sets the ceiling for everything downstream. The most careful work on <a href=\"/blog/2026-07-28-choosing-embedding-model-rag/\">choosing an embedding model</a> or <a href=\"/blog/2026-08-28-contextual-retrieval-rag/\">contextual retrieval</a> is capped by what the parser managed to recover from the page. Get a clean, structured, correctly ordered representation out of the PDF, and the rest of the RAG stack has something real to work with. Feed it flattened tables and scrambled columns, and no reranker will save you. The same instinct is behind the <a href=\"/project/archi/\">Archi ingestion pipeline</a>: retrieval is only ever as good as the parse that fed it.</p>\n<hr>\n<p><em>Diagrams by M. Hassan Ahmed, released under CC0. No external image was used for this post; the figures are original work by the author.</em></p>","frontmatter":{"title":"Parsing PDFs for RAG: Text, Tables, and OCR","date":"2026-09-07T00:00:00.000Z","description":"How to parse PDFs for a RAG pipeline: detect born-digital vs scanned files, pull out tables without flattening them, and reach for OCR only when you must.","thumbnail":{"childImageSharp":{"fluid":{"base64":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAABQAAAALCAIAAADwazoUAAAACXBIWXMAAAsSAAALEgHS3X78AAACAklEQVQoz0WQWVPbMBSF81JI7HiT5D22LG+yvMUEShzbIQkM9KGFKXT6QP//D+kNzLQz35w5vjpndK2ZwUejGFExGnzQst5pT+L4lg7P+fgCmo3Pye57Nr4kux/l8c1p79Wsh+RHfpxJqiervqR4OmbYTm1frK+PdbcX9a5eT931kVe9aHZF1Zft6IU1tjOdMMhDayajUDbOLBFVCFuiaKGtFloAOgdUf6GfvaSHcy1Y4ugzIxkhMJP1QEGRiiIFMwX6OFI/DHIyyxfEL0yXY5crKNQIHAEUFDwkYW1fj2+s7hvKelKdMB9IfSLijhR7a/1Eyr07/LLSr4wVIc1plEdxSWkaJ6XtxVB2UTF5+3dSHuztT7N5cO/e3cMfs3tyx9+4PDj7d7ucymqTZnWeN1xs4jiv2q23Ss4PptMrzCcjuUVpD5jrRwBWwNmA0i1pHkh6w5iIYs4Yj5MqDNM4rV24ea44qpVi2mqeQKtKdbjuCc0tlnaOglrzChQ2qpOFQeT71LRXlDJs+nEUY3M1WygO/DOpjrC81ZyS6TWdXs+6f7Pbexga4uBVQyt4U9bT9vo43NxedbUosziBsrV0CzXq1LA12MbKt1Z+C9i8N+INDBXaobCiQcRYniY8+SCiiesGs/nSmsvmJ5cSuVjgf8Dn/7lkXkjkC7AATy5l80I2/wIPgFRej56GUAAAAABJRU5ErkJggg==","aspectRatio":1.899441340782123,"src":"/static/2cd1baef737ff704c51e7dc49d568c7e/40a76/hero.png","srcSet":"/static/2cd1baef737ff704c51e7dc49d568c7e/c972b/hero.png 340w,\n/static/2cd1baef737ff704c51e7dc49d568c7e/27625/hero.png 680w,\n/static/2cd1baef737ff704c51e7dc49d568c7e/40a76/hero.png 1360w,\n/static/2cd1baef737ff704c51e7dc49d568c7e/ed396/hero.png 2000w","sizes":"(max-width: 1360px) 100vw, 1360px"}}}}}},"pageContext":{"slug":"/2026-09-07-parsing-pdfs-for-rag/","previous":"blog/2026-09-08-ivf-index-vector-search/","next":"blog/2026-09-11-cosine-dot-product-euclidean-vector-search/"}},"staticQueryHashes":["32046230"]}