{"componentChunkName":"component---src-templates-blog-post-js","path":"/blog/2026-09-19-minhash-lsh-near-duplicate-detection-rag/","result":{"data":{"site":{"siteMetadata":{"title":"M.Hassan Ahmed","author":"Hassan11196"}},"markdownRemark":{"id":"bfb5012f-35ea-5f68-9ad7-06dc5a14f989","excerpt":"The first time retrieval quality on Archi dropped for no obvious reason, the cause turned out to be embarrassingly mundane: duplicates. The crawler that indexes…","html":"<p>The first time retrieval quality on <a href=\"/project/archi/\">Archi</a> dropped for no obvious reason, the cause turned out to be embarrassingly mundane: duplicates. The crawler that indexes CERN’s internal portals, JIRA tickets, and operator logbooks was doing its job a little too well. It pulled in:</p>\n<ul>\n<li>the same runbook, which lives on three different wiki spaces;</li>\n<li>the original ticket plus two forwarded copies, each with a reply quoted underneath;</li>\n<li>a status page that regenerates hourly with one timestamp changed.</li>\n</ul>\n<p>None of these were byte-for-byte identical, so an md5 hash of each document came out different. As far as the index knew, they were four separate sources of truth.</p>\n<p>That is the problem this post tackles. When a corpus is full of near-duplicates, retrieval spends its top-k slots (the k best-scoring results) returning the same content in slightly different clothes. The reranker wastes effort ordering copies of one answer, and the context window fills with redundancy instead of coverage. This post is for anyone building an ingestion pipeline for RAG (retrieval-augmented generation) or search who has more documents than they can eyeball. You have probably noticed that “just hash it” quietly fails the moment two documents are 98% the same instead of 100%.</p>\n<p>The tools that fix this are MinHash and Locality-Sensitive Hashing (LSH). They are old: MinHash comes from Andrei Broder’s work on deduplicating the AltaVista web index in the late 1990s (<a href=\"https://ieeexplore.ieee.org/document/666900\">Broder, 1997</a>). They hold up because the problem never went away. Below I go through the pipeline in order: turning documents into sets, compressing the sets into MinHash signatures, using LSH to find candidate pairs, and turning those pairs into clean clusters.</p>\n<h2>Why exact hashing and pairwise comparison both fail</h2>\n<p>Exact hashing is the obvious first move, and it catches exactly one case: documents that are identical down to the last byte. Change a single character (a trailing newline, a re-rendered timestamp, a <code class=\"language-text\">Fwd:</code> in the subject) and the hash is completely different. Cryptographic hashes are designed to behave this way. It is the opposite of what you want when the question is “are these two documents basically the same thing?”</p>\n<p>So you reach for a similarity measure. The natural one for text is the <a href=\"https://en.wikipedia.org/wiki/Jaccard_index\">Jaccard index</a>. Represent each document as a set of features, then measure the overlap as the size of the intersection divided by the size of the union:</p>\n<div class=\"gatsby-highlight\" data-language=\"text\"><pre class=\"language-text\"><code class=\"language-text\">J(A, B) = |A ∩ B| / |A ∪ B|</code></pre></div>\n<p>If two documents share most of their features, Jaccard is close to 1. If they share almost nothing, it is close to 0. That is exactly the “basically the same” signal that exact hashing throws away.</p>\n<p>The catch is cost. Computing Jaccard for every pair of documents takes O(n²) comparisons, and each comparison touches two full feature sets. For a few thousand documents that is fine. A real ops corpus accumulates hundreds of thousands, and at that size n² means tens of billions of set operations; you will not finish before the next crawl starts. MinHash and LSH exist to get the answer of that O(n²) scan without paying for it.</p>\n<h2>Step one: turn documents into sets of shingles</h2>\n<p>Before any hashing, you need the sets. The standard construction is <strong>shingling</strong> (also called k-grams): slide a window of <code class=\"language-text\">k</code> consecutive tokens across the document and collect every window. For example, the set of word 3-grams for “the disk was full again” is <code class=\"language-text\">{&quot;the disk was&quot;, &quot;disk was full&quot;, &quot;was full again&quot;}</code>. The function below builds that set from lowercased, whitespace-split words, and returns the whole text as one shingle when it is shorter than <code class=\"language-text\">k</code>.</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\"><span class=\"token keyword\">def</span> <span class=\"token function\">shingles</span><span class=\"token punctuation\">(</span>text<span class=\"token punctuation\">:</span> <span class=\"token builtin\">str</span><span class=\"token punctuation\">,</span> k<span class=\"token punctuation\">:</span> <span class=\"token builtin\">int</span> <span class=\"token operator\">=</span> <span class=\"token number\">3</span><span class=\"token punctuation\">)</span> <span class=\"token operator\">-</span><span class=\"token operator\">></span> <span class=\"token builtin\">set</span><span class=\"token punctuation\">[</span><span class=\"token builtin\">str</span><span class=\"token punctuation\">]</span><span class=\"token punctuation\">:</span>\n    tokens <span class=\"token operator\">=</span> text<span class=\"token punctuation\">.</span>lower<span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">.</span>split<span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span>\n    <span class=\"token keyword\">if</span> <span class=\"token builtin\">len</span><span class=\"token punctuation\">(</span>tokens<span class=\"token punctuation\">)</span> <span class=\"token operator\">&lt;</span> k<span class=\"token punctuation\">:</span>\n        <span class=\"token keyword\">return</span> <span class=\"token punctuation\">{</span><span class=\"token string\">\" \"</span><span class=\"token punctuation\">.</span>join<span class=\"token punctuation\">(</span>tokens<span class=\"token punctuation\">)</span><span class=\"token punctuation\">}</span>\n    <span class=\"token keyword\">return</span> <span class=\"token punctuation\">{</span><span class=\"token string\">\" \"</span><span class=\"token punctuation\">.</span>join<span class=\"token punctuation\">(</span>tokens<span class=\"token punctuation\">[</span>i <span class=\"token punctuation\">:</span> i <span class=\"token operator\">+</span> k<span class=\"token punctuation\">]</span><span class=\"token punctuation\">)</span> <span class=\"token keyword\">for</span> i <span class=\"token keyword\">in</span> <span class=\"token builtin\">range</span><span class=\"token punctuation\">(</span><span class=\"token builtin\">len</span><span class=\"token punctuation\">(</span>tokens<span class=\"token punctuation\">)</span> <span class=\"token operator\">-</span> k <span class=\"token operator\">+</span> <span class=\"token number\">1</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">}</span></code></pre></div>\n<p>The choice of <code class=\"language-text\">k</code> is the first real tradeoff:</p>\n<ul>\n<li><strong>Small shingles</strong> (<code class=\"language-text\">k=1</code>, individual words) make almost every document overlap, because English documents share common words. Unrelated texts get high Jaccard scores.</li>\n<li><strong>Large shingles</strong> (<code class=\"language-text\">k=8</code>) only match on long verbatim runs, so a paraphrase or a reformatted paragraph looks completely new.</li>\n</ul>\n<p>For prose I have found word 3-grams to 5-grams a reasonable default. The <a href=\"http://www.mmds.org/\"><em>Mining of Massive Datasets</em></a> textbook (chapter 3, the definitive treatment of this whole pipeline) suggests picking <code class=\"language-text\">k</code> large enough that any given shingle has a low probability of appearing in a document. For short, noisy text like log lines, where word boundaries are unreliable, consider character shingles instead.</p>\n<p>Whatever you pick, apply it consistently and normalize first. Lowercase the text, collapse whitespace, and strip the boilerplate header and footer that every wiki page carries, all <em>before</em> shingling. Otherwise a shared navigation bar makes two unrelated pages look similar. Boilerplate is the single most common reason a dedup pass produces nonsense, and it deserves a dedicated cleaning step.</p>\n<h2>Step two: compress each set into a MinHash signature</h2>\n<p>Storing and intersecting full shingle sets is still expensive. MinHash replaces each set with a short, fixed-length <strong>signature</strong>, and comparing two signatures estimates their Jaccard similarity directly.</p>\n<p>Here is the trick. Apply a random hash function to every element of a set and keep the minimum value. For two sets A and B, the probability that they produce the <em>same</em> minimum equals their Jaccard similarity. Repeat with <code class=\"language-text\">num_perm</code> independent hash functions and each document gets a signature of <code class=\"language-text\">num_perm</code> integers. The fraction of positions where two signatures agree is then an unbiased estimate of J(A, B), and more permutations give a tighter estimate. (<a href=\"https://en.wikipedia.org/wiki/MinHash\">Wikipedia’s MinHash article</a> has the proof if you want to see why the minimum trick works.)</p>\n<p>You do not implement this by hand in production. The <a href=\"https://ekzhu.github.io/datasketch/\"><code class=\"language-text\">datasketch</code></a> library is the well-worn Python choice. The code below builds a MinHash from a text’s 3-gram shingles, then compares two logbook-style sentences that differ only at the end:</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\"><span class=\"token keyword\">from</span> datasketch <span class=\"token keyword\">import</span> MinHash\n\n<span class=\"token keyword\">def</span> <span class=\"token function\">minhash</span><span class=\"token punctuation\">(</span>text<span class=\"token punctuation\">:</span> <span class=\"token builtin\">str</span><span class=\"token punctuation\">,</span> num_perm<span class=\"token punctuation\">:</span> <span class=\"token builtin\">int</span> <span class=\"token operator\">=</span> <span class=\"token number\">128</span><span class=\"token punctuation\">)</span> <span class=\"token operator\">-</span><span class=\"token operator\">></span> MinHash<span class=\"token punctuation\">:</span>\n    m <span class=\"token operator\">=</span> MinHash<span class=\"token punctuation\">(</span>num_perm<span class=\"token operator\">=</span>num_perm<span class=\"token punctuation\">)</span>\n    <span class=\"token keyword\">for</span> sh <span class=\"token keyword\">in</span> shingles<span class=\"token punctuation\">(</span>text<span class=\"token punctuation\">,</span> k<span class=\"token operator\">=</span><span class=\"token number\">3</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n        m<span class=\"token punctuation\">.</span>update<span class=\"token punctuation\">(</span>sh<span class=\"token punctuation\">.</span>encode<span class=\"token punctuation\">(</span><span class=\"token string\">\"utf-8\"</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">)</span>\n    <span class=\"token keyword\">return</span> m\n\na <span class=\"token operator\">=</span> minhash<span class=\"token punctuation\">(</span><span class=\"token string\">\"the disk on node 14 was full again this morning\"</span><span class=\"token punctuation\">)</span>\nb <span class=\"token operator\">=</span> minhash<span class=\"token punctuation\">(</span><span class=\"token string\">\"the disk on node 14 was full again last night\"</span><span class=\"token punctuation\">)</span>\n<span class=\"token keyword\">print</span><span class=\"token punctuation\">(</span>a<span class=\"token punctuation\">.</span>jaccard<span class=\"token punctuation\">(</span>b<span class=\"token punctuation\">)</span><span class=\"token punctuation\">)</span>  <span class=\"token comment\"># estimated Jaccard, ~0.5 here</span></code></pre></div>\n<p><code class=\"language-text\">num_perm</code> is the dial between accuracy and memory. 128 permutations is a common starting point. The estimate’s standard error falls roughly with the square root of <code class=\"language-text\">num_perm</code>, so doubling it does not halve the error, and going past 256 rarely earns its cost for dedup. The important win is already in hand: every document is now a fixed signature of 128 integers, however long the original was.</p>\n<p>The diagram below shows the whole pipeline, including the LSH stage that comes next.</p>\n<p><img src=\"/276cea2569264485a770c05cfae8219c/pipeline.svg\" alt=\"Pipeline diagram showing five stages left to right. Stage one, Documents, lists ticket 4021, ticket 4021 forwarded, wiki page A, and wiki page A copy, labelled full text. Stage two, Shingles, produces overlapping k-word windows such as the disk was and disk was full, labelled a set per document. Stage three, MinHash, produces a signature of num_perm integers like 7 2 9 1 that approximates Jaccard similarity, labelled fixed-size vector. Stage four, LSH banding, splits each signature into b bands and hashes each band into a bucket, labelled no O of n squared scan. Stage five, Pairs, yields candidate near-duplicates to verify and then keep one of, labelled tiny set. A note explains MinHash turns set similarity into cheap integer comparison, and LSH makes two documents a candidate pair only when at least one of their b bands hashes to the same bucket.\"></p>\n<h2>Step three: LSH finds candidates without comparing every pair</h2>\n<p>MinHash shrinks each document, but comparing every signature to every other is still O(n²), just with cheaper comparisons. LSH removes the n² entirely:</p>\n<ol>\n<li>Split each signature into <code class=\"language-text\">b</code> bands of <code class=\"language-text\">r</code> rows each, so that <code class=\"language-text\">b * r = num_perm</code>.</li>\n<li>Hash each band to a bucket.</li>\n<li>Treat two documents as a <strong>candidate pair</strong> if they land in the same bucket for at least one band.</li>\n</ol>\n<p>The intuition: two very similar signatures will probably agree across a whole band somewhere. Two dissimilar ones are unlikely to agree on an entire band by chance.</p>\n<p>The probability that a pair with Jaccard similarity <code class=\"language-text\">s</code> becomes a candidate is:</p>\n<div class=\"gatsby-highlight\" data-language=\"text\"><pre class=\"language-text\"><code class=\"language-text\">P(candidate) = 1 - (1 - s^r)^b</code></pre></div>\n<p>That function is an S-curve. It stays near zero below a threshold, sits near one above it, and the transition is sharp. The threshold, the point where the probability crosses 0.5, is approximately <code class=\"language-text\">(1/b)^(1/r)</code>. You tune <code class=\"language-text\">b</code> and <code class=\"language-text\">r</code> to put that threshold where your definition of “duplicate” lives. The chart below plots the curve for 16 bands of 8 rows.</p>\n<p><img src=\"/6d757ce8626afce89c36fafd6143539a/lsh-s-curve.svg\" alt=\"Line chart of the LSH banding probability curve for b equals 16 bands of r equals 8 rows, with num_perm 128. The x-axis is Jaccard similarity from 0 to 1, the y-axis is the probability a pair becomes a candidate from 0 to 1. The curve stays near zero until about 0.6, rises steeply through a threshold near 0.71 marked with a dashed red line and the formula one over b to the power one over r, then flattens near one past 0.85. Annotations note that unrelated documents are almost never paired while near-duplicates are almost always paired, and that changing b and r slides and sharpens the curve.\"></p>\n<p>In practice you let the library solve for <code class=\"language-text\">b</code> and <code class=\"language-text\">r</code> from a target threshold. <code class=\"language-text\">datasketch</code> does this when you build the index. The code below inserts every document’s signature into an index set for Jaccard of 0.8 or more, then asks for the near-duplicates of one ticket:</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\"><span class=\"token keyword\">from</span> datasketch <span class=\"token keyword\">import</span> MinHashLSH\n\n<span class=\"token comment\"># Index near-duplicates at Jaccard >= 0.8</span>\nlsh <span class=\"token operator\">=</span> MinHashLSH<span class=\"token punctuation\">(</span>threshold<span class=\"token operator\">=</span><span class=\"token number\">0.8</span><span class=\"token punctuation\">,</span> num_perm<span class=\"token operator\">=</span><span class=\"token number\">128</span><span class=\"token punctuation\">)</span>\n\nsignatures <span class=\"token operator\">=</span> <span class=\"token punctuation\">{</span>doc_id<span class=\"token punctuation\">:</span> minhash<span class=\"token punctuation\">(</span>text<span class=\"token punctuation\">)</span> <span class=\"token keyword\">for</span> doc_id<span class=\"token punctuation\">,</span> text <span class=\"token keyword\">in</span> corpus<span class=\"token punctuation\">.</span>items<span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">}</span>\n<span class=\"token keyword\">for</span> doc_id<span class=\"token punctuation\">,</span> m <span class=\"token keyword\">in</span> signatures<span class=\"token punctuation\">.</span>items<span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n    lsh<span class=\"token punctuation\">.</span>insert<span class=\"token punctuation\">(</span>doc_id<span class=\"token punctuation\">,</span> m<span class=\"token punctuation\">)</span>\n\n<span class=\"token comment\"># For any document, retrieve its near-duplicate candidates in ~constant time</span>\ndupes_of_4021 <span class=\"token operator\">=</span> lsh<span class=\"token punctuation\">.</span>query<span class=\"token punctuation\">(</span>signatures<span class=\"token punctuation\">[</span><span class=\"token string\">\"ticket-4021\"</span><span class=\"token punctuation\">]</span><span class=\"token punctuation\">)</span></code></pre></div>\n<p><code class=\"language-text\">query</code> returns candidates in roughly constant time per document, so a full dedup pass over the corpus is close to linear instead of quadratic. That is the difference between a job that finishes during ingestion and one that never runs.</p>\n<h2>Turning candidate pairs into a clean corpus</h2>\n<p>LSH gives you candidate <em>pairs</em>, not final clusters, and you deliberately tune it to over-return slightly. Two steps close the gap.</p>\n<p><strong>First, verify.</strong> LSH is only a filter. Confirm each candidate with the actual MinHash estimate (or the exact Jaccard, if you kept the shingle sets) and drop pairs below your real threshold. This removes the small number of false positives that banding lets through.</p>\n<p><strong>Second, cluster.</strong> In practice duplication is transitive: if A duplicates B and B duplicates C, all three are the same document. Build a graph with the confirmed pairs as edges and take its connected components. A <a href=\"https://en.wikipedia.org/wiki/Disjoint-set_data_structure\">union-find</a> structure does this in near-linear time. Each component becomes one cluster, and you keep a single canonical member. My rule for “canonical” is the longest, most recently updated version, since for an ops runbook the fullest copy is usually the one worth indexing. Everything else is dropped from the index or collapsed into a pointer to the canonical document.</p>\n<h2>Failure modes worth knowing before you ship</h2>\n<p><strong>Semantic duplicates are invisible to this.</strong> MinHash measures <em>lexical</em> overlap, meaning shared wording rather than shared meaning. “The node ran out of memory” and “the machine hit an OOM (out-of-memory) condition” mean the same thing but share almost no shingles, so their Jaccard is near zero and LSH will never pair them. That is not a bug; it is the boundary of the technique. Same-meaning, different-words duplicates are a job for embeddings and vector similarity, a separate pass with different tradeoffs that I covered in <a href=\"/blog/2026-07-28-choosing-embedding-model-rag/\">choosing an embedding model for RAG</a>. Use MinHash for the cheap lexical dupes and embeddings for the semantic ones; they solve different halves of the same problem.</p>\n<p><strong>Short documents are unstable.</strong> A one-line log message has very few shingles, so changing a single word swings its Jaccard wildly. Below some length the estimate is too noisy to trust. Set a minimum token count, and fall back to exact-match or normalized-string comparison for anything shorter.</p>\n<p><strong>Boilerplate dominates similarity if you let it.</strong> This is the one that will actually burn you. If every page shares a 200-word footer, two otherwise unrelated pages can clear an 0.8 threshold on the footer alone. Strip templated chrome before shingling, and when Jaccard scores come back suspiciously high for obviously different documents, suspect boilerplate first.</p>\n<p><strong>The threshold is a product decision, not a default.</strong> An 0.8 threshold treats “80% overlapping” as duplicate. Whether that is right depends on whether your corpus holds legitimately similar-but-distinct documents; two release notes for consecutive versions might be 85% identical and both worth keeping. Sample the pairs near your threshold and read them before you trust the number.</p>\n<h2>What I would do differently</h2>\n<p>On the first pass I treated dedup as a one-time batch job over the whole corpus. That is the wrong shape for a crawler that runs continuously. The <code class=\"language-text\">MinHashLSH</code> index supports incremental <code class=\"language-text\">insert</code>, so the better design checks each newly crawled document against the standing index at ingestion time. That is the moment to decide whether it is a new source or a copy of one already indexed. It is the same instinct behind doing <a href=\"/blog/2026-06-29-incremental-rag-indexing/\">RAG indexing incrementally</a> rather than rebuilding from scratch. I would also persist the signatures instead of recomputing them every run: they are small, and shingling the full corpus again and again is wasted work.</p>\n<p>The other thing I underweighted early was measurement. It is tempting to pick a threshold, watch the duplicate count drop, and call it done. But the number that matters is whether retrieval got better. The only way to know is to hold out a labelled query set and measure recall and answer quality before and after, the discipline I leaned on in <a href=\"/blog/2026-07-04-measuring-rag-retrieval-quality/\">measuring RAG retrieval quality</a>. A dedup pass that quietly drops a document you needed is worse than the redundancy it removed.</p>\n<h2>Reach for the cheap, old algorithm first</h2>\n<p>Near-duplicate detection is unglamorous plumbing, and it is exactly the kind of thing that decides whether a RAG system feels sharp or sludgy. A deduped corpus returns diverse, non-redundant context. A corpus that has not been deduped spends its best retrieval slots saying the same thing three times. MinHash and LSH let you dedupe at a scale where reading the documents yourself stopped being an option long ago, and they do it in close to linear time.</p>\n<p>I kept coming back to this pattern across the data plumbing behind <a href=\"/project/archi/\">Archi</a> and the <a href=\"/project/cms-workflow-operations/\">CMS workflow operations</a> tooling at CERN. The cheap, approximate, old algorithm that gets you 95% of the way for 5% of the cost is usually the one worth reaching for first. Save the expensive semantic pass for the cases the cheap one provably cannot catch.</p>\n<hr>\n<p><em>Diagrams by M. Hassan Ahmed, released under CC0. No external image was used for this post; the figures are original work by the author.</em></p>","frontmatter":{"title":"Dedupe a RAG Corpus with MinHash and LSH","date":"2026-09-19T00:00:00.000Z","description":"Exact md5 hashing misses near-duplicate documents; comparing every pair is O(n²). How MinHash and LSH find and cluster them across a RAG corpus.","thumbnail":{"childImageSharp":{"fluid":{"base64":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAABQAAAALCAIAAADwazoUAAAACXBIWXMAAAsSAAALEgHS3X78AAACDUlEQVQoz02PW2+bMBiGuVhOgDEHGxswhnAMJKQkKLRJIUnTbqu0dZo0bVOlSpN2sV3tF+zPz+m0atJj6/2Ory2VGStSbz7zT/um3y537bzbVoeuPh2aw3UtwtentttV15eLY7/qrqoy94uMlc9IIy0cggCSYlbfpFUfz7t0sRe62rwp6lNUdvP1XTzvk0VfrE5ptddJOQBcTAkk1eQCYAWyzgR/Q8XwhVYMIbisc8UMZSMQWjUDgMJz/3ObJGoTTYz5qhGImuXkNpubNJMhkzVnrDnISVgQO35ikhS7M6AHE+CMIFN0X/KSTV7f2MHKdCud5JjXdriGJEPBJjv+rJuv3z9sH59+3z7+KHef+fEBtJfl7be3ix6argTtxKA5wImGE/E8RLNiegGsqYaiMS5cP46ikLI8jGPKogmOFDsEJDFQJENPcqMmXe6j8tpiS4BTxJZuuNZwavhNd3Gq63df7tv797/6j49V+7CqP4XlvmmfdkWr6Y4ESWo6M2EunCGOkJOLD5t2Au2M+wvCLqoi6bZXSdUgf0XZ2nBKyi8ZzRToSGOVvKDpLsKBafmWxRWNvpLxREU68jgnGNtAs6BhQh3phgF0PAFUEueFMSBDxR6pZ8aAyiKjkjz113VaFZ7j8ek05EGYxKHrehNAhLP9H0TM/7vPGbFFBhhA4WMPFTpQ6FAhA5kO5HPDH8eyT3LwxOXyAAAAAElFTkSuQmCC","aspectRatio":1.899441340782123,"src":"/static/4aa6b11a00da437300ba4f3109b953d1/40a76/hero.png","srcSet":"/static/4aa6b11a00da437300ba4f3109b953d1/c972b/hero.png 340w,\n/static/4aa6b11a00da437300ba4f3109b953d1/27625/hero.png 680w,\n/static/4aa6b11a00da437300ba4f3109b953d1/40a76/hero.png 1360w,\n/static/4aa6b11a00da437300ba4f3109b953d1/ed396/hero.png 2000w","sizes":"(max-width: 1360px) 100vw, 1360px"}}}}}},"pageContext":{"slug":"/2026-09-19-minhash-lsh-near-duplicate-detection-rag/","previous":"blog/2026-09-20-mixture-of-experts-explained/","next":"blog/2026-09-24-roaring-bitmaps-fast-filters/"}},"staticQueryHashes":["32046230"]}