{"componentChunkName":"component---src-templates-blog-post-js","path":"/blog/2026-09-01-metadata-filtering-rag-vector-search/","result":{"data":{"site":{"siteMetadata":{"title":"M.Hassan Ahmed","author":"Hassan11196"}},"markdownRemark":{"id":"6d99677f-e8e4-5f96-8abd-559f92d363f6","excerpt":"Most RAG bugs I have chased were not embedding-quality problems. They were filter problems. Someone asks the copilot about a transfer error at a specific site…","html":"<p>Most RAG bugs I have chased were not embedding-quality problems. They were filter problems. Someone asks the copilot about a transfer error at a specific site, and the answer quotes a fix from a different site two years earlier. Everyone concludes the retrieval is bad. It usually is not: the nearest vectors were fine. The system just failed to restrict the search to the rows that were actually relevant.</p>\n<p>This post is for engineers building retrieval on a vector store who need results scoped by structured attributes: a site, a date range, a document type, an access level. It covers:</p>\n<ul>\n<li>the three ways an engine can combine a metadata filter with an approximate-nearest-neighbor (ANN) search, and why the naive one quietly hurts recall;</li>\n<li>why filtering an <a href=\"/blog/2026-07-09-hnsw-vector-search-explained/\">HNSW</a> graph is harder than it looks;</li>\n<li>how the choice plays out in Postgres, Qdrant, and OpenSearch.</li>\n</ul>\n<p>The running example is <a href=\"/project/archi/\">Archi</a>, the retrieval copilot I worked on for CMS computing operations at CERN. There, “only entries from this site, this year” is part of almost every query.</p>\n<h2>Why vector distance alone is not enough</h2>\n<p>A vector index answers one question well: given this query vector, which stored vectors are closest? It knows nothing about where a chunk came from, such as a 2021 logbook for a different site. So once your data has structure (and operations data always does), you need to satisfy two constraints at once: <em>close in embedding space</em> and <em>matches these attributes</em>.</p>\n<p>The three strategies below differ only in <strong>when</strong> the engine applies the filter relative to the search. That ordering decides whether you get correct results, fast results, or neither.</p>\n<p><img src=\"/0fe0897c5346de52595e5214264ee7ae/filtering-strategies.svg\" alt=\"Three columns comparing post-filter, exact pre-filter, and filtered ANN. Post-filter searches all vectors then drops non-matching rows and loses recall; exact pre-filter selects matching rows then brute-force scores them, correct but O(n); filtered ANN applies the filter during HNSW traversal, fast and filter-aware. A footer rule of thumb says never post-filter in production.\"></p>\n<h2>Strategy 1: post-filter, and why it fails quietly</h2>\n<p>The tempting version, especially when you bolt a filter onto an existing search, is to run the vector query, take the top <em>k</em>, then throw away the rows that do not match. This is a post-filter, sometimes called <em>filter-after-search</em>. The query below is the kind that looks correct but behaves this way:</p>\n<div class=\"gatsby-highlight\" data-language=\"sql\"><pre class=\"language-sql\"><code class=\"language-sql\"><span class=\"token comment\">-- pgvector: search first, filter the result set. Looks fine, isn't.</span>\n<span class=\"token keyword\">SELECT</span> id<span class=\"token punctuation\">,</span> body\n<span class=\"token keyword\">FROM</span> chunks\n<span class=\"token keyword\">WHERE</span> embedding <span class=\"token operator\">&lt;=></span> :query_vec <span class=\"token operator\">&lt;</span> <span class=\"token number\">1.0</span>        <span class=\"token comment\">-- distance, computed on ALL rows</span>\n  <span class=\"token operator\">AND</span> site <span class=\"token operator\">=</span> <span class=\"token string\">'T2_US_MIT'</span> <span class=\"token operator\">AND</span> <span class=\"token keyword\">year</span> <span class=\"token operator\">>=</span> <span class=\"token number\">2024</span>   <span class=\"token comment\">-- filter applied to top-k output</span>\n<span class=\"token keyword\">ORDER</span> <span class=\"token keyword\">BY</span> embedding <span class=\"token operator\">&lt;=></span> :query_vec\n<span class=\"token keyword\">LIMIT</span> <span class=\"token number\">10</span><span class=\"token punctuation\">;</span></code></pre></div>\n<p>The failure mode is not an error. It is silent recall loss (recall being the share of truly relevant documents that come back). Here is how it happens:</p>\n<ol>\n<li>Say <code class=\"language-text\">site = &#39;T2_US_MIT&#39;</code> covers 3% of your corpus, and the query is generic enough that most of the closest vectors come from busier sites.</li>\n<li>The ANN search returns its ten nearest neighbors, and eight of them come from other sites.</li>\n<li>The filter removes those eight, leaving two rows.</li>\n<li>The genuinely relevant MIT document sat at rank fourteen. It was never fetched, so it never had a chance.</li>\n</ol>\n<p>You asked for <code class=\"language-text\">LIMIT 10</code> and got two. Nobody sees a stack trace; you just see a worse answer.</p>\n<p>You can paper over this by over-fetching: request <code class=\"language-text\">k = 200</code>, filter down, and hope enough rows survive. But that is a guess, and the more selective the filter, the worse the guess gets. For a filter matching 0.1% of rows, you might fetch thousands and still come up short. Post-filtering is fine for a demo and a liability in production. Of the three options in the diagram, it is the one I would tell you never to ship.</p>\n<h2>Strategy 2: exact pre-filter</h2>\n<p>The opposite order is to filter first, then run an exact (brute-force) distance scan over whatever survives:</p>\n<div class=\"gatsby-highlight\" data-language=\"sql\"><pre class=\"language-sql\"><code class=\"language-sql\"><span class=\"token comment\">-- Filter first, then score exactly on the subset.</span>\n<span class=\"token keyword\">SELECT</span> id<span class=\"token punctuation\">,</span> body\n<span class=\"token keyword\">FROM</span> chunks\n<span class=\"token keyword\">WHERE</span> site <span class=\"token operator\">=</span> <span class=\"token string\">'T2_US_MIT'</span> <span class=\"token operator\">AND</span> <span class=\"token keyword\">year</span> <span class=\"token operator\">>=</span> <span class=\"token number\">2024</span>\n<span class=\"token keyword\">ORDER</span> <span class=\"token keyword\">BY</span> embedding <span class=\"token operator\">&lt;=></span> :query_vec\n<span class=\"token keyword\">LIMIT</span> <span class=\"token number\">10</span><span class=\"token punctuation\">;</span></code></pre></div>\n<p>This is always correct. Every returned row matches the filter, and because the engine scores the whole subset exactly, you get the true nearest neighbors within it.</p>\n<p>The catch is cost. A brute-force scan is O(n) in the size of the filtered subset. If <code class=\"language-text\">site = &#39;T2_US_MIT&#39; AND year &gt;= 2024</code> leaves a few thousand rows, scanning them is nothing, often faster than touching the ANN index at all. If your filter is loose and leaves several million rows, you compute millions of distances per query while the ANN index you built sits unused.</p>\n<p>So exact pre-filter is the right tool precisely when the filter is <em>selective</em>. Postgres knows this: with pgvector, if the planner estimates that the filtered set is small, it skips the vector index and does exactly this scan. The trouble starts in the middle ground, with filters that are neither tiny nor loose. That is where the third strategy earns its place.</p>\n<h2>Strategy 3: filtered ANN, the one you actually want</h2>\n<p>The strategy that scales pushes the filter <em>into</em> the ANN search itself, so the graph traversal only ever considers nodes that satisfy the predicate. The engine wastes no distance computations on rows it will discard, and it never brute-forces a huge subset.</p>\n<h3>Why filtering an HNSW graph is hard</h3>\n<p>This sounds obvious, and it is genuinely hard to implement well. HNSW is a navigable small-world graph: search works by greedily hopping to closer and closer neighbors. When you forbid most nodes, you tear holes in that graph.</p>\n<p>The path from the entry point to the best matching node may run through nodes that fail the filter. If the search cannot step on them, it cannot reach the destination. Recall collapses even though the matching node is in the index.</p>\n<p>The engines that do this well spend their effort on exactly this problem: staying connected while honoring the filter. Qdrant’s team wrote a good explanation of why filtered HNSW is not just “skip the bad nodes” and how they keep the graph traversable (<a href=\"https://qdrant.tech/articles/filtrable-hnsw/\">filtrable HNSW</a>). Research systems like <a href=\"https://arxiv.org/abs/2403.04871\">ACORN</a> attack the same problem by searching an expanded neighborhood, so the traversal can route around filtered-out nodes.</p>\n<h3>How each engine exposes it</h3>\n<p>The practical point: you do not implement this yourself. You use an engine that treats the filter as a first-class part of the query, and you verify its recall on your own data.</p>\n<p><strong>pgvector (0.8+).</strong> Give the planner the filter alongside <code class=\"language-text\">ORDER BY embedding &lt;=&gt; :query_vec</code>. <strong>Iterative index scans</strong> then keep pulling from the HNSW index until enough rows pass the filter, instead of stopping at the first <code class=\"language-text\">k</code> (<a href=\"https://github.com/pgvector/pgvector#filtering\">pgvector filtering docs</a>). The query is the same as in Strategy 2; the <code class=\"language-text\">SET</code> line is what changes the behavior:</p>\n<div class=\"gatsby-highlight\" data-language=\"sql\"><pre class=\"language-sql\"><code class=\"language-sql\"><span class=\"token comment\">-- pgvector 0.8+: iterative scan keeps reading the index until the</span>\n<span class=\"token comment\">-- filter is satisfied, rather than post-filtering a fixed k.</span>\n<span class=\"token keyword\">SET</span> hnsw<span class=\"token punctuation\">.</span>iterative_scan <span class=\"token operator\">=</span> <span class=\"token string\">'relaxed_order'</span><span class=\"token punctuation\">;</span>\n<span class=\"token keyword\">SELECT</span> id<span class=\"token punctuation\">,</span> body\n<span class=\"token keyword\">FROM</span> chunks\n<span class=\"token keyword\">WHERE</span> site <span class=\"token operator\">=</span> <span class=\"token string\">'T2_US_MIT'</span> <span class=\"token operator\">AND</span> <span class=\"token keyword\">year</span> <span class=\"token operator\">>=</span> <span class=\"token number\">2024</span>\n<span class=\"token keyword\">ORDER</span> <span class=\"token keyword\">BY</span> embedding <span class=\"token operator\">&lt;=></span> :query_vec\n<span class=\"token keyword\">LIMIT</span> <span class=\"token number\">10</span><span class=\"token punctuation\">;</span></code></pre></div>\n<p><strong>OpenSearch.</strong> The k-NN query takes a <code class=\"language-text\">filter</code> clause and applies it during the search rather than after. The engine decides between exact and approximate search based on how restrictive the filter is (<a href=\"https://opensearch.org/docs/latest/search-plugins/knn/filter-search-knn/\">efficient k-NN filtering</a>). Archi’s retrieval runs on OpenSearch, which is why I keep coming back to it in these posts. The same index backs both the <a href=\"/blog/2026-07-02-hybrid-search-rag-bm25-vectors/\">BM25 half of hybrid search</a> and the metadata filters here. In the request below, notice that the <code class=\"language-text\">filter</code> sits inside the <code class=\"language-text\">knn</code> clause, not beside it:</p>\n<div class=\"gatsby-highlight\" data-language=\"json\"><pre class=\"language-json\"><code class=\"language-json\"><span class=\"token punctuation\">{</span>\n  <span class=\"token property\">\"size\"</span><span class=\"token operator\">:</span> <span class=\"token number\">10</span><span class=\"token punctuation\">,</span>\n  <span class=\"token property\">\"query\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">{</span>\n    <span class=\"token property\">\"knn\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">{</span>\n      <span class=\"token property\">\"embedding\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">{</span>\n        <span class=\"token property\">\"vector\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">[</span><span class=\"token number\">0.11</span><span class=\"token punctuation\">,</span> <span class=\"token number\">-0.04</span><span class=\"token punctuation\">,</span> <span class=\"token string\">\"...\"</span><span class=\"token punctuation\">]</span><span class=\"token punctuation\">,</span>\n        <span class=\"token property\">\"k\"</span><span class=\"token operator\">:</span> <span class=\"token number\">10</span><span class=\"token punctuation\">,</span>\n        <span class=\"token property\">\"filter\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">{</span>\n          <span class=\"token property\">\"bool\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">{</span>\n            <span class=\"token property\">\"must\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">[</span>\n              <span class=\"token punctuation\">{</span> <span class=\"token property\">\"term\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">{</span> <span class=\"token property\">\"site\"</span><span class=\"token operator\">:</span> <span class=\"token string\">\"T2_US_MIT\"</span> <span class=\"token punctuation\">}</span> <span class=\"token punctuation\">}</span><span class=\"token punctuation\">,</span>\n              <span class=\"token punctuation\">{</span> <span class=\"token property\">\"range\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">{</span> <span class=\"token property\">\"year\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">{</span> <span class=\"token property\">\"gte\"</span><span class=\"token operator\">:</span> <span class=\"token number\">2024</span> <span class=\"token punctuation\">}</span> <span class=\"token punctuation\">}</span> <span class=\"token punctuation\">}</span>\n            <span class=\"token punctuation\">]</span>\n          <span class=\"token punctuation\">}</span>\n        <span class=\"token punctuation\">}</span>\n      <span class=\"token punctuation\">}</span>\n    <span class=\"token punctuation\">}</span>\n  <span class=\"token punctuation\">}</span>\n<span class=\"token punctuation\">}</span></code></pre></div>\n<p><strong>Qdrant.</strong> The filter is a top-level part of the request. Usefully, Qdrant lets you build a <strong>payload index</strong> on the fields you filter by, so the predicate is cheap to evaluate during traversal (<a href=\"https://qdrant.tech/documentation/concepts/filtering/\">Qdrant filtering</a>). The same site-and-year filter looks like this in the Python client:</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\"><span class=\"token keyword\">from</span> qdrant_client <span class=\"token keyword\">import</span> QdrantClient<span class=\"token punctuation\">,</span> models\n\nclient<span class=\"token punctuation\">.</span>query_points<span class=\"token punctuation\">(</span>\n    collection_name<span class=\"token operator\">=</span><span class=\"token string\">\"chunks\"</span><span class=\"token punctuation\">,</span>\n    query<span class=\"token operator\">=</span>query_vec<span class=\"token punctuation\">,</span>\n    query_filter<span class=\"token operator\">=</span>models<span class=\"token punctuation\">.</span>Filter<span class=\"token punctuation\">(</span>must<span class=\"token operator\">=</span><span class=\"token punctuation\">[</span>\n        models<span class=\"token punctuation\">.</span>FieldCondition<span class=\"token punctuation\">(</span>key<span class=\"token operator\">=</span><span class=\"token string\">\"site\"</span><span class=\"token punctuation\">,</span> match<span class=\"token operator\">=</span>models<span class=\"token punctuation\">.</span>MatchValue<span class=\"token punctuation\">(</span>value<span class=\"token operator\">=</span><span class=\"token string\">\"T2_US_MIT\"</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">,</span>\n        models<span class=\"token punctuation\">.</span>FieldCondition<span class=\"token punctuation\">(</span>key<span class=\"token operator\">=</span><span class=\"token string\">\"year\"</span><span class=\"token punctuation\">,</span> <span class=\"token builtin\">range</span><span class=\"token operator\">=</span>models<span class=\"token punctuation\">.</span>Range<span class=\"token punctuation\">(</span>gte<span class=\"token operator\">=</span><span class=\"token number\">2024</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">,</span>\n    <span class=\"token punctuation\">]</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">,</span>\n    limit<span class=\"token operator\">=</span><span class=\"token number\">10</span><span class=\"token punctuation\">,</span>\n<span class=\"token punctuation\">)</span></code></pre></div>\n<h2>Failure modes to know before they bite</h2>\n<p><strong>Selective filters quietly lower recall.</strong> Filtered ANN does not escape the connectivity problem. The more of the graph your filter removes, the more likely the traversal misses the true nearest matching node. When a filter gets extreme, matching well under 1% of rows, most engines fall back to exact search for that query, which is correct but slower. Know where your engine draws that line, because query latency will jump there.</p>\n<p><strong>High-cardinality filters need their own index.</strong> Filtering on a field the engine must check candidate by candidate is slow. Build the field index so the predicate becomes a lookup, not a scan: a Qdrant payload index, an OpenSearch mapping with the field indexed, or a Postgres B-tree on <code class=\"language-text\">site</code>/<code class=\"language-text\">year</code>.</p>\n<p><strong>Range filters on timestamps are the common trap.</strong> <code class=\"language-text\">year &gt;= 2024</code> is cheap. <code class=\"language-text\">created_at BETWEEN two arbitrary instants</code> across a huge table is not, unless the field is indexed and the engine can use it during traversal. In Archi, we filter on timestamps quantized to a <code class=\"language-text\">year</code> (and sometimes <code class=\"language-text\">month</code>) field instead of raw timestamps, and that made these queries predictable.</p>\n<p><strong>Test recall; do not assume it.</strong> The only way to trust a filtered search is to measure it. Take a set of queries with known-relevant filtered documents and check how often they come back. That is the same <a href=\"/blog/2026-07-04-measuring-rag-retrieval-quality/\">retrieval-quality harness</a> you would use for the unfiltered case, run with the filter on. If recall drops with the filter applied, suspect your engine’s filtered traversal, not your embeddings.</p>\n<h2>What I would do differently</h2>\n<p>Early on, I treated the metadata filter as a detail to add after the retrieval “worked.” That is backwards. On operations data, the filter is often the more important half of the query: an operator asking about their site does not want a semantically similar answer from someone else’s.</p>\n<p>If I were starting again, I would design the schema around the filters first:</p>\n<ol>\n<li>Decide which fields get filtered.</li>\n<li>Index those fields.</li>\n<li>Quantize timestamps to something coarse and indexable.</li>\n<li>Only then tune the vector side.</li>\n</ol>\n<p>Getting the filter right removed more bad answers than any embedding-model swap I tried.</p>\n<p>If you are building retrieval for a system where scope matters (a specific tenant, site, product, or time window), reach for your engine’s filtered ANN path, index the fields you filter on, and measure recall with the filter applied. It is a small amount of plumbing that decides whether the copilot answers about the right thing. You can see how these pieces fit together in <a href=\"/project/archi/\">Archi</a> and the rest of my <a href=\"/\">AI and RAG work</a>.</p>","frontmatter":{"title":"Metadata Filtering in RAG Vector Search","date":"2026-09-01T00:00:00.000Z","description":"Post-filtering vector search silently drops good results. Here is pre-filter vs filtered ANN for RAG, why filtered HNSW is hard, and how to choose.","thumbnail":{"childImageSharp":{"fluid":{"base64":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAABQAAAALCAIAAADwazoUAAAACXBIWXMAAAsSAAALEgHS3X78AAAB30lEQVQoz02O666iMBRG+XMUW0oLlEuh5X4TFQV1xOPxGOPMZN7/iWaryWSSRfPt9lu0mlEPBGgGVPb+cKluf4qv38l0B2LgdE/Pj/q1CQFWsb/iagsKiBo2JYBIaFpJVozb4dZ2U92eVuvLcv3ZdFOznNbb66r/grHtzkm6o1aKydPSsB0TJ3OiJVcrJ1lZoiW8QHaC3Aw7GYaek8G4YDG2EygE7Z6JGiwYNUQl5bnIDkFxdI4XkR9lcQrzozNMbnMM4jHIDlxuqF9bfsPLHe9/ML+mXoWZesqIScOKMZOIJxDsoLb9ing5thQ0CFzCFNQwi007M1j8VF6AHGErhqcyr3LiFQ1q5leWX9mqs0QDGY4IzzBXBo9pWNAgx84zg6ghU8BtxC1d2cly5FEn0q1IelEMnloDpluZQUmLJuj3+fWeTle7WZO4fMoLUwCYiAUL55tRN8UMe3Pk6t1W9xIdefB3bEX0/El/Psj3t3m7sV8PUncLEmjwoRfYluk4RWEViSoMynA1hllHaQQFw5Fmt2G70T2d+OFI+x1J67fsv8E0VKIOo0apZSRbKSolW+rEOnYRCw1PmUHq5i1Pa8jYVaBouuH9Y4Zd4APxGXI/8JP5f6c69mYLDrwzrH8BI7BHFvWFfcIAAAAASUVORK5CYII=","aspectRatio":1.899441340782123,"src":"/static/1547a1eef22b6c181246160d46518c3d/40a76/hero.png","srcSet":"/static/1547a1eef22b6c181246160d46518c3d/c972b/hero.png 340w,\n/static/1547a1eef22b6c181246160d46518c3d/27625/hero.png 680w,\n/static/1547a1eef22b6c181246160d46518c3d/40a76/hero.png 1360w,\n/static/1547a1eef22b6c181246160d46518c3d/ed396/hero.png 2000w","sizes":"(max-width: 1360px) 100vw, 1360px"}}}}}},"pageContext":{"slug":"/2026-09-01-metadata-filtering-rag-vector-search/","previous":"blog/2026-08-29-circuit-breaker-llm-api-calls/","next":"blog/2026-08-31-continuous-batching-llm-serving/"}},"staticQueryHashes":["32046230"]}