{"componentChunkName":"component---src-templates-blog-post-js","path":"/blog/2026-07-02-hybrid-search-rag-bm25-vectors/","result":{"data":{"site":{"siteMetadata":{"title":"M.Hassan Ahmed","author":"Hassan11196"}},"markdownRemark":{"id":"c9272a05-b97b-5fec-a63c-586853b48c38","excerpt":"Someone asks the copilot “what fixed the  transfer error last spring?” The retrieval step returns three paragraphs about transfer errors in general and nothing…","html":"<p>Someone asks the copilot “what fixed the <code class=\"language-text\">T1_US_FNAL</code> transfer error last spring?” The retrieval step returns three paragraphs about transfer errors in general and nothing about <code class=\"language-text\">T1_US_FNAL</code>. The embedding model saw a string that looks like every other site name and put it in roughly the same region of vector space as the rest. The one token that mattered got averaged away.</p>\n<p>This failure is what pushes most RAG (retrieval-augmented generation) systems from pure vector search to hybrid search: running keyword search and vector search side by side and merging the results. Dense retrieval is good at meaning and bad at exact strings. Keyword search is the reverse. If your corpus is full of identifiers, error codes, config keys, or proper nouns (and internal operations docs always are), you want both, plus a principled way to merge them.</p>\n<p>This post is for engineers who already have a working vector RAG pipeline and keep hitting queries it retrieves badly. It covers what each retriever is actually good at, how to fuse two ranked lists without hand-tuning weights, why a reranker belongs at the end, and where the whole thing still breaks.</p>\n<h2>Two retrievers that fail in opposite directions</h2>\n<p><a href=\"https://en.wikipedia.org/wiki/Okapi_BM25\">BM25</a> is the lexical (keyword-matching) baseline, and it has held up for thirty years. It scores a document by how often the query terms appear in it, with two adjustments:</p>\n<ul>\n<li>Repeats are damped, so the tenth occurrence of a word counts for less than the first.</li>\n<li>Rare words count for more than common ones.</li>\n</ul>\n<p>BM25 runs over an inverted index (a map from each token to the documents that contain it), so it is fast and matches tokens exactly. Search for <code class=\"language-text\">T1_US_FNAL</code> and BM25 finds the documents that literally contain <code class=\"language-text\">T1_US_FNAL</code>. Robertson and Zaragoza give the full statistical account in <a href=\"https://www.staff.city.ac.uk/~sbrp622/papers/foundations_bm25_review.pdf\"><em>The Probabilistic Relevance Framework: BM25 and Beyond</em></a>.</p>\n<p>Dense retrieval works differently. You embed each chunk into a vector with a model like <a href=\"https://www.sbert.net/\">Sentence-Transformers</a>, embed the query the same way, and return the chunks whose vectors sit closest to the query’s. Closeness in that space means <em>similar meaning</em>. A query about “moving data between sites” can match a chunk that says “transfers between grid tiers” with no shared words.</p>\n<p>That is the whole appeal, and it is also the weakness. The model is trained to collapse surface differences, so it collapses the differences you care about too. <code class=\"language-text\">T1_US_FNAL</code> and <code class=\"language-text\">T2_DE_DESY</code> are near-identical to an embedding model and completely different to an operator.</p>\n<p>Neither retriever is wrong. They are good at different queries, and real questions are a mix. So run both and combine them.</p>\n<h2>Why you can’t just add the scores</h2>\n<p>The naive version is to run both searches, add the scores, and sort. It does not work, because the scores are not comparable. BM25 produces an unbounded positive number that depends on term statistics. Cosine similarity lives in a fixed range near zero to one. Add them, and BM25 either dominates or vanishes depending on the query. You end up hand-tuning a weight per corpus, and it drifts the moment the data changes.</p>\n<p>There are two ways out. Different engines default to different ones, so it is worth knowing both.</p>\n<p><strong>Score normalization</strong> rescales each retriever’s scores onto a common range, then combines them. <a href=\"https://docs.opensearch.org/latest/search-plugins/search-pipelines/normalization-processor/\">OpenSearch’s normalization processor</a> does this inside a hybrid query. It supports min-max and L2 normalization, then combines the results with an arithmetic, geometric, or harmonic mean. It works, but you still choose a normalization method and a combination weight. Min-max in particular is sensitive to a single outlier score stretching the range.</p>\n<p><strong>Rank fusion</strong> throws the scores away entirely and combines the lists by rank, meaning each document’s position in each list. That is what the next section covers.</p>\n<h2>Reciprocal Rank Fusion (RRF)</h2>\n<p><a href=\"https://plg.uwaterloo.ca/~gvcormack/cormacksigir09-rrf.pdf\">Reciprocal Rank Fusion</a> (RRF) is the method that quietly ended up inside most hybrid search implementations, including Elasticsearch and newer OpenSearch. The formula fits on one line: a document’s fused score is the sum of <code class=\"language-text\">1 / (k + rank)</code> over every ranked list it appears in.</p>\n<p>The function below implements it. Each input is one retriever’s output, ordered best first.</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\"><span class=\"token keyword\">def</span> <span class=\"token function\">reciprocal_rank_fusion</span><span class=\"token punctuation\">(</span>ranked_lists<span class=\"token punctuation\">,</span> k<span class=\"token operator\">=</span><span class=\"token number\">60</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n    <span class=\"token triple-quoted-string string\">\"\"\"ranked_lists: a list of retriever outputs, each an ordered list of doc ids.\"\"\"</span>\n    scores <span class=\"token operator\">=</span> <span class=\"token punctuation\">{</span><span class=\"token punctuation\">}</span>\n    <span class=\"token keyword\">for</span> ranking <span class=\"token keyword\">in</span> ranked_lists<span class=\"token punctuation\">:</span>\n        <span class=\"token keyword\">for</span> rank<span class=\"token punctuation\">,</span> doc_id <span class=\"token keyword\">in</span> <span class=\"token builtin\">enumerate</span><span class=\"token punctuation\">(</span>ranking<span class=\"token punctuation\">,</span> start<span class=\"token operator\">=</span><span class=\"token number\">1</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n            scores<span class=\"token punctuation\">[</span>doc_id<span class=\"token punctuation\">]</span> <span class=\"token operator\">=</span> scores<span class=\"token punctuation\">.</span>get<span class=\"token punctuation\">(</span>doc_id<span class=\"token punctuation\">,</span> <span class=\"token number\">0</span><span class=\"token punctuation\">)</span> <span class=\"token operator\">+</span> <span class=\"token number\">1</span> <span class=\"token operator\">/</span> <span class=\"token punctuation\">(</span>k <span class=\"token operator\">+</span> rank<span class=\"token punctuation\">)</span>\n    <span class=\"token keyword\">return</span> <span class=\"token builtin\">sorted</span><span class=\"token punctuation\">(</span>scores<span class=\"token punctuation\">,</span> key<span class=\"token operator\">=</span>scores<span class=\"token punctuation\">.</span>get<span class=\"token punctuation\">,</span> reverse<span class=\"token operator\">=</span><span class=\"token boolean\">True</span><span class=\"token punctuation\">)</span>\n\nfused <span class=\"token operator\">=</span> reciprocal_rank_fusion<span class=\"token punctuation\">(</span><span class=\"token punctuation\">[</span>bm25_hits<span class=\"token punctuation\">,</span> vector_hits<span class=\"token punctuation\">]</span><span class=\"token punctuation\">)</span>\ntop_k <span class=\"token operator\">=</span> fused<span class=\"token punctuation\">[</span><span class=\"token punctuation\">:</span><span class=\"token number\">8</span><span class=\"token punctuation\">]</span></code></pre></div>\n<p>That is the whole algorithm. Notice what is missing: no scores from either retriever, no normalization, no per-corpus weight. A document only needs its <em>position</em> in each list. For example, a chunk ranked first by BM25 and eighth by vector search scores <code class=\"language-text\">1/61 + 1/68</code>.</p>\n<p>A chunk that both retrievers rank near the top climbs to the top of the fused list, which is exactly the behavior you want. Agreement between two independent methods is a strong signal.</p>\n<p>The <code class=\"language-text\">k = 60</code> is not arbitrary. In the <a href=\"https://plg.uwaterloo.ca/~gvcormack/cormacksigir09-rrf.pdf\">original 2009 paper by Cormack, Clarke, and Grossman</a>, 60 was the constant that scored best across their benchmarks, and it has been the default everywhere since. Its job is to flatten the curve so the gap between rank 1 and rank 2 is not enormous. A larger <code class=\"language-text\">k</code> makes the fusion more forgiving of exact position. A smaller one rewards the very top ranks more heavily.</p>\n<p>I have never needed to move it. The paper’s headline result was a parameter-free formula beating trained learning-to-rank methods, which is a good reminder that the simple thing is often enough.</p>\n<p>The diagram below shows the pipeline end to end.</p>\n<p><img src=\"/8be0ae2ebc57fee6af18ca3485348132/hybrid-retrieval.svg\" alt=\"Hybrid retrieval pipeline: a query fans out to a BM25 lexical search and a dense vector search running in parallel, both feed into a fusion step using RRF or score normalization, and the fused top-k is reranked before going to the LLM\"></p>\n<h2>Fuse in the search engine, not in Python</h2>\n<p>The code above is the right way to <em>understand</em> RRF. It is not always the right way to ship it. Pulling a few hundred candidates from each retriever across the network just to fuse them in your app adds a round trip and moves ranking logic away from the data. If your search engine already does hybrid queries, let it.</p>\n<p>In <a href=\"https://docs.opensearch.org/latest/vector-search/ai-search/hybrid-search/index/\">OpenSearch</a>, you register a search pipeline once. After that, a single <code class=\"language-text\">hybrid</code> query carries both a lexical clause and a k-NN (nearest-neighbor vector) clause. The pipeline normalizes and combines the two result sets server-side and hands you one ranked list. Elasticsearch exposes RRF directly as a retriever.</p>\n<p>The Python version still earns its place in two cases:</p>\n<ul>\n<li>You are combining more than two sources, say lexical, dense, and a metadata filter from a different store.</li>\n<li>You want to unit-test the ranking without standing up a cluster.</li>\n</ul>\n<h2>Reranking: the step people skip</h2>\n<p>Fusion gives you a good <em>candidate set</em>, not a final order. RRF knows only about ranks. It has never looked at whether a chunk actually answers the question.</p>\n<p>A <a href=\"https://www.sbert.net/examples/applications/retrieve_rerank/README.html\">cross-encoder reranker</a> does look. It is a model that takes the query and one candidate chunk <em>together</em> and scores their relevance directly. Vector search, by contrast, compares two vectors that were embedded separately and never saw each other.</p>\n<p>The code below scores the top 30 fused candidates with a cross-encoder, re-sorts them, and keeps the best 5 as context.</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\"><span class=\"token keyword\">from</span> sentence_transformers <span class=\"token keyword\">import</span> CrossEncoder\n\nreranker <span class=\"token operator\">=</span> CrossEncoder<span class=\"token punctuation\">(</span><span class=\"token string\">\"cross-encoder/ms-marco-MiniLM-L-6-v2\"</span><span class=\"token punctuation\">)</span>\n\npairs <span class=\"token operator\">=</span> <span class=\"token punctuation\">[</span><span class=\"token punctuation\">(</span>query<span class=\"token punctuation\">,</span> doc_text<span class=\"token punctuation\">[</span>doc_id<span class=\"token punctuation\">]</span><span class=\"token punctuation\">)</span> <span class=\"token keyword\">for</span> doc_id <span class=\"token keyword\">in</span> fused<span class=\"token punctuation\">[</span><span class=\"token punctuation\">:</span><span class=\"token number\">30</span><span class=\"token punctuation\">]</span><span class=\"token punctuation\">]</span>\nrerank_scores <span class=\"token operator\">=</span> reranker<span class=\"token punctuation\">.</span>predict<span class=\"token punctuation\">(</span>pairs<span class=\"token punctuation\">)</span>\n\nreranked <span class=\"token operator\">=</span> <span class=\"token punctuation\">[</span>doc <span class=\"token keyword\">for</span> _<span class=\"token punctuation\">,</span> doc <span class=\"token keyword\">in</span> <span class=\"token builtin\">sorted</span><span class=\"token punctuation\">(</span><span class=\"token builtin\">zip</span><span class=\"token punctuation\">(</span>rerank_scores<span class=\"token punctuation\">,</span> fused<span class=\"token punctuation\">[</span><span class=\"token punctuation\">:</span><span class=\"token number\">30</span><span class=\"token punctuation\">]</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">,</span> reverse<span class=\"token operator\">=</span><span class=\"token boolean\">True</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">]</span>\nfinal_context <span class=\"token operator\">=</span> reranked<span class=\"token punctuation\">[</span><span class=\"token punctuation\">:</span><span class=\"token number\">5</span><span class=\"token punctuation\">]</span></code></pre></div>\n<p>The pattern is <em>retrieve wide, rerank narrow</em>. Fusion hands you thirty candidates cheaply. The cross-encoder does the expensive, careful read on just those thirty, and only the top five reach the model’s context.</p>\n<p>It costs one extra model call. It is usually the single biggest quality jump after adding lexical search, because it is the first stage that reads the query and the chunk as one thing.</p>\n<h2>Where it breaks</h2>\n<p><strong>Tokenization mismatch on the lexical side.</strong> BM25 only matches what your analyzer (the step that splits text into tokens) produced. If the indexer splits <code class=\"language-text\">T1_US_FNAL</code> into <code class=\"language-text\">t1</code>, <code class=\"language-text\">us</code>, <code class=\"language-text\">fnal</code> but the query analyzer keeps it whole, they never meet. Your keyword branch then silently returns nothing for the exact query it was supposed to save. Index-time and query-time analysis have to agree, and that deserves an explicit test.</p>\n<p><strong>Fusing lists of wildly different lengths.</strong> If BM25 returns four hits and vector search returns two hundred, RRF still works, but the fused ranking is dominated by whichever list is longer in its tail. Cap both retrievers at the same depth before fusing. The top 50 to 100 from each is plenty, since anything past that rarely survives reranking anyway.</p>\n<p><strong>Blaming retrieval when the problem is chunking.</strong> Hybrid search cannot retrieve a fact that got split across two chunks so that neither holds the whole answer. When a query fails, check that the answer exists intact in a single chunk before you tune fusion weights. This sits upstream of everything here, and it is where I have wasted the most time.</p>\n<p>Keeping that index fresh is its own problem. I wrote about detecting changed files and re-embedding only what moved in <a href=\"/blog/2026-06-29-incremental-rag-indexing/\">incremental indexing for RAG pipelines</a>.</p>\n<p><strong>Latency creep.</strong> Two retrievers plus a reranker is more work than one vector lookup. The retrievers run in parallel, so that part is close to free, but the cross-encoder works through the candidates sequentially. If you stream answers to a user, budget for it, and try reranking fewer candidates before you consider dropping the step.</p>\n<h2>What I would do differently</h2>\n<p>I reached for a reranker too late. On early RAG systems, my instinct was to tune the retrievers: better embeddings, a bigger candidate pool, fiddling with weights. The cheapest large gain was bolting a cross-encoder onto the end and letting the retrievers stay rough. Retrieve wide and rerank narrow is a better default than retrieve precise, and it is less work.</p>\n<p>I also used to normalize and weight scores by hand, one setting per corpus, and then relearn the setting every time the data shifted. RRF made that whole category of tuning disappear. Ranks are stable in a way raw scores are not, and giving up the scores cost me nothing I could measure.</p>\n<h2>Where this runs</h2>\n<p>Hybrid retrieval is the search layer under <a href=\"/project/archi/\">Archi</a>, the RAG copilot I worked on for CMS computing operations at CERN. Lexical search catches the site names, error codes, and JIRA ticket IDs that operators actually type. Vector search handles the “how did we handle something like this before” questions that never share vocabulary with the docs.</p>\n<p>The same two-retriever pattern showed up in <a href=\"/project/argusa-ai-challenge-2025/\">ARGRAG</a>, the enterprise-document RAG system my teammate and I built for the Argusa AI Challenge. There the corpus was a pile of PDFs, spreadsheets, and source files, and no single retrieval method covered all of it.</p>\n<p>One retriever is a bet that your queries will look like your training data. Hybrid search is what you do when you know they won’t.</p>\n<p><em>Image credit: hybrid search diagrams by M. Hassan Ahmed, created for this post, released under CC0 (public domain).</em></p>","frontmatter":{"title":"Hybrid Search for RAG: BM25 + Vectors","date":"2026-07-02T00:00:00.000Z","description":"Vector search alone misses exact IDs and error codes. Here's how to combine BM25 keyword search with dense retrieval, fuse the rankings with RRF, and rerank.","thumbnail":{"childImageSharp":{"fluid":{"base64":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAABQAAAALCAIAAADwazoUAAAACXBIWXMAAAsSAAALEgHS3X78AAAB40lEQVQoz02PWY/aMBRG89LJQpwEKAnxTUycfSGELANMYKCqAKmsbacPfen//xs1g0pHOrqyfb/ja3MmrAy8pP52Wv8sqh/Z5DwuLvn0ytZ5cWHbsvk1Kb/fDosry2STS9m8jehXZnGSBoJqagalcZlM22S6TIo2mryk5Wpcr2+1emXnYb7I6/Wk2bhJ7caVplNmcZ0u6XRHkmbzCD8hS1DJE7J5xUY9Imnkk4xZS+5SQYEnhHkF2IJHwFrM4jqa/Q/SN5zuwAGgaeAF1PFGxHML21oQ+8WlSwNnPd3TPrsym/eucJJqMTqqJSDAtpdGUVsmdR6NwMqns/bIvroGs0njXTE++cGybwSiAneLyfBAQDiP/WkaWhgC6j/vzvX5NNvtXWdumlUW79N4z8vGI89JCr4jIqz2YFHE1TikFo6CbHa6bv78nh8PYbCwYZ7F2zTaqX0qouFd4UTFvCMgE2k49UehSxwLIj9ujufi7dwcvvneDPCz6yw957WjgqgM7wrHrnnAy8Pbg4nlAKaAy/ZLfTpU7cYyazBrAgtdH/PyQETmPc9k44GADEkxqG2z2Q5xAi+erbaetwBoiD0HXL07//OcIOsf4TsDSbU1PejjhNXuIOgN494w7BoB+yQb+zH8F4mGUab5x2CsAAAAAElFTkSuQmCC","aspectRatio":1.899441340782123,"src":"/static/dd2bcaa9601c27b33c0e7510855e8462/40a76/hero.png","srcSet":"/static/dd2bcaa9601c27b33c0e7510855e8462/c972b/hero.png 340w,\n/static/dd2bcaa9601c27b33c0e7510855e8462/27625/hero.png 680w,\n/static/dd2bcaa9601c27b33c0e7510855e8462/40a76/hero.png 1360w,\n/static/dd2bcaa9601c27b33c0e7510855e8462/ed396/hero.png 2000w","sizes":"(max-width: 1360px) 100vw, 1360px"}}}}}},"pageContext":{"slug":"/2026-07-02-hybrid-search-rag-bm25-vectors/","previous":"blog/2026-07-03-argocd-sync-waves-ordering-rollout/","next":"blog/2026-07-01-opensearch-workflow-monitoring/"}},"staticQueryHashes":["32046230"]}