{"componentChunkName":"component---src-templates-blog-post-js","path":"/blog/2026-09-02-reciprocal-rank-fusion-hybrid-search/","result":{"data":{"site":{"siteMetadata":{"title":"M.Hassan Ahmed","author":"Hassan11196"}},"markdownRemark":{"id":"09c22b50-9382-5b6f-bc0f-83c0ecc02e8f","excerpt":"When I wrote about hybrid search for RAG, I ran a keyword search and a vector search side by side, then hand-waved the last step: “fuse the two ranked lists…","html":"<p>When I <a href=\"/blog/2026-07-02-hybrid-search-rag-bm25-vectors/\">wrote about hybrid search for RAG</a>, I ran a keyword search and a vector search side by side, then hand-waved the last step: “fuse the two ranked lists.” That step is where most hybrid setups quietly go wrong. BM25 (the standard keyword-ranking function) hands you one ordered list, a vector index hands you another, and you need a single ranking out of them. The obvious move is to add the scores and sort. That produces worse results than either retriever alone, and it took me a while to understand why.</p>\n<p>This post is for engineers who already run two retrievers and need to combine their outputs without babysitting a weight parameter for every corpus. Reciprocal Rank Fusion (RRF) is the method that made this problem go away for me on <a href=\"/project/archi/\">Archi</a>, the retrieval copilot I worked on for CMS computing operations at CERN. It is close to the simplest thing that could work, which is exactly why it holds up. I cover why adding scores fails, the RRF formula with a worked example, what its one constant does, how to run it in OpenSearch, and where it breaks.</p>\n<h2>Why adding the two scores fails</h2>\n<p>The two retrievers score on scales that have nothing to do with each other:</p>\n<ul>\n<li><strong>BM25.</strong> Run a keyword query in <a href=\"https://opensearch.org/\">OpenSearch</a> and BM25 hands back a relevance score. It is an unbounded positive number built from term frequencies and document lengths. A good match for a rare term might score 14; another query might top out at 3.</li>\n<li><strong>Vector search.</strong> Run the same question against a vector index and you get <a href=\"https://en.wikipedia.org/wiki/Cosine_similarity\">cosine similarity</a> (how closely two embedding vectors point in the same direction), which lives in a fixed range around 0 to 1.</li>\n</ul>\n<p>Put those two numbers in the same sum and the result depends on the term statistics of each query. On some queries the BM25 score swamps the cosine score; on others it vanishes under it. So you reach for a weight, <code class=\"language-text\">0.6 * bm25 + 0.4 * cosine</code>. Now you are tuning that weight per corpus, and it drifts the moment your documents change. I did this once. It felt like tuning a radio by feel, and every reindex knocked it back out.</p>\n<p>The deeper problem is that a score only means something <em>inside its own list</em>. BM25’s 14 and cosine’s 0.82 answer different questions on different scales, so comparing them directly is a category error. What both lists do agree on is <strong>order</strong>: which document each retriever thinks is best, second-best, and so on. Rank is the common currency. RRF is built entirely on rank and throws the raw scores away.</p>\n<h2>The RRF formula</h2>\n<p>For a document <code class=\"language-text\">d</code>, add up one term per retriever. Each term is one divided by a constant <code class=\"language-text\">k</code> plus that document’s rank in that retriever’s list:</p>\n<div class=\"gatsby-highlight\" data-language=\"text\"><pre class=\"language-text\"><code class=\"language-text\">score(d) = Σ  1 / (k + rank_r(d))\n          r</code></pre></div>\n<p><code class=\"language-text\">rank_r(d)</code> is 1 for the top result of retriever <code class=\"language-text\">r</code>, 2 for the next, and so on. If a document does not appear in a retriever’s list at all, that retriever contributes nothing for it. The constant <code class=\"language-text\">k</code> is a smoothing term. The original paper uses <code class=\"language-text\">k = 60</code>, and almost everyone has kept that default since.</p>\n<p>That is the whole algorithm. It comes from a 2009 SIGIR paper by Cormack, Clarke, and Büttcher, <a href=\"https://cormack.uwaterloo.ca/cormacksigir09-rrf.pdf\"><em>Reciprocal Rank Fusion Outperforms Condorcet and Individual Rank Learning Methods</em></a>. Their finding was that this one-line rule beat considerably more elaborate rank-aggregation schemes. Seventeen years later it is the default fusion step in most hybrid search stacks, which tells you something about how well the fancier methods generalized.</p>\n<p>In Python it is short enough to read in one sitting. The function walks each ranked list, adds <code class=\"language-text\">1 / (k + rank)</code> to each document’s running total, and sorts by the totals:</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\"><span class=\"token keyword\">def</span> <span class=\"token function\">rrf</span><span class=\"token punctuation\">(</span>ranked_lists<span class=\"token punctuation\">,</span> k<span class=\"token operator\">=</span><span class=\"token number\">60</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n    <span class=\"token triple-quoted-string string\">\"\"\"ranked_lists: iterable of lists of doc ids, each already sorted best-first.\"\"\"</span>\n    scores <span class=\"token operator\">=</span> <span class=\"token punctuation\">{</span><span class=\"token punctuation\">}</span>\n    <span class=\"token keyword\">for</span> ranked <span class=\"token keyword\">in</span> ranked_lists<span class=\"token punctuation\">:</span>\n        <span class=\"token keyword\">for</span> rank<span class=\"token punctuation\">,</span> doc_id <span class=\"token keyword\">in</span> <span class=\"token builtin\">enumerate</span><span class=\"token punctuation\">(</span>ranked<span class=\"token punctuation\">,</span> start<span class=\"token operator\">=</span><span class=\"token number\">1</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n            scores<span class=\"token punctuation\">[</span>doc_id<span class=\"token punctuation\">]</span> <span class=\"token operator\">=</span> scores<span class=\"token punctuation\">.</span>get<span class=\"token punctuation\">(</span>doc_id<span class=\"token punctuation\">,</span> <span class=\"token number\">0.0</span><span class=\"token punctuation\">)</span> <span class=\"token operator\">+</span> <span class=\"token number\">1.0</span> <span class=\"token operator\">/</span> <span class=\"token punctuation\">(</span>k <span class=\"token operator\">+</span> rank<span class=\"token punctuation\">)</span>\n    <span class=\"token keyword\">return</span> <span class=\"token builtin\">sorted</span><span class=\"token punctuation\">(</span>scores<span class=\"token punctuation\">,</span> key<span class=\"token operator\">=</span>scores<span class=\"token punctuation\">.</span>get<span class=\"token punctuation\">,</span> reverse<span class=\"token operator\">=</span><span class=\"token boolean\">True</span><span class=\"token punctuation\">)</span></code></pre></div>\n<p>There is no score normalization, no per-corpus weight, and no assumption that the two retrievers produce comparable numbers. You hand it ranked lists of IDs, and it hands back one fused list.</p>\n<h2>A worked example</h2>\n<p>Say a shifter (an operator on duty in CMS computing operations) asks Archi about a transfer failure. The two retrievers come back with these top-four lists.</p>\n<p><img src=\"/e31398c82f7808c55d0d346c29480b8e/rrf-fusion.svg\" alt=\"Two ranked lists, one from BM25 and one from a vector index, feed into Reciprocal Rank Fusion. Each document earns 1/(k+rank) from each list it appears in, and the sums are added. Documents D3 and D7 rise to the top of the fused list because both retrievers rank them highly, while documents that appear in only one list fall below them.\"></p>\n<p>BM25 ranks <code class=\"language-text\">D7, D3, D9, D2</code>. The vector index ranks <code class=\"language-text\">D3, D5, D7, D8</code>. With <code class=\"language-text\">k = 60</code>:</p>\n<ul>\n<li><strong>D3</strong> sits at rank 2 in BM25 and rank 1 in vectors: <code class=\"language-text\">1/62 + 1/61 = 0.03252</code>.</li>\n<li><strong>D7</strong> is rank 1 in BM25 and rank 3 in vectors: <code class=\"language-text\">1/61 + 1/63 = 0.03226</code>.</li>\n<li><strong>D5</strong> appears only in the vector list at rank 2: <code class=\"language-text\">1/62 = 0.01613</code>.</li>\n<li><strong>D9</strong> appears only in BM25 at rank 3: <code class=\"language-text\">1/63 = 0.01587</code>.</li>\n<li><strong>D2</strong> and <strong>D8</strong> each appear once at rank 4: <code class=\"language-text\">1/64 = 0.01563</code>.</li>\n</ul>\n<p>D3 wins not because either retriever loved it most, but because <em>both</em> retrievers put it near the top. D7 comes second for the same reason. Every document that showed up in only one list falls below the two that both retrievers agreed on.</p>\n<p>That agreement-wins behavior is the entire point. It is why RRF tends to be more stable than either retriever alone: a document has to be plausible on two independent grounds to reach the top.</p>\n<p>Notice the numbers are tiny and clustered, with everything a hair above <code class=\"language-text\">1/64</code>. RRF scores are not calibrated probabilities, so do not read them as confidence. They exist only to induce an order.</p>\n<h2>What the k constant controls</h2>\n<p><code class=\"language-text\">k</code> sets how steeply rank matters. It sits in the denominator next to the rank. A small <code class=\"language-text\">k</code> makes the gap between rank 1 and rank 2 large, and a large <code class=\"language-text\">k</code> flattens the whole list toward equal weight. Compare the two ends:</p>\n<ul>\n<li><strong>With <code class=\"language-text\">k = 0</code>,</strong> the top result contributes <code class=\"language-text\">1/1 = 1.0</code> and the second <code class=\"language-text\">1/2 = 0.5</code>. Being first is worth twice being second, and a single retriever’s top pick can dominate.</li>\n<li><strong>With <code class=\"language-text\">k = 60</code>,</strong> rank 1 contributes <code class=\"language-text\">1/61 ≈ 0.0164</code> and rank 2 contributes <code class=\"language-text\">1/62 ≈ 0.0161</code>, nearly the same.</li>\n</ul>\n<p>A large <code class=\"language-text\">k</code> says “I trust that these documents are all roughly relevant, but I do not fully trust the exact order within each list.” That is usually the right stance for BM25 and dense retrieval, whose orderings are noisy past the first few hits.</p>\n<p>The 60 default is a reasonable prior, not a law. If one retriever’s top-1 is almost always the right answer, a smaller <code class=\"language-text\">k</code> lets it carry more weight. Treat <code class=\"language-text\">k</code> as the one knob worth sweeping, and sweep it against a real metric like <a href=\"/blog/2026-07-04-measuring-rag-retrieval-quality/\">recall@k or MRR</a> rather than by eye.</p>\n<h2>Running RRF inside OpenSearch</h2>\n<p>You do not have to implement RRF yourself if your engine has it. This matters to me because Archi’s retrieval sits on OpenSearch. Doing the fusion inside a <a href=\"https://docs.opensearch.org/latest/search-plugins/search-pipelines/index/\">search pipeline</a> keeps a whole round trip and a chunk of glue code out of the backend.</p>\n<p>OpenSearch ships two ways to combine hybrid results, and they map exactly onto the score-versus-rank split from earlier:</p>\n<ul>\n<li><strong>The <a href=\"https://docs.opensearch.org/latest/search-plugins/search-pipelines/normalization-processor/\">normalization processor</a> is score-based.</strong> It rescales each clause’s scores to a shared range and combines them with a mean.</li>\n<li><strong>The <a href=\"https://docs.opensearch.org/latest/search-plugins/search-pipelines/score-ranker-processor/\">score-ranker processor</a> is rank-based.</strong> It was added in OpenSearch 2.19 and implements RRF directly.</li>\n</ul>\n<p>You register the score-ranker once as a pipeline. The <code class=\"language-text\">rank_constant</code> field is the <code class=\"language-text\">k</code> from the formula:</p>\n<div class=\"gatsby-highlight\" data-language=\"json\"><pre class=\"language-json\"><code class=\"language-json\">PUT /_search/pipeline/rrf-pipeline\n<span class=\"token punctuation\">{</span>\n  <span class=\"token property\">\"phase_results_processors\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">[</span>\n    <span class=\"token punctuation\">{</span>\n      <span class=\"token property\">\"score-ranker-processor\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">{</span>\n        <span class=\"token property\">\"combination\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">{</span>\n          <span class=\"token property\">\"technique\"</span><span class=\"token operator\">:</span> <span class=\"token string\">\"rrf\"</span><span class=\"token punctuation\">,</span>\n          <span class=\"token property\">\"rank_constant\"</span><span class=\"token operator\">:</span> <span class=\"token number\">60</span>\n        <span class=\"token punctuation\">}</span>\n      <span class=\"token punctuation\">}</span>\n    <span class=\"token punctuation\">}</span>\n  <span class=\"token punctuation\">]</span>\n<span class=\"token punctuation\">}</span></code></pre></div>\n<p>Then point a hybrid query at that pipeline, and OpenSearch fuses the sub-query rankings with RRF before it returns hits. <a href=\"https://www.elastic.co/docs/reference/elasticsearch/rest-apis/retrievers#rrf-retriever\">Elasticsearch</a> exposes the same idea through its <code class=\"language-text\">rrf</code> retriever with a <code class=\"language-text\">rank_constant</code> parameter. On an engine without native support, running the ten-line function above in your service is genuinely fine. The fusion is cheap next to the two searches feeding it.</p>\n<h2>Where RRF breaks</h2>\n<p>RRF has sharp edges. Ignoring them leads to the same “why is retrieval worse now” debugging session I have already been through.</p>\n<p><strong>It only sees the top of each list.</strong> In practice you fuse the top 50 or 100 from each retriever, not the whole index. A document ranked 300th by BM25 and 4th by vectors contributes only its vector rank, because it fell outside the BM25 window entirely. Wider windows improve recall at the cost of latency and memory, and this cutoff, not the formula, is where most “RRF missed the obvious answer” cases actually come from.</p>\n<p><strong>Rank throws away margin.</strong> Because RRF ignores scores, it cannot tell a runaway top hit from a near-tie. If BM25’s rank 1 scored 40 and its rank 2 scored 2, RRF still treats them as adjacent ranks one apart. Usually that stability is what you want, but occasionally a retriever is genuinely, hugely confident and you have discarded that signal. If your top hits are often decisive, the score-based normalization path may serve you better.</p>\n<p><strong>More retrievers is not automatically better.</strong> Every list you fuse in gets an equal vote, so a weak third retriever can pull mediocre documents up simply by ranking them at all. I would add a retriever only when I can show, on an eval set, that it recovers queries the other two miss, not on the theory that more sources must help.</p>\n<p><strong>Ties are real and common.</strong> Documents that appear once at the same rank across different lists get identical scores, as D2 and D8 did above. Decide the tiebreak deliberately, with a stable sort by original rank or a fallback score, rather than letting insertion order decide it for you.</p>\n<h2>What I would do differently</h2>\n<p>The first time I reached for this, I over-thought the fusion and under-thought everything around it. My advice, in order:</p>\n<ol>\n<li><strong>Start with <code class=\"language-text\">k = 60</code></strong> and do not touch it until you have an evaluation set to move it against.</li>\n<li><strong>Get the retrieval windows right before you tune anything.</strong> A too-small top-<code class=\"language-text\">n</code> per retriever hurts far more than a suboptimal <code class=\"language-text\">k</code>.</li>\n<li><strong>Keep RRF in its lane.</strong> It merges rankings; it does not judge relevance. It will happily fuse two bad lists into one bad list, so the retrievers underneath still have to be good.</li>\n</ol>\n<p>On Archi I follow the fusion with a <a href=\"/blog/2026-07-08-cross-encoder-reranking-rag/\">cross-encoder reranker</a> over the fused top results. A cross-encoder reads the query and each candidate together, and that is where the real score-based judgment happens. RRF gets a strong candidate set into the reranker cheaply, and the reranker does the careful part.</p>\n<p>That division of labor is the takeaway. RRF is the cheap, corpus-agnostic step that turns “two lists on incompatible scales” into “one sensible ranking” with a single constant and no per-corpus tuning. It is not the whole retrieval stack. It is the join in the middle that I stopped having to think about once I understood it, which is the best thing you can say about a piece of infrastructure.</p>\n<p>If you want the layers on either side of it, the <a href=\"/blog/2026-07-02-hybrid-search-rag-bm25-vectors/\">hybrid search post</a> covers the two retrievers that feed RRF, and the <a href=\"/blog/2026-07-08-cross-encoder-reranking-rag/\">reranking post</a> covers what comes after. Together they are most of how retrieval works on <a href=\"/project/archi/\">Archi</a>.</p>","frontmatter":{"title":"Reciprocal Rank Fusion for Hybrid Search","date":"2026-09-02T00:00:00.000Z","description":"Reciprocal Rank Fusion merges two ranked lists without tuning score weights. The RRF formula, why the k constant matters, and how to run it in OpenSearch.","thumbnail":{"childImageSharp":{"fluid":{"base64":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAABQAAAALCAIAAADwazoUAAAACXBIWXMAAAsSAAALEgHS3X78AAACI0lEQVQoz02QvW/TQBiHvcSfZ9+d7xyffbZjOx9N4qSkSRqSNp+tUJUSqYpECUPUAgKpYUPdWNjYmJDYmBnKwsDEwMIfxqWFCumZfvo9d+/7SlmheE8tTJpJebbfn3R6o73uUW8waLamnYezbn/Y6ozb+424JDr3fUmGkSBnBjpJLXcHsmqpOSzvjuNaf6c1SeoHaXZYeTDehs0R9uuiI5qyFSowkgCOTBILIE1RvgidFNiFW2IDhf8R6XZk0QTRVJSFJZB0GLCkXWs/KjYnSTaq7B2Xdmcs6WK+i1ndQMGdLGqk1KFhnfpV6teEKUIhc+iUKM94mPlhxqOGHzZsVrNo0SSJcP7BEasjt4ycVEx399xWVgFTDdsjeUZcFzsEUlm3NcvTof/3WxwIRKJBT0NcKGILA4eSZjLTjodPvxy9/EGiQdp9frL51Zhcq8DRTFfWoKIjVSEW4AhFCIVAZ4pq5xRTVi1JBXngpL3N19H1T1qbpcdvJu9+V5fvcyp0Cv3Zi5vB2Sd2cha83eRPH3vPVt7VZWG4ml3ctBcfJA24wObukzm7XNppmfY63utzZzpUFOTEhwfrb53F5/x8Sa7W9umcnC/Jq1U4vRiuv7cWHyXVdA2LxW5Y8UKKPY/yKo8Y8XWLyQalLMY0UDQaFTJVd2xSYKycU7DLixb2tmNvfSwOw2XgatAHt7fZ5oYDxHkspgIXkkgkBvJNO1AMx8Tb/A8dqVfhOKMvqAAAAABJRU5ErkJggg==","aspectRatio":1.899441340782123,"src":"/static/03c69b7b1f3f10ef17416f7f382600ae/40a76/hero.png","srcSet":"/static/03c69b7b1f3f10ef17416f7f382600ae/c972b/hero.png 340w,\n/static/03c69b7b1f3f10ef17416f7f382600ae/27625/hero.png 680w,\n/static/03c69b7b1f3f10ef17416f7f382600ae/40a76/hero.png 1360w,\n/static/03c69b7b1f3f10ef17416f7f382600ae/ed396/hero.png 2000w","sizes":"(max-width: 1360px) 100vw, 1360px"}}}}}},"pageContext":{"slug":"/2026-09-02-reciprocal-rank-fusion-hybrid-search/","previous":"blog/2026-09-03-opensearch-ism-index-lifecycle/","next":"blog/2026-09-06-kubernetes-poddisruptionbudget-node-drains/"}},"staticQueryHashes":["32046230"]}