Reciprocal Rank Fusion for Hybrid Search
Reciprocal Rank Fusion merges two ranked lists without tuning score weights. The RRF formula, why the k constant matters, and how to run it in OpenSearch.
When I wrote about hybrid search for RAG, I ran a keyword search and a vector search side by side, then hand-waved the last step: “fuse the two ranked lists.” That step is where most hybrid setups quietly go wrong. BM25 (the standard keyword-ranking function) hands you one ordered list, a vector index hands you another, and you need a single ranking out of them. The obvious move is to add the scores and sort. That produces worse results than either retriever alone, and it took me a while to understand why.
This post is for engineers who already run two retrievers and need to combine their outputs without babysitting a weight parameter for every corpus. Reciprocal Rank Fusion (RRF) is the method that made this problem go away for me on Archi, the retrieval copilot I worked on for CMS computing operations at CERN. It is close to the simplest thing that could work, which is exactly why it holds up. I cover why adding scores fails, the RRF formula with a worked example, what its one constant does, how to run it in OpenSearch, and where it breaks.
Why adding the two scores fails
The two retrievers score on scales that have nothing to do with each other:
- BM25. Run a keyword query in OpenSearch and BM25 hands back a relevance score. It is an unbounded positive number built from term frequencies and document lengths. A good match for a rare term might score 14; another query might top out at 3.
- Vector search. Run the same question against a vector index and you get cosine similarity (how closely two embedding vectors point in the same direction), which lives in a fixed range around 0 to 1.
Put those two numbers in the same sum and the result depends on the term statistics of each query. On some queries the BM25 score swamps the cosine score; on others it vanishes under it. So you reach for a weight, 0.6 * bm25 + 0.4 * cosine. Now you are tuning that weight per corpus, and it drifts the moment your documents change. I did this once. It felt like tuning a radio by feel, and every reindex knocked it back out.
The deeper problem is that a score only means something inside its own list. BM25’s 14 and cosine’s 0.82 answer different questions on different scales, so comparing them directly is a category error. What both lists do agree on is order: which document each retriever thinks is best, second-best, and so on. Rank is the common currency. RRF is built entirely on rank and throws the raw scores away.
The RRF formula
For a document d, add up one term per retriever. Each term is one divided by a constant k plus that document’s rank in that retriever’s list:
score(d) = Σ 1 / (k + rank_r(d))
rrank_r(d) is 1 for the top result of retriever r, 2 for the next, and so on. If a document does not appear in a retriever’s list at all, that retriever contributes nothing for it. The constant k is a smoothing term. The original paper uses k = 60, and almost everyone has kept that default since.
That is the whole algorithm. It comes from a 2009 SIGIR paper by Cormack, Clarke, and Büttcher, Reciprocal Rank Fusion Outperforms Condorcet and Individual Rank Learning Methods. Their finding was that this one-line rule beat considerably more elaborate rank-aggregation schemes. Seventeen years later it is the default fusion step in most hybrid search stacks, which tells you something about how well the fancier methods generalized.
In Python it is short enough to read in one sitting. The function walks each ranked list, adds 1 / (k + rank) to each document’s running total, and sorts by the totals:
def rrf(ranked_lists, k=60):
"""ranked_lists: iterable of lists of doc ids, each already sorted best-first."""
scores = {}
for ranked in ranked_lists:
for rank, doc_id in enumerate(ranked, start=1):
scores[doc_id] = scores.get(doc_id, 0.0) + 1.0 / (k + rank)
return sorted(scores, key=scores.get, reverse=True)There is no score normalization, no per-corpus weight, and no assumption that the two retrievers produce comparable numbers. You hand it ranked lists of IDs, and it hands back one fused list.
A worked example
Say a shifter (an operator on duty in CMS computing operations) asks Archi about a transfer failure. The two retrievers come back with these top-four lists.
BM25 ranks D7, D3, D9, D2. The vector index ranks D3, D5, D7, D8. With k = 60:
- D3 sits at rank 2 in BM25 and rank 1 in vectors:
1/62 + 1/61 = 0.03252. - D7 is rank 1 in BM25 and rank 3 in vectors:
1/61 + 1/63 = 0.03226. - D5 appears only in the vector list at rank 2:
1/62 = 0.01613. - D9 appears only in BM25 at rank 3:
1/63 = 0.01587. - D2 and D8 each appear once at rank 4:
1/64 = 0.01563.
D3 wins not because either retriever loved it most, but because both retrievers put it near the top. D7 comes second for the same reason. Every document that showed up in only one list falls below the two that both retrievers agreed on.
That agreement-wins behavior is the entire point. It is why RRF tends to be more stable than either retriever alone: a document has to be plausible on two independent grounds to reach the top.
Notice the numbers are tiny and clustered, with everything a hair above 1/64. RRF scores are not calibrated probabilities, so do not read them as confidence. They exist only to induce an order.
What the k constant controls
k sets how steeply rank matters. It sits in the denominator next to the rank. A small k makes the gap between rank 1 and rank 2 large, and a large k flattens the whole list toward equal weight. Compare the two ends:
- With
k = 0, the top result contributes1/1 = 1.0and the second1/2 = 0.5. Being first is worth twice being second, and a single retriever’s top pick can dominate. - With
k = 60, rank 1 contributes1/61 ≈ 0.0164and rank 2 contributes1/62 ≈ 0.0161, nearly the same.
A large k says “I trust that these documents are all roughly relevant, but I do not fully trust the exact order within each list.” That is usually the right stance for BM25 and dense retrieval, whose orderings are noisy past the first few hits.
The 60 default is a reasonable prior, not a law. If one retriever’s top-1 is almost always the right answer, a smaller k lets it carry more weight. Treat k as the one knob worth sweeping, and sweep it against a real metric like recall@k or MRR rather than by eye.
Running RRF inside OpenSearch
You do not have to implement RRF yourself if your engine has it. This matters to me because Archi’s retrieval sits on OpenSearch. Doing the fusion inside a search pipeline keeps a whole round trip and a chunk of glue code out of the backend.
OpenSearch ships two ways to combine hybrid results, and they map exactly onto the score-versus-rank split from earlier:
- The normalization processor is score-based. It rescales each clause’s scores to a shared range and combines them with a mean.
- The score-ranker processor is rank-based. It was added in OpenSearch 2.19 and implements RRF directly.
You register the score-ranker once as a pipeline. The rank_constant field is the k from the formula:
PUT /_search/pipeline/rrf-pipeline
{
"phase_results_processors": [
{
"score-ranker-processor": {
"combination": {
"technique": "rrf",
"rank_constant": 60
}
}
}
]
}Then point a hybrid query at that pipeline, and OpenSearch fuses the sub-query rankings with RRF before it returns hits. Elasticsearch exposes the same idea through its rrf retriever with a rank_constant parameter. On an engine without native support, running the ten-line function above in your service is genuinely fine. The fusion is cheap next to the two searches feeding it.
Where RRF breaks
RRF has sharp edges. Ignoring them leads to the same “why is retrieval worse now” debugging session I have already been through.
It only sees the top of each list. In practice you fuse the top 50 or 100 from each retriever, not the whole index. A document ranked 300th by BM25 and 4th by vectors contributes only its vector rank, because it fell outside the BM25 window entirely. Wider windows improve recall at the cost of latency and memory, and this cutoff, not the formula, is where most “RRF missed the obvious answer” cases actually come from.
Rank throws away margin. Because RRF ignores scores, it cannot tell a runaway top hit from a near-tie. If BM25’s rank 1 scored 40 and its rank 2 scored 2, RRF still treats them as adjacent ranks one apart. Usually that stability is what you want, but occasionally a retriever is genuinely, hugely confident and you have discarded that signal. If your top hits are often decisive, the score-based normalization path may serve you better.
More retrievers is not automatically better. Every list you fuse in gets an equal vote, so a weak third retriever can pull mediocre documents up simply by ranking them at all. I would add a retriever only when I can show, on an eval set, that it recovers queries the other two miss, not on the theory that more sources must help.
Ties are real and common. Documents that appear once at the same rank across different lists get identical scores, as D2 and D8 did above. Decide the tiebreak deliberately, with a stable sort by original rank or a fallback score, rather than letting insertion order decide it for you.
What I would do differently
The first time I reached for this, I over-thought the fusion and under-thought everything around it. My advice, in order:
- Start with
k = 60and do not touch it until you have an evaluation set to move it against. - Get the retrieval windows right before you tune anything. A too-small top-
nper retriever hurts far more than a suboptimalk. - Keep RRF in its lane. It merges rankings; it does not judge relevance. It will happily fuse two bad lists into one bad list, so the retrievers underneath still have to be good.
On Archi I follow the fusion with a cross-encoder reranker over the fused top results. A cross-encoder reads the query and each candidate together, and that is where the real score-based judgment happens. RRF gets a strong candidate set into the reranker cheaply, and the reranker does the careful part.
That division of labor is the takeaway. RRF is the cheap, corpus-agnostic step that turns “two lists on incompatible scales” into “one sensible ranking” with a single constant and no per-corpus tuning. It is not the whole retrieval stack. It is the join in the middle that I stopped having to think about once I understood it, which is the best thing you can say about a piece of infrastructure.
If you want the layers on either side of it, the hybrid search post covers the two retrievers that feed RRF, and the reranking post covers what comes after. Together they are most of how retrieval works on Archi.