{"componentChunkName":"component---src-templates-blog-post-js","path":"/blog/2026-07-08-cross-encoder-reranking-rag/","result":{"data":{"site":{"siteMetadata":{"title":"M.Hassan Ahmed","author":"Hassan11196"}},"markdownRemark":{"id":"8ec3862c-86da-53e4-97ff-f1494db5c70d","excerpt":"An operator asks the copilot “what fixed the  transfer error last spring?” The chunk that answers it is in the index, and retrieval finds it. But it comes back…","html":"<p>An operator asks the copilot “what fixed the <code class=\"language-text\">T1_US_FNAL</code> transfer error last spring?” The chunk that answers it is in the index, and retrieval finds it. But it comes back at rank 8, and the prompt only carries the top 4 chunks into the model. The answer was retrieved, then thrown away before the model ever saw it.</p>\n<p>Reranking closes this gap. Your first-stage retriever is tuned to be fast and to not miss things, so it casts a wide net and orders the catch only roughly. A reranker takes that shortlist and puts the actually relevant chunk on top, so it survives the cut into the prompt.</p>\n<p>This post is for engineers who already have retrieval working, hybrid or not, and keep seeing the right document come back somewhere in the top 20 but not the top 3. It covers why the retrieval score is coarse, what a cross-encoder does differently, working code for both a local model and a hosted API, and the failure modes that show up once reranking is in the pipeline.</p>\n<h2>Why the retrieval score is coarse</h2>\n<p>Dense retrieval runs on a <a href=\"https://sbert.net/examples/sentence_transformer/applications/retrieve_rerank/README.html\">bi-encoder</a>, a model that encodes the query and each document separately. It embeds every document into a vector ahead of time, embeds the query at search time, and ranks documents by cosine similarity. The reason this scales is exactly the reason it loses precision: the query and the document never meet. Each is compressed into a fixed-length vector on its own, and at query time the model only compares two points in that space.</p>\n<p>That compression throws away information ranking needs. The document vector has to summarize the whole chunk before it knows what you are going to ask. Take a chunk that mentions <code class=\"language-text\">T1_US_FNAL</code> once, in passing, alongside four other sites. Its vector ends up mostly about “transfer errors in general,” and the one token you cared about is averaged into the background.</p>\n<p>I wrote about the keyword side of this failure in the post on <a href=\"/blog/2026-07-02-hybrid-search-rag-bm25-vectors/\">hybrid search for RAG</a>. Adding BM25 keyword search recovers the exact-match cases, but even then the fusion step is combining two coarse scores. Neither retriever ever reads the query and the chunk side by side and asks: does this specific passage answer this specific question?</p>\n<h2>What a cross-encoder does differently</h2>\n<p>A cross-encoder asks exactly that. Instead of embedding the two texts separately, it concatenates them as <code class=\"language-text\">[query] [SEP] [document]</code> and runs the pair through a transformer in one pass. Every token in the query can attend to (directly weigh itself against) every token in the document. The output is a single number: how relevant this passage is to this query.</p>\n<p>The idea comes from <a href=\"https://arxiv.org/abs/1901.04085\">Nogueira and Cho’s 2019 “Passage Re-ranking with BERT”</a>. They applied it to the MS MARCO passage ranking task and beat the previous best by a wide margin, doing nothing more exotic than this. The diagram below contrasts the two approaches.</p>\n<p><img src=\"/2aea86b1984b15c0bcabe043c2627b69/bi-vs-cross-encoder.svg\" alt=\"Two ways to score a query against a document: the bi-encoder embeds the query and the document separately and compares the vectors with cosine similarity, while the cross-encoder concatenates both into one input and runs a full transformer pass to produce a single relevance score\"></p>\n<p>The accuracy comes at a cost. Because the score depends on the pair, there is nothing to precompute: you cannot embed your corpus once and reuse it. Every query-document pair is a fresh forward pass through the model. Scoring a million documents this way at query time is a non-starter, which is why you never use a cross-encoder as your retriever. You use it as a second stage, on a shortlist the first stage has already narrowed down.</p>\n<h2>The two-stage pattern: retrieve wide, rerank narrow</h2>\n<p>The pattern is to retrieve wide and cheap, then rerank narrow and precise:</p>\n<ol>\n<li><strong>Stage one</strong> pulls the top 50 or so candidates with the bi-encoder (or hybrid search).</li>\n<li><strong>Stage two</strong> runs the cross-encoder over just those 50, rescores them, and keeps the top 5 for the prompt.</li>\n</ol>\n<p>Fifty forward passes is a latency you can pay. A million is not.</p>\n<p>The property to hold onto is that stage one alone sets the recall ceiling. The reranker can only reorder what it is given. If the answer is not among the 50 candidates, no amount of reranking invents it.</p>\n<p>So the two stages are tuned for different things. Stage one is tuned for recall: getting the right chunk into the shortlist at all. Stage two is tuned for precision: getting it to the top of the shortlist. Measure them separately, because they fail separately. Recall@50 (how often the correct chunk appears anywhere in the top 50) tells you whether stage one is even giving the reranker a chance. No reranking metric means anything until that number is high.</p>\n<h2>A local reranker with sentence-transformers</h2>\n<p>The lowest-friction way to add reranking is the <code class=\"language-text\">CrossEncoder</code> class from <a href=\"https://sbert.net/docs/cross_encoder/usage/usage.html\">sentence-transformers</a>. The <code class=\"language-text\">ms-marco-MiniLM</code> models are small, run fine on CPU for modest shortlists, and are trained for exactly this scoring task.</p>\n<p>The <code class=\"language-text\">rerank</code> function below scores each (query, candidate) pair, attaches that score to the candidate, and returns the <code class=\"language-text\">top_n</code> highest.</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\"><span class=\"token keyword\">from</span> sentence_transformers <span class=\"token keyword\">import</span> CrossEncoder\n\n<span class=\"token comment\"># small, fast, trained on MS MARCO passage ranking</span>\nreranker <span class=\"token operator\">=</span> CrossEncoder<span class=\"token punctuation\">(</span><span class=\"token string\">\"cross-encoder/ms-marco-MiniLM-L6-v2\"</span><span class=\"token punctuation\">)</span>\n\n<span class=\"token keyword\">def</span> <span class=\"token function\">rerank</span><span class=\"token punctuation\">(</span>query<span class=\"token punctuation\">,</span> candidates<span class=\"token punctuation\">,</span> top_n<span class=\"token operator\">=</span><span class=\"token number\">5</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n    <span class=\"token comment\"># candidates: list of dicts from stage one, each with a \"text\" field</span>\n    pairs <span class=\"token operator\">=</span> <span class=\"token punctuation\">[</span><span class=\"token punctuation\">(</span>query<span class=\"token punctuation\">,</span> c<span class=\"token punctuation\">[</span><span class=\"token string\">\"text\"</span><span class=\"token punctuation\">]</span><span class=\"token punctuation\">)</span> <span class=\"token keyword\">for</span> c <span class=\"token keyword\">in</span> candidates<span class=\"token punctuation\">]</span>\n    scores <span class=\"token operator\">=</span> reranker<span class=\"token punctuation\">.</span>predict<span class=\"token punctuation\">(</span>pairs<span class=\"token punctuation\">)</span>          <span class=\"token comment\"># one score per pair</span>\n    <span class=\"token keyword\">for</span> c<span class=\"token punctuation\">,</span> s <span class=\"token keyword\">in</span> <span class=\"token builtin\">zip</span><span class=\"token punctuation\">(</span>candidates<span class=\"token punctuation\">,</span> scores<span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n        c<span class=\"token punctuation\">[</span><span class=\"token string\">\"rerank_score\"</span><span class=\"token punctuation\">]</span> <span class=\"token operator\">=</span> <span class=\"token builtin\">float</span><span class=\"token punctuation\">(</span>s<span class=\"token punctuation\">)</span>\n    ranked <span class=\"token operator\">=</span> <span class=\"token builtin\">sorted</span><span class=\"token punctuation\">(</span>candidates<span class=\"token punctuation\">,</span> key<span class=\"token operator\">=</span><span class=\"token keyword\">lambda</span> c<span class=\"token punctuation\">:</span> c<span class=\"token punctuation\">[</span><span class=\"token string\">\"rerank_score\"</span><span class=\"token punctuation\">]</span><span class=\"token punctuation\">,</span> reverse<span class=\"token operator\">=</span><span class=\"token boolean\">True</span><span class=\"token punctuation\">)</span>\n    <span class=\"token keyword\">return</span> ranked<span class=\"token punctuation\">[</span><span class=\"token punctuation\">:</span>top_n<span class=\"token punctuation\">]</span></code></pre></div>\n<p>Wire it in between retrieval and prompt assembly. Stage one returns 50 candidates, and stage two cuts them to 5:</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\">candidates <span class=\"token operator\">=</span> hybrid_search<span class=\"token punctuation\">(</span>query<span class=\"token punctuation\">,</span> k<span class=\"token operator\">=</span><span class=\"token number\">50</span><span class=\"token punctuation\">)</span>       <span class=\"token comment\"># stage one: wide net</span>\ntop <span class=\"token operator\">=</span> rerank<span class=\"token punctuation\">(</span>query<span class=\"token punctuation\">,</span> candidates<span class=\"token punctuation\">,</span> top_n<span class=\"token operator\">=</span><span class=\"token number\">5</span><span class=\"token punctuation\">)</span>      <span class=\"token comment\"># stage two: precise</span>\ncontext <span class=\"token operator\">=</span> <span class=\"token string\">\"\\n\\n\"</span><span class=\"token punctuation\">.</span>join<span class=\"token punctuation\">(</span>c<span class=\"token punctuation\">[</span><span class=\"token string\">\"text\"</span><span class=\"token punctuation\">]</span> <span class=\"token keyword\">for</span> c <span class=\"token keyword\">in</span> top<span class=\"token punctuation\">)</span>\nanswer <span class=\"token operator\">=</span> llm<span class=\"token punctuation\">.</span>generate<span class=\"token punctuation\">(</span>build_prompt<span class=\"token punctuation\">(</span>query<span class=\"token punctuation\">,</span> context<span class=\"token punctuation\">)</span><span class=\"token punctuation\">)</span></code></pre></div>\n<p><code class=\"language-text\">predict</code> batches internally, so the 50 pairs go through as a handful of batches, not 50 separate calls. On a MiniLM model, a shortlist of 50 short chunks takes tens of milliseconds on a GPU and a few hundred on CPU. Measure that number on your own hardware before you decide how big a shortlist you can afford.</p>\n<h2>A hosted reranker</h2>\n<p>If you do not want to run the model yourself, the reranker is one of the few RAG components where a hosted API is genuinely convenient, because it is stateless. You send the query and the candidate texts, and you get back scores. <a href=\"https://docs.cohere.com/reference/rerank\">Cohere’s Rerank endpoint</a> is the common choice. The function below has the same inputs and outputs as the local version:</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\"><span class=\"token keyword\">import</span> cohere\n\nco <span class=\"token operator\">=</span> cohere<span class=\"token punctuation\">.</span>ClientV2<span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span>\n\n<span class=\"token keyword\">def</span> <span class=\"token function\">rerank_hosted</span><span class=\"token punctuation\">(</span>query<span class=\"token punctuation\">,</span> candidates<span class=\"token punctuation\">,</span> top_n<span class=\"token operator\">=</span><span class=\"token number\">5</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n    docs <span class=\"token operator\">=</span> <span class=\"token punctuation\">[</span>c<span class=\"token punctuation\">[</span><span class=\"token string\">\"text\"</span><span class=\"token punctuation\">]</span> <span class=\"token keyword\">for</span> c <span class=\"token keyword\">in</span> candidates<span class=\"token punctuation\">]</span>\n    resp <span class=\"token operator\">=</span> co<span class=\"token punctuation\">.</span>rerank<span class=\"token punctuation\">(</span>\n        model<span class=\"token operator\">=</span><span class=\"token string\">\"rerank-v3.5\"</span><span class=\"token punctuation\">,</span>\n        query<span class=\"token operator\">=</span>query<span class=\"token punctuation\">,</span>\n        documents<span class=\"token operator\">=</span>docs<span class=\"token punctuation\">,</span>\n        top_n<span class=\"token operator\">=</span>top_n<span class=\"token punctuation\">,</span>\n    <span class=\"token punctuation\">)</span>\n    <span class=\"token comment\"># results come back sorted, each with an index into `docs`</span>\n    <span class=\"token keyword\">return</span> <span class=\"token punctuation\">[</span>\n        <span class=\"token punctuation\">{</span><span class=\"token operator\">**</span>candidates<span class=\"token punctuation\">[</span>r<span class=\"token punctuation\">.</span>index<span class=\"token punctuation\">]</span><span class=\"token punctuation\">,</span> <span class=\"token string\">\"rerank_score\"</span><span class=\"token punctuation\">:</span> r<span class=\"token punctuation\">.</span>relevance_score<span class=\"token punctuation\">}</span>\n        <span class=\"token keyword\">for</span> r <span class=\"token keyword\">in</span> resp<span class=\"token punctuation\">.</span>results\n    <span class=\"token punctuation\">]</span></code></pre></div>\n<p>The API returns results already sorted, each with a <code class=\"language-text\">relevance_score</code> and an <code class=\"language-text\">index</code> back into the list you sent, which the code uses to recover the original candidate. <code class=\"language-text\">rerank-v3.5</code> has a 4096-token context per document and truncates longer ones, a detail worth remembering when your chunks are large.</p>\n<p>The tradeoff against the local model is the usual one. You have no GPU to manage and no model to host, but you pay per call and add a network hop to your critical path.</p>\n<h2>Choosing the shortlist size</h2>\n<p>One knob decides both quality and latency: <code class=\"language-text\">k</code>, the number of candidates you hand the reranker.</p>\n<ul>\n<li><strong>Too small</strong>, and you clip the answer out before reranking can help. The whole point of the opening example was that the answer sat at rank 8, so a <code class=\"language-text\">k</code> of 5 would have missed it.</li>\n<li><strong>Too large</strong>, and you pay for forward passes on candidates that had no chance, adding latency the user feels.</li>\n</ul>\n<p>Don’t guess. Retrieve a generous number, say 100, and measure against a set of real queries how often the correct chunk falls inside the top <code class=\"language-text\">k</code> for a range of <code class=\"language-text\">k</code> values. You are looking for the point where recall stops climbing. If recall@50 and recall@100 are the same, there is no reason to rerank 100; the extra 50 are pure cost.</p>\n<p>In practice, a <code class=\"language-text\">k</code> somewhere between 20 and 50 covers most of the recall for a well-tuned first stage, and reranking down to the top 3 to 5 is plenty for the prompt.</p>\n<h2>Failure modes</h2>\n<p><strong>Reranking a bad shortlist.</strong> This is the one that wastes the most time. If retrieval never surfaced the right chunk, the reranker reorders garbage, and you get a confidently wrong answer with a high rerank score. When a query fails, check whether the answer chunk is anywhere in the pre-rerank candidates before you touch the reranker. If it is not, the bug is upstream, in retrieval or in <a href=\"/blog/2026-07-06-chunking-strategies-for-rag/\">chunking</a>, and no reranker will save it.</p>\n<p><strong>Truncation on long chunks.</strong> Like any transformer, a cross-encoder has an input limit, and the query plus a long document can blow past it. When that happens, the model silently reads only a prefix of the chunk. If the relevant sentence lives in the tail, the reranker scores the pair as if that sentence were not there. Keep chunks inside the reranker’s context window, or the model is grading a document it only half read.</p>\n<p><strong>Treating the score as a probability.</strong> The output is a relevance score, not a calibrated confidence, and the scale differs between models. A threshold you tuned for <code class=\"language-text\">MiniLM-L6</code> means nothing for <code class=\"language-text\">rerank-v3.5</code>. If you want a cutoff that drops weak candidates, calibrate it per model against your own data, and re-check it whenever you swap the model.</p>\n<p><strong>Latency you forgot to budget.</strong> Reranking adds a synchronous step between retrieval and generation. It is small on its own, but it lands right before the model call, so it stacks on top of the slowest part of the request. If you stream responses (I covered <a href=\"/blog/2026-06-30-fastapi-sse-streaming-llm/\">SSE (Server-Sent Events) streaming from FastAPI</a> in an earlier post), the rerank has to finish before the first token can go out. The user waits through it with nothing on screen. Measure it as part of time-to-first-token, not as a free add-on.</p>\n<h2>What I would do differently</h2>\n<p>The mistake I would warn against is adding a reranker before you have measured whether retrieval is your actual bottleneck. It is a satisfying component to bolt on, and it does help. But if your recall@50 is already poor, the reranker cannot do anything with candidates that do not contain the answer.</p>\n<p>Measure the first stage first. Fix chunking and retrieval until the right document reliably lands in the top 50, and only then add reranking to sharpen the order. Done in that sequence, the reranker is the last few points of precision. Done first, it is a distraction that hides the real problem behind a plausible-looking score.</p>\n<h2>Closing: where reranking fits</h2>\n<p>Reranking is the stage that decides which retrieved chunks actually reach the model. It is the difference between a copilot that finds the right transfer-error fix and one that returns four paragraphs about transfer errors in general.</p>\n<p>This is the third piece I have written about the retrieval pipeline behind <a href=\"/project/archi/\">Archi</a>, the RAG copilot I worked on for CMS computing operations at CERN. It follows <a href=\"/blog/2026-07-06-chunking-strategies-for-rag/\">chunking</a>, which sets the ceiling, and <a href=\"/blog/2026-07-02-hybrid-search-rag-bm25-vectors/\">hybrid search</a>, which widens the net. The order matters: chunk so the answer stays whole, retrieve so it makes the shortlist, rerank so it reaches the prompt. Get all three right and the model finally sees what it needed.</p>\n<hr>\n<p><em>Diagrams by M. Hassan Ahmed, released under CC0. Image credit: original work by the author.</em></p>","frontmatter":{"title":"Cross-Encoder Reranking for RAG Pipelines","date":"2026-07-08T00:00:00.000Z","description":"Retrieval puts the right chunk at rank 8, but the generator only reads the top few. How a cross-encoder reranker reorders RAG candidates, and where it fails.","thumbnail":{"childImageSharp":{"fluid":{"base64":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAABQAAAALCAIAAADwazoUAAAACXBIWXMAAAsSAAALEgHS3X78AAABzElEQVQoz4WQ6Y6bMBSFUTWZCTvYgLGNzb4kEAgQMglJJn3/p+ptRqpU9UelT9bR9TnHi7IT6bLrvsblMZ7WYboN07U7Pqfz2k9rP8LwfpyB60sDz/lzqvZ7mbVxrmwsYQYly0dZLbI+18dntl+r4Rk3l7y718efxeFe9g9Y092F51NUzBYpIfVhS0V3IsMVuisMV1peAkJ3XrgSJiA0V25NrsMZKNFs2JKmG+v2b79iegXNTsnuxvOFZYvHDwZKoPG7QnOEZUcsK+g65etS3M/xdSZjZ/IU8orLOtncw3R2aYtYh6MB6lSLaXak2pGFBPejKG/E+Sa6mdS9V3aobG1ZqRZXHLoXzYMkJxwdPTECJs5eYW7YXBBpOEzDKc+vtLhgPkA7eODZ4FHcoIybG00mzFqPd4CJJGyE2D+lZE5DirEB4fLKy0/EO8RaFB0sL9NMqhh+FRZ3LCabHhzWA/BDGyMMMJOEOy41HQpdBoqdILf93A4K0JrNVIsqBY+/+nYoq1rKWohGSh9z5NLQ5+9G+GGEWzNUv7Go+ke/UBgmO8ELmcSUyZAyzw8Q8RD4yN+E/2rlTQ9+aORNIxs9hNu+6eRdDzZ6sDX+zy/csEbw2u67jwAAAABJRU5ErkJggg==","aspectRatio":1.899441340782123,"src":"/static/8653045cf4c590d50d2952d56a05590d/40a76/hero.png","srcSet":"/static/8653045cf4c590d50d2952d56a05590d/c972b/hero.png 340w,\n/static/8653045cf4c590d50d2952d56a05590d/27625/hero.png 680w,\n/static/8653045cf4c590d50d2952d56a05590d/40a76/hero.png 1360w,\n/static/8653045cf4c590d50d2952d56a05590d/ed396/hero.png 2000w","sizes":"(max-width: 1360px) 100vw, 1360px"}}}}}},"pageContext":{"slug":"/2026-07-08-cross-encoder-reranking-rag/","previous":"blog/2026-07-07-kubernetes-liveness-readiness-startup-probes/","next":"blog/2026-07-10-fastapi-blocking-event-loop/"}},"staticQueryHashes":["32046230"]}