Late Interaction Retrieval for RAG with ColBERT
A bi-encoder averages token detail away; a cross-encoder is too slow to rank a corpus. Late interaction with ColBERT sits between them. Here is how it works.
Two of the retrieval posts on this site describe opposite failures. In hybrid search, a query for T1_US_FNAL came back with three paragraphs about transfer errors in general and nothing about that specific site. The embedding model had squashed the identifier into the same region as every other site name. In cross-encoder reranking, the fix was to read the query and each candidate side by side, one transformer pass per pair. That is accurate, but far too expensive to run over a whole corpus.
Between those two lives a third option that most RAG (retrieval-augmented generation) stacks skip: late interaction. It keeps a separate vector for every token instead of pooling each chunk into one. It compares a query and a document at the token level, yet it still precomputes the document side offline. That is the design behind ColBERT.
This post is for engineers who already run dense retrieval and keep losing on exact identifiers, error codes, and rare terms. Those are the queries where a single pooled vector throws away the one token you cared about. I explain what late interaction computes, why it can still be precomputed, how to run it with a library, and what it costs once it is in your index.
Where a single vector loses information
A standard dense retriever is a bi-encoder: it encodes the query and each chunk separately into vectors. It runs a chunk through a transformer, then pools the per-token outputs into one fixed-length vector, often 384 or 768 dimensions. Pooling usually means taking a mean, or keeping the output of the special [CLS] token. At query time you embed the query the same way and rank by cosine similarity between two points. The HNSW post covers how the nearest-neighbor search over those points stays fast.
The pooling step is where the detail goes. Take a chunk that mentions T1_US_FNAL once, alongside four other site names and a paragraph of context. Its vector is mostly “transfer errors in general.” The token you searched for is one contribution out of a few hundred, averaged into the background before you ever asked your question. The document had to commit to a single summary of itself without knowing the query. That summary is lossy in exactly the direction that matters for precise lookups.
Late interaction removes the pooling. ColBERT keeps every token’s contextualized vector (the transformer’s output for that token, shaped by the words around it), for both the query and the document. It defers the comparison to a step that runs token against token. The name comes from that ordering. The query and document are still encoded independently, so document vectors can be precomputed. Their interaction happens late: after encoding, at retrieval time.
MaxSim: how ColBERT scores a document
The comparison ColBERT uses is called MaxSim. It works in two steps:
- For each query token, find its highest similarity against any single document token.
- Sum those maxima over the query tokens.
That sum is the document’s score.
Once you have the two sets of vectors, the operator is almost trivial in code. One matrix multiply gives every query-token-to-document-token similarity, then a row-wise max and a sum give the score:
import numpy as np
def maxsim(query_vecs, doc_vecs):
# query_vecs: (Nq, dim), doc_vecs: (Nd, dim), both L2-normalized
sim = query_vecs @ doc_vecs.T # (Nq, Nd) cosine similarities
return sim.max(axis=1).sum() # best doc token per query token, summedThe behavior falls out of that max. Each query token finds its own best evidence anywhere in the document, independently of where the other query tokens matched. So T1_US_FNAL can lock onto the one document token that is literally T1_US_FNAL and score it near 1.0. Meanwhile transfer matches a different token elsewhere, and neither dilutes the other. There is no averaging step to wash the rare token out.
The original ColBERT paper (Khattab and Zaharia, SIGIR 2020) introduced this mechanism. It is the whole reason late interaction beats a single pooled vector on exact-match queries.
Why the document side can still be precomputed
A cross-encoder can’t serve as a retriever because its score depends on the query and document jointly. Nothing precomputes, so every candidate needs a fresh forward pass. ColBERT dodges that:
- At index time, the document tokens are encoded with no knowledge of the query. You run every chunk through the encoder once and store its token vectors.
- At query time, you encode only the short query and run MaxSim. That is a matrix multiply and a max, not a transformer pass.
The result has the accuracy shape of a cross-encoder (token-level matching, exact terms preserved), with retrieval-time cost closer to dense search than to reranking. The cross-encoder post said you couldn’t have that there. ColBERT recovers it by keeping the two encoders separate and pushing the interaction to a cheap operator.
Running it without building an index by hand
You do not have to implement the storage and search yourself. The most direct path is RAGatouille, which wraps the reference ColBERT implementation. This snippet indexes a list of chunks and runs one query against them:
from ragatouille import RAGPretrainedModel
rag = RAGPretrainedModel.from_pretrained("colbert-ir/colbertv2.0")
rag.index(
collection=[chunk.text for chunk in chunks],
document_ids=[chunk.id for chunk in chunks],
index_name="cms-ops",
)
hits = rag.search("what fixed the T1_US_FNAL export buffer stall?", k=10)Each call does one job:
from_pretrainedloads a checkpoint that already knows how to produce token vectors.indexencodes the collection and writes the token vectors plus the search structures to disk.searchencodes the query and ranks by MaxSim.
Under the hood are ColBERTv2 (Santhanam et al., NAACL 2022) and the PLAID engine. PLAID prunes candidates by centroid (a representative vector for a cluster of token vectors) before running full MaxSim, so the search scales past a handful of documents.
You can also run late interaction inside a vector database that supports multi-vector documents. Vespa and Qdrant both do, which saves you standing up a second index.
The cost: one vector per token
Late interaction is not free, and the price is storage. A bi-encoder stores one vector per chunk. ColBERT stores one vector per token, so a 200-token chunk becomes 200 vectors. Before any compression, an index of a few hundred thousand chunks turns into tens of millions of small vectors. Those land on disk and, for the parts you search hot, in memory.
ColBERTv2 was built to attack this problem. It compresses each token vector against a set of learned centroids: it stores the nearest centroid’s id plus a small quantized residual (the leftover difference, stored at low precision) instead of the full vector. The reference implementation reports that this cuts the space footprint of late interaction by 6 to 10x while holding retrieval quality. Even so, budget for a late-interaction index being larger on disk than a single-vector index of the same corpus. It buys precision with bytes.
There is a second cost: the reference stack has more moving parts than index.search(vec, k). You are running an encoder, a centroid index, residual decoding, and the MaxSim pass. When something is slow or a result looks wrong, there are more layers between you and the answer than in a plain HNSW lookup.
When I would reach for ColBERT
I would not rip out a working bi-encoder to put ColBERT everywhere. It earns its keep on corpora where the queries are the identifiers. CMS computing operations, the domain behind Archi, is exactly that. Site names like T1_US_FNAL, dataset paths, error codes, config keys, and JIRA ids are the tokens a pooled vector averages away. They are also the tokens an operator is most often searching for. Late interaction keeps them.
For a mixed corpus, the honest comparison is against the two-stage pipeline the reranking post already describes: cheap dense retrieval to a shortlist, then a cross-encoder to reorder it. That stack is simpler to operate and often good enough.
ColBERT is the move when the recall stage itself fails on exact terms. If the right chunk is not landing in the top 50 at all, no reranker downstream can save it. Once you have measured that failure (the post on measuring retrieval quality covers how), late interaction is the tool aimed straight at it.
The three encoders are one spectrum, not three unrelated tricks. The bi-encoder pools everything into a point and is cheapest. The cross-encoder reads the pair jointly and is most accurate. Late interaction keeps the tokens and pushes the interaction to a cheap operator, landing in between on both axes. Knowing which failure you have tells you which one to reach for.