How to Choose an Embedding Model for RAG
The embedding model sets the ceiling on RAG retrieval quality. How to choose one by task fit, sequence length, dimensions, and domain, plus the silent bugs.
Most RAG (retrieval-augmented generation) tuning happens on the wrong end of the pipeline. People spend a week on chunk sizes, bolt on a cross-encoder reranker, and swap in a bigger model to write the answer. All of that helps. But all of it runs on the set of chunks your embedding model handed back. If the right chunk never made that set, none of the downstream work can recover it.
Yet the embedding model is picked more casually than anything else in the stack. Someone grabs whatever sat at the top of a leaderboard, or the default in the vector database quickstart, and moves on. That choice quietly sets the recall ceiling for the whole system.
This post is for engineers building a retrieval system who want a real basis for that decision instead of a leaderboard rank. I walk through the axes that actually change retrieval quality, the ones people skip, and the silent bugs that make a good model score badly. My running example is Archi, the retrieval copilot I worked on for CMS computing operations at CERN. Its corpus is internal logbooks and tickets, not the clean web text these models were trained on, and that is exactly where the default choice starts to hurt.
The embedding model is the recall ceiling
Here is the shape of the problem before the details.
A reranker reorders the top-k candidates. The LLM reads whatever made it into context. Neither can pull in a document that vector search never returned. So the embedding model and the similarity metric almost entirely decide which chunks come back for a query. Everything after them is polishing a list it can no longer add to. That is why this choice deserves more care than it usually gets, and why measuring it well matters more than measuring the parts downstream.
Read the leaderboard for your task, not the average
The reflex is to open the MTEB leaderboard and take the top row. MTEB, the Massive Text Embedding Benchmark (Muennighoff et al., 2022), is genuinely useful. But its headline number is an average across eight task types: retrieval, semantic similarity, classification, clustering, reranking, and more. A model can top the average by excelling at classification and clustering while sitting mid-pack on retrieval. Retrieval is the only column that describes what RAG does.
So filter to the retrieval task and sort on that. Retrieval is scored with nDCG@10, which rewards putting relevant documents near the top of the first ten results, and that ranking looks different from the overall board.
Watch for one more distinction. A model built for symmetric similarity, where the two texts are the same kind of thing, is not automatically good at retrieval. In retrieval, a short question has to match a long passage that never repeats its words.
Treat the leaderboard as a shortlist, not a verdict. It ranks models on public academic corpora, and your corpus is not those.
The axes that actually move retrieval
Once you have three or four candidates off the retrieval column, the choice comes down to a handful of properties. None of them is the average score.
Sequence length has to fit your chunk
Every embedding model has a maximum input length, and it silently truncates anything longer. Older BERT-family encoders cap at 512 tokens. Several newer models take 8192 (the E5 line and others).
Suppose your chunks run 1,000 tokens and you feed them to a 512-token model. The back half of every chunk is dropped before it is ever embedded. You will see it as unexplained recall misses on content that happens to live late in a chunk.
Pick the model and the chunk size together, not separately. The chunk has to fit inside the model’s window with the query prefix included.
Dimensions are a storage and latency bill
The number of dimensions in each vector decides how much you pay to store and search the index. A 3072-dim vector at float32 is 12 KB. At ten million chunks, that is 120 GB before any HNSW graph overhead, and every query does similarity math across all those dimensions. Halving the dimensions roughly halves the memory and makes search noticeably faster.
You used to have to retrain to get a smaller vector. Now many models are trained with Matryoshka Representation Learning (Kusupati et al., 2022). That training packs the important signal into the early dimensions, so you can truncate the vector and keep most of the quality. OpenAI’s text-embedding-3 models use this. Their own numbers show a text-embedding-3-large vector cut to 256 dimensions still beating the older 1536-dim ada-002 on MTEB (announcement).
The catch is in the footnote of that diagram: truncation is only safe on a model trained for it. Chop an ordinary embedding to a third of its length and you get a broken vector, not a smaller one, because nothing forced the signal toward the front. And after you truncate, re-normalize the vector to unit length. Otherwise your cosine distances quietly stop meaning what your index assumes they mean.
Symmetric or asymmetric: the prefix people forget
This one costs more retrieval quality than any other single mistake I see, and it produces no error.
Retrieval is asymmetric: a short query has to match a long passage that phrases things differently. Several strong open models are trained for exactly that, and they expect you to label which side of the pair each text is:
- The E5 models want a literal
query:prefix on the query andpassage:on each document. - The BGE family wants an instruction like
Represent this sentence for searching relevant passages:on the query and plain text on the documents.
These prefixes are not optional decoration. The papers and model cards are explicit that omitting them degrades retrieval, because the asymmetry is what the model learned. Here is the E5 pattern with sentence_transformers; note that documents and queries go through the same model with different prefixes:
from sentence_transformers import SentenceTransformer
model = SentenceTransformer("intfloat/e5-large-v2")
# Index time: every chunk goes in with the passage prefix.
doc_vecs = model.encode(
[f"passage: {chunk}" for chunk in chunks],
normalize_embeddings=True,
)
# Query time: the SAME model, but the query prefix.
q_vec = model.encode(
"query: why did the transfer agent back off?",
normalize_embeddings=True,
)Get the prefixes backwards, or drop them, and the model still returns a vector and search still runs. The results are just quietly worse. Because nothing throws an error, this can hide for months. When someone reports that an open model “underperforms the API,” check the prefixes first.
Domain fit beats the benchmark
MTEB is mostly clean, general text: Wikipedia, web pages, news. Your corpus may be nothing like it. Archi indexes shift logbooks full of site names, error codes, and internal shorthand. A model that tops the general retrieval board can still fumble a query where the signal is a workflow ID and a stack trace.
General-purpose models are a strong default. But domain-adapted or instruction-tuned models sometimes win on specialized corpora, by a margin that dwarfs the leaderboard gaps between the top general models. The only way to know is to test on your own text, which is the next section.
Hosted API or self-hosted
A hosted embedding API (OpenAI, Cohere, Voyage) is one call, with no GPU to run. It also means every chunk you ever index, and every query, leaves your network. For a CERN-internal corpus, that alone is often a hard stop.
Self-hosting an open model like E5 or BGE keeps the data in-house and removes the per-token bill. The cost is that you run the inference yourself and own the throughput. For Archi, the data-locality question decides this before performance does. If your corpus is public and you want the fastest path to a working system, a hosted model is the easier start.
The cheap evaluation that settles it
None of the axes above tells you which model wins on your data. A small labeled set does. It is a couple of hours of work that saves you from shipping the wrong model:
- Collect 50 to 100 real queries. Real ones, from logs or from the people who will use the thing, not questions you invented to flatter the system.
- For each query, mark the chunks that genuinely answer it.
- For each candidate model, embed the corpus, run the queries, and measure recall@k: of the chunks you marked relevant, how many showed up in the top k.
The code is small. recall_at_k scores one query, and the average over the eval set is the number you compare across models:
def recall_at_k(retrieved_ids, relevant_ids, k):
hits = len(set(retrieved_ids[:k]) & set(relevant_ids))
return hits / len(relevant_ids)
# Average recall@k across the eval set is the number that decides it.
scores = [recall_at_k(r, rel, k=10) for r, rel in zip(runs, gold)]
print(sum(scores) / len(scores))Recall@k is the honest metric for the embedding stage. It asks the one question the reranker cannot fix later: did the right chunk make the candidate set at all? I went deeper on retrieval metrics, including where recall and nDCG diverge, in measuring RAG retrieval quality.
Run each candidate at full and truncated dimensions, with and without the correct prefix. You will usually find that the ranking on your data disagrees with the leaderboard, which is the entire point of doing it.
Silent failure modes to check first
When retrieval is worse than expected, these are the usual causes, roughly in the order I check them. All four fail without an error.
- Prefix mismatch. The query and documents were encoded with the wrong prefix, or none. This is the single most common one for open models, covered above.
- Metric and normalization mismatch. Cosine similarity equals a dot product only on unit-length vectors. If your index computes dot product but you did not normalize, or you normalized at index time and not at query time, distances mean the wrong thing. Pick one metric and normalize consistently on both sides.
- Truncation at max length. Chunks longer than the model’s window lose their tail before embedding. Log token counts during ingestion and you will see it.
- The re-embedding tax on model swaps. Vectors from two different models are not comparable, so changing the embedding model means re-embedding the whole corpus, not a rolling upgrade. On a large or steadily changing index, that is a real batch job. It is one more reason to run the eval before you commit, and part of why incremental indexing is worth designing for up front.
What I’d reach for
For a general English corpus with no data-locality constraint, I would use a top hosted retrieval model truncated with Matryoshka to 512 or 1024 dimensions. It is a strong, boring default; I would ship it and move on.
For anything that has to stay in-house, or anything with heavy domain vocabulary, I start from a self-hosted E5 or BGE model. I get the prefixes right and let the recall@k eval on real queries pick the finalist. Then I add hybrid search with BM25 (keyword ranking alongside the vectors). That way exact identifiers, the error codes and workflow IDs a dense model paraphrases away, still land.
That is the order that matters. Chunking, reranking, and the answer model are all worth tuning, but they operate on whatever the embedding model retrieved. Spend your first careful hour on the model that sets the ceiling, measure it on your own text, and check the prefix before you blame the model. On Archi, the corpus is CERN’s internal operations knowledge and the queries are real operator questions. There, that ordering is the difference between a copilot that finds the past incident and one that confidently answers from the wrong three chunks.
Diagrams by M. Hassan Ahmed, released under CC0. No external image was used for this post; the figures are original work by the author.