Matryoshka Embeddings: Smaller Vectors for RAG

Matryoshka embeddings front-load meaning into a vector's first dimensions, so you can truncate them for cheaper, faster RAG search without losing much recall.

Every vector you store in a RAG (retrieval-augmented generation) system costs you twice: once on disk, and again in RAM. An approximate nearest-neighbor index wants its graph resident in memory to be fast. Switch to a modern embedding model and the per-vector cost jumps. OpenAI’s text-embedding-3-large returns 3072 floats per chunk, double what ada-002 gave you. On a corpus of a few hundred thousand documents, that is real money and real latency. The usual answer, “pick a smaller model,” means giving up retrieval quality to get it.

Matryoshka embeddings let you keep the good model and shrink the vector instead. You store the first few hundred numbers of each embedding rather than all of them, and retrieval barely notices. This post is for engineers building RAG who want a smaller vector index and faster search, without dropping to a weaker model or re-embedding a corpus they already have.

I explain why truncating a normal embedding wrecks it, what Matryoshka training changes, the exact steps to use one (including the one people forget), and where it bites. The context is Archi, the retrieval copilot I worked on for CMS computing operations at CERN. Its corpus is logbooks and tickets, and index size there is a standing constraint, not an afterthought.

Why you cannot just truncate a normal embedding

The obvious move is to keep the first 256 numbers of a 3072-dim vector and throw the rest away. With an ordinary embedding model, this destroys the vector. Nothing in standard contrastive training says the first 256 dimensions should carry more meaning than dimensions 1000 through 1256. The information is spread across the whole vector with no ordering, so a prefix is a random subset of half-formed features. Cosine similarity over those prefixes is close to noise, and your recall falls off a cliff.

So the naive optimization fails for a real reason: dimension order is meaningless in a normal embedding. Matryoshka Representation Learning (MRL), introduced by Kusupati et al. at NeurIPS 2022, changes exactly that.

What Matryoshka training changes

The idea is in the name: Russian nesting dolls, each complete inside the next. MRL trains the model so that nested prefixes of one embedding (say the first 64, 128, 256, 512, … dimensions) are each a usable representation on their own.

The mechanism is a small change to the loss. Standard training computes the loss only on the full 3072-dim output. MRL computes it at several nested lengths at once and adds them up, as in this sketch:

# Sketch of the Matryoshka loss. `z` is the full embedding for a batch;
# each prefix z[:, :m] is asked to solve the same task on its own.
loss = 0.0
for m in [64, 128, 256, 512, 1024, 2048, 3072]:
    loss += criterion(head(z[:, :m]), labels)   # e.g. contrastive / classification loss
loss.backward()

Training grades the 256-dim prefix directly, so gradients push the most useful, most general features into the earliest coordinates. Later coordinates get the finer detail that only helps when you can afford it. The paper shows this needs only O(log d) of these cut points to make every prefix in between behave. It also adds no cost at inference: you still do one forward pass and get one vector. The nesting is a property of that vector, not extra work.

A 3072-dimension embedding drawn as a bar shaded from dark on the left to light on the right, where the leftmost coordinates hold the most information and later ones add finer detail. Below it, four boxes show keeping a 256-, 512-, 1536-, or full 3072-dimension prefix, trading index size against recall, with a note that the short vector is just the first N numbers of the long one re-scaled to unit length.

The one number worth remembering comes from the MTEB benchmark. There, text-embedding-3-large truncated to 256 dimensions still scores higher than the old ada-002 at its full 1536 (about 62.0 versus 61.0, per OpenAI’s release notes). A vector one-twelfth the size, beating last year’s default. That is the whole pitch.

Using them: truncate, then re-normalize

Two rules are easy to get wrong:

  1. If your model was trained with MRL, you truncate the vector yourself. You do not ask for a different model.
  2. After truncating, you must re-normalize to unit length before comparing. Your similarity metric assumes unit-length vectors.

There are three common ways to apply those rules, depending on where your embeddings come from.

With the OpenAI API, the dimensions parameter does both for you, on the server:

from openai import OpenAI
client = OpenAI()

resp = client.embeddings.create(
    model="text-embedding-3-large",
    input="rucio transfer stuck at T2_CH_CERN, exit code 8021",
    dimensions=256,          # truncate to 256 and re-normalize, server-side
)
vec = resp.data[0].embedding # already unit-length, ready to index

With an open model through Sentence Transformers, pass truncate_dim and it handles the renormalization. Nomic’s nomic-embed-text-v1.5, used below, is Matryoshka-trained and goes down to 64 dims:

from sentence_transformers import SentenceTransformer

model = SentenceTransformer("nomic-ai/nomic-embed-text-v1.5",
                            truncate_dim=256, trust_remote_code=True)
vecs = model.encode(chunks, normalize_embeddings=True)

With full-length vectors you already have sitting in a table, you can get the short version without re-embedding anything. It takes three lines of NumPy: slice, compute the norms, divide. This is the case that saved me a re-index on Archi. The full embeddings were already computed, so shrinking the index was a transform over data I had.

import numpy as np

def shrink(vecs: np.ndarray, dim: int) -> np.ndarray:
    prefix = vecs[:, :dim]                        # keep the first `dim` coordinates
    norms = np.linalg.norm(prefix, axis=1, keepdims=True)
    return prefix / norms                         # back to unit length -- do not skip this

Skip that last division and everything still runs, which is the trap. Truncating changes each vector’s magnitude, so un-normalized cosine scores are quietly wrong. Your recall sags for a reason no error message will point at. Re-normalize.

One caveat surprises people. For a model like OpenAI’s, asking the API for dimensions=256 is not always identical to embedding at full size and slicing the first 256 yourself, because the service re-normalizes against the full-length norm. Pick one path, and keep the query and the corpus on the same path. Mixing a sliced-and-locally-normalized query with a server-truncated index is the kind of bug that shows up as “retrieval feels a bit off” and costs an afternoon.

Adaptive retrieval: cheap first pass, precise second

Truncating everything to 256 dims trades a little accuracy for a lot of savings. Adaptive retrieval, which the MRL paper calls funnel retrieval, tries to win back the accuracy while keeping most of the savings. It is a two-pass search.

A two-pass retrieval diagram. A corpus of one million chunks keeps short 256-dimension vectors in the in-memory approximate-nearest-neighbor index and full 3072-dimension vectors on disk. Pass one runs a fast coarse search over the short vectors and returns about 200 candidates. Pass two rescores only those candidates with the full-dimension dot product and keeps the top ten, which go to the language model. The expensive full-dimension math runs on hundreds of rows, not millions.

  1. Pass one searches the whole corpus using the short 256-dim vectors. It is fast, and the index that has to live in RAM is small. It returns a generous shortlist of a few hundred candidates.
  2. Pass two takes only that shortlist and rescores it with the full 3072-dim vectors, which you kept on disk, to get exact-quality ordering.

The expensive high-dimensional math runs on a few hundred rows instead of a few million, so you pay for full precision only where it changes the answer.

If that shape looks familiar, it should. It is the same move as cross-encoder reranking, one layer earlier. There, a cheap bi-encoder (which embeds query and document separately) proposes, and an expensive cross-encoder (which reads them together) reorders. Here, a cheap short vector proposes and an expensive long vector reorders.

Cheap-then-precise is the reliable structure for retrieval that has to be both fast and good. You can also stack both stages: short-vector recall, then full-vector rescore, then a cross-encoder on the final handful. Weaviate and Supabase both document this against pgvector-style stores if you want a worked setup.

Where it bites

The model has to be trained for it. This is not a trick you apply to any embedding; truncating a non-MRL model gives you noise, full stop. Check the model card before you plan an index around a dimension count. text-embedding-3-* and nomic-embed-text-v1.5 are trained for it, and plenty of otherwise-good models are not.

Shorter is not free, and the curve is not linear. Recall holds up well down to a point and then drops faster. Where that knee sits depends on your corpus and your queries, so measure it rather than assuming 256 is fine because a benchmark said so. Hold out a set of real queries with known-good answers and plot recall@k at 128, 256, 512, and full.

On a narrow technical corpus, the safe dimension is often higher than the general-purpose benchmark suggests. Your documents are less spread out, and near-duplicates are easy to confuse once you throw away the fine-detail coordinates.

Re-normalization, again. It is the single most common mistake, so it is worth stating twice. Every truncation needs a re-normalize, and your query vector has to be produced the same way as your stored vectors.

Adaptive retrieval needs the full vectors kept somewhere. The two-pass win assumes you still have the 3072-dim vectors on disk for the rescore. If you truncated at ingest and threw the rest away to save space, you cannot do the precise second pass. Keep the full vectors in cheap storage (a column you do not index) if you want that option later.

It does not fix a bad chunker. A smaller vector retrieves the chunks you gave it. If your chunking splits a table away from its caption, no dimension count rescues that. Matryoshka is an efficiency lever on retrieval, not a quality fix for ingestion.

Tradeoffs, and when I reach for it

Put plainly, Matryoshka embeddings buy you a storage-and-latency dial. That dial used to require re-embedding the whole corpus with a different model. It is worth a lot when the index is large and memory-bound, and almost nothing when it is small. If your corpus is ten thousand chunks, the full vectors fit in RAM without complaint, and this is premature optimization: store the full thing and move on. The lever earns its keep once index memory is a line item you actually notice.

For Archi that point arrives naturally, because operational corpora only grow and the index sits next to everything else the ops team runs. I truncate to a shorter vector for the first-pass search and keep the full vectors for a rescore on the shortlist. That gets most of the recall for a fraction of the resident memory, on embeddings I had already computed. It follows the same instinct as the rest of that retrieval stack, hybrid search and reranking. Spend the expensive computation only on the handful of candidates where it changes the answer, and keep the first, widest pass as cheap as you can make it.


Diagrams by M. Hassan Ahmed, created for this post and released under CC0 (public domain). No external image was used.