Binary Quantization for Vector Search
Float32 embeddings make a vector index expensive to keep in RAM. Binary quantization cuts them 32x; Hamming search plus rescoring keeps recall high.
Float32 embeddings make a vector index expensive to keep in RAM. Binary quantization cuts them 32x; Hamming search plus rescoring keeps recall high.
Cosine, dot product, and Euclidean rank vector search results the same on normalized embeddings, differently otherwise. How to pick and configure one.
Small chunks search precisely but read like fragments. Parent document retrieval matches on child chunks, then feeds the whole parent to the model.
Brute-force vector search compares every embedding. How an IVF index partitions vectors into cells, probes only the nearest, and where the nprobe knob bites.
Vector search misses chunks that lost their document context. Contextual Retrieval prepends a short LLM-written summary to each chunk before you embed it.
HNSW makes vector search fast, but the embeddings still fill your RAM. Product quantization compresses them ~32x with a small recall hit. Here is how it works.
Matryoshka embeddings front-load meaning into a vector's first dimensions, so you can truncate them for cheaper, faster RAG search without losing much recall.
The embedding model sets the ceiling on RAG retrieval quality. How to choose one by task fit, sequence length, dimensions, and domain, plus the silent bugs.
Exact-match caching misses paraphrases, so LLM bills stay high. Here is how to build a semantic cache with embeddings, a similarity threshold, and its traps.