{"componentChunkName":"component---src-templates-blog-post-js","path":"/blog/2026-08-18-product-quantization-vector-search/","result":{"data":{"site":{"siteMetadata":{"title":"M.Hassan Ahmed","author":"Hassan11196"}},"markdownRemark":{"id":"2700db83-5625-52c5-9c8e-a3d808b301f2","excerpt":"The HNSW post ended on a bill I did not pay. HNSW makes the search fast by walking a graph in a handful of hops instead of scanning the whole corpus. But it…","html":"<p>The <a href=\"/blog/2026-07-09-hnsw-vector-search-explained/\">HNSW post</a> ended on a bill I did not pay. HNSW makes the <em>search</em> fast by walking a graph in a handful of hops instead of scanning the whole corpus. But it still keeps every embedding resident in memory as raw <code class=\"language-text\">float32</code>. On a small index nobody notices. On a corpus where every log line, ticket, and doc page becomes a chunk, the vectors themselves stop fitting in the box. They run out of room before query latency ever becomes a problem.</p>\n<p>This post is about shrinking the vectors, not the search over them. The technique is product quantization (PQ), first described by <a href=\"https://inria.hal.science/inria-00514462\">Jégou, Douze, and Schmid in 2011</a>. It is what sits underneath the “compressed” index types in <a href=\"https://github.com/facebookresearch/faiss\">FAISS</a>, Milvus, and Qdrant. The post is for engineers who already have vector search working, have watched the memory number climb, and want to know what the index is doing when it trades exactness for room to grow.</p>\n<h2>The problem is memory, not speed</h2>\n<p>Do the arithmetic once and it sticks. A million chunks at 768 dimensions, four bytes per dimension, is <code class=\"language-text\">1e6 × 768 × 4 ≈ 3 GB</code> of vectors alone. The HNSW graph adds more on top, because every node stores a list of neighbor ids. Push the corpus to ten million and you are near 30 GB.</p>\n<p>All of it has to stay resident. Approximate nearest neighbor search is random access by nature, so there is no hot working set to page in and out. You can shard across machines, but each shard still holds its slice in RAM.</p>\n<p>So the question changes. HNSW answered “how do I look at fewer vectors?” Product quantization answers a different one: “how do I make each vector smaller?” The goal is small enough that the whole index stays in memory, instead of spilling to disk or onto a bigger, pricier instance.</p>\n<h2>Scalar quantization gets you 4x, then stops</h2>\n<p>The first move most people make is scalar quantization: map each <code class=\"language-text\">float32</code> dimension to an <code class=\"language-text\">int8</code>, one byte instead of four. That is a clean 4x, it is cheap, and recall barely moves for most embedding models, which is why several vector databases now default to it. But it treats every dimension on its own, so it bottoms out at one byte per dimension. To go further you have to stop quantizing dimensions independently, and that is exactly the door PQ walks through.</p>\n<h2>The idea: quantize chunks of the vector</h2>\n<p>PQ encodes a vector in three steps:</p>\n<ol>\n<li>Split the <code class=\"language-text\">D</code>-dimensional vector into <code class=\"language-text\">m</code> contiguous sub-vectors.</li>\n<li>In each subspace, run <a href=\"https://en.wikipedia.org/wiki/K-means_clustering\">k-means</a> on a sample of your data to learn a small codebook of <code class=\"language-text\">k</code> centroids (representative points for that slice). Pick <code class=\"language-text\">k = 256</code>, and a centroid id fits in a single byte.</li>\n<li>Replace each sub-vector with the id of its nearest centroid.</li>\n</ol>\n<p>The whole vector is now <code class=\"language-text\">m</code> bytes.</p>\n<p><img src=\"/8a28542199bc27ea6b542337ea0eb083/pq-encoding.svg\" alt=\"A 768-dimensional embedding is split into four sub-vectors; each sub-vector is snapped to the nearest of 256 centroids in its own learned codebook, and the four centroid ids are stored as a four-byte code that stands in for the original 3072-byte vector.\"></p>\n<p>The word “product” is the interesting part. With <code class=\"language-text\">m</code> codebooks of 256 entries each, you can represent <code class=\"language-text\">256^m</code> distinct vectors: the Cartesian <em>product</em> of the per-subspace choices. Yet you only ever store <code class=\"language-text\">m × 256</code> centroids. That combinatorial reach from a tiny table is the whole trick.</p>\n<p>The compression follows from <code class=\"language-text\">m</code>. Take a 768-dim embedding, which is <code class=\"language-text\">3072</code> bytes as raw floats. Split it into <code class=\"language-text\">m = 96</code> sub-vectors of 8 dimensions each, at one byte per sub-quantizer, and you land at 96 bytes: about 32x smaller. Drop to <code class=\"language-text\">m = 48</code> and you get 48 bytes and 64x, with more error, because each byte now has to summarize a wider slice of the vector. The diagram uses <code class=\"language-text\">m = 4</code> so the codebooks are legible. In practice the knobs are <code class=\"language-text\">m</code> (how many pieces) and <code class=\"language-text\">nbits</code> (how many centroids per piece, set as bits per code; almost always 8).</p>\n<h2>Computing distances without decompressing</h2>\n<p>This is the part that makes PQ fast at query time and not just small on disk: you never rebuild the vector. The method is <a href=\"https://inria.hal.science/inria-00514462\">asymmetric distance computation</a> (ADC):</p>\n<ol>\n<li>Keep the query at full precision, and split it the same way.</li>\n<li>For each subspace, precompute a small table of distances from the query’s sub-vector to all 256 centroids of that codebook. That is an <code class=\"language-text\">m × 256</code> table, built once per query.</li>\n<li>The distance from the query to any stored code is now a sum of <code class=\"language-text\">m</code> lookups. For each sub-quantizer, use the code’s stored byte as the index, read one cell, and add the cells up.</li>\n</ol>\n<p><img src=\"/2d9f8529562a83e358ae2c866d8afc3f/pq-adc-lookup.svg\" alt=\"The query stays full precision and is split into sub-vectors; a distance table is precomputed from each query sub-vector to all 256 centroids of its codebook, so the approximate distance to a stored code is four table reads plus three additions, with no vector ever decompressed.\"></p>\n<p>With the diagram’s <code class=\"language-text\">m = 4</code>, scanning a million codes becomes a million rounds of “four reads and three adds.” There are no 768-dimension dot products and no decompression anywhere in the loop.</p>\n<p>There is also a symmetric variant that quantizes the query too. It lets you reuse tables across queries, but it folds the query’s own quantization error into every distance. “Asymmetric” keeps that error out of the estimate, so ADC is the default worth reaching for.</p>\n<h2>PQ alone still scans everything, so pair it with IVF</h2>\n<p>One honest caveat: PQ shrinks the vectors, but by itself it does not reduce how many you compare against. A million lookups is fast, but it is still <code class=\"language-text\">O(N)</code>. So in practice PQ rides underneath a coarse index that narrows the candidate set first.</p>\n<p>The classic pairing is an inverted file index, <a href=\"https://github.com/facebookresearch/faiss/wiki/Faster-search\">IVF</a>:</p>\n<ul>\n<li>A coarse k-means partitions the space into cells.</li>\n<li>The query probes only the <code class=\"language-text\">nprobe</code> nearest cells.</li>\n<li>PQ encodes the <em>residual</em> (the vector minus its cell’s centroid), which has smaller magnitude and therefore quantizes more accurately.</li>\n</ul>\n<p>That combination, IVF-PQ, is the workhorse of large FAISS indexes. It shows up in the <a href=\"https://github.com/facebookresearch/faiss/wiki/The-index-factory\">index factory</a> strings people paste around, like <code class=\"language-text\">IVF4096,PQ96</code>. You can also put PQ codes under an HNSW graph instead of IVF: same division of labor, different coarse structure.</p>\n<h2>Where it bites</h2>\n<p>The happy path is short. The engineering is in the edges, and PQ has a few sharp ones.</p>\n<p><strong>Recall drops, and you feel it at the top of the list.</strong> Approximate distances reorder near-ties, so the true nearest neighbor sometimes comes back ranked third. The standard fix is a two-stage read: pull a larger shortlist with PQ, then rerank it with exact full-precision vectors. That only works if you kept the full vectors somewhere. FAISS’s <code class=\"language-text\">IndexRefineFlat</code> does exactly this: it buys the memory back, but only for the shortlist, not the whole corpus. <a href=\"/blog/2026-07-04-measuring-rag-retrieval-quality/\">Measure recall against a flat baseline</a> before and after. The whole point of PQ is the trade (how much recall you give up for the memory you save), and you cannot judge that trade without the number.</p>\n<p><strong>Codebooks are trained, so they go stale.</strong> Those centroids were learned by k-means from a sample of your embeddings. Swap the embedding model, or let the corpus drift into a new domain, and the codebooks no longer describe the vectors you are storing. Encoding error then climbs quietly. Retraining means re-encoding everything, which runs straight into the <a href=\"/blog/2026-06-29-incremental-rag-indexing/\">incremental indexing</a> tradeoffs: you cannot just append; you have to rebuild.</p>\n<p><strong>Variance is uneven across dimensions, and naive splits waste bits.</strong> Chop a vector into contiguous slices and one sub-quantizer may get all the high-variance dimensions while another gets nearly constant ones. Both spend the same 256 centroids on very different amounts of information. Optimized product quantization (OPQ), from <a href=\"https://www.microsoft.com/en-us/research/publication/optimized-product-quantization/\">Ge, He, Ke, and Sun</a>, learns a rotation first, so variance spreads evenly across subspaces before quantizing. It is usually close to a free recall bump, and FAISS exposes it as an <code class=\"language-text\">OPQ</code> pre-transform you prepend to the index string.</p>\n<p><strong>The metric has to match.</strong> Textbook PQ assumes Euclidean (L2) distance. Most modern text embeddings are compared with cosine or inner product, so normalize the vectors first. Otherwise, the geometry the codebooks were built around is not the geometry you are querying in.</p>\n<p><strong>Small corpora do not need it.</strong> If the raw vectors fit in RAM (a few hundred thousand 768-dim vectors is one to two gigabytes), PQ only adds error for no real saving. It earns its place when memory is the binding constraint, not as a reflex.</p>\n<h2>What I would reach for, in order</h2>\n<p>The order of moves as an index grows is fairly settled:</p>\n<ol>\n<li>Keep flat or HNSW with full <code class=\"language-text\">float32</code> while it fits.</li>\n<li>Reach for scalar <code class=\"language-text\">int8</code> when you want a cheap 4x and can spare a little recall.</li>\n<li>Move to IVF-PQ, or HNSW over PQ codes, once you are genuinely memory-bound in the millions. When you do, add OPQ and keep an exact-rerank step for the shortlist.</li>\n</ol>\n<p>At every step, hold the <a href=\"/blog/2026-07-04-measuring-rag-retrieval-quality/\">flat baseline’s recall@k</a> (how many of the true top <code class=\"language-text\">k</code> the index returns) next to the compressed index’s. That gap is the only thing that tells you whether the memory you saved cost you answers.</p>\n<p>For <a href=\"/project/archi/\">Archi</a>, the retrieval copilot I worked on for CMS computing operations at CERN, this is not a hypothetical. The knowledge base has no natural ceiling: every new logbook entry and ticket is another chunk. It is the vectors, not the graph traversal, that threaten to price the index out of the memory it runs in. PQ is the lever that keeps a growing dense index resident. It sits quietly under the dense half of <a href=\"/blog/2026-07-02-hybrid-search-rag-bm25-vectors/\">hybrid search</a>, and the whole point is that once the recall math checks out, nobody upstream has to think about it again.</p>\n<hr>\n<p><em>Diagrams by M. Hassan Ahmed, released under CC0. No external image was used for this post; the figures are original work by the author.</em></p>","frontmatter":{"title":"Product Quantization for Vector Search","date":"2026-08-18T00:00:00.000Z","description":"HNSW makes vector search fast, but the embeddings still fill your RAM. Product quantization compresses them ~32x with a small recall hit. Here is how it works.","thumbnail":{"childImageSharp":{"fluid":{"base64":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAABQAAAALCAIAAADwazoUAAAACXBIWXMAAAsSAAALEgHS3X78AAAB9klEQVQozz1Q2W7bMBDUmy1KlERRsg7qMnVbR2w39SH5StLEboIANdrH9lP6713ZQIHBYriYmeWuoOc9KXq92Knphm/e66dfxeHKNx/F/ke8/Ziu36vTT76G5zV4fFPTrV70JB8AREDYRdgBYM2Pkvms6au2T8vVrO3yap2WX8tmm83W9XwfxguVRqCXFBcpUJkg6aFsTLGVWVFr+JXOSmUSSyQQSYBIIJEQaQMRtUCmETamoP8PQXcqVr642cHLjizdqSaXSaAYsWIm+FaBS6qHKVcMjkk4EDOBdEnzBeq2QfPdK569/IllRxlGqQwbXJ2kxM6hQgRSPZlEmteQZK35rWLG4IQmmBu/vrD8aTCne1kfzLI+Nf120X3qzkymMdI8WWHGwyv//Ze9/MF0ejMzgTq1X59ZdvKykwtm2FBxZX1YT3cL+ALw4ZxwFxobfq1aGb4tAjLYGcyX+7fd9CARX1RsM1y6+cFNOpb2wMfYglEOXwTFKixW2iQRFQcgELv0ymeLd07cO3EHkSK2/OKYPF6j5pwsP4GPkKEED/byYs+/2YtXo9wjjYmyJcAcYDz/YrF8jG1xgAWp1EnzZkvtRMSOiCfI5FbVBauzVe/UsAEXyISxPAFoRogJG0nm/QkEYYuYESjuzRGiY0QlGkL6SNTvsn9m7U0QrDp1ngAAAABJRU5ErkJggg==","aspectRatio":1.899441340782123,"src":"/static/4cc29e9488098aed49e7246f0af47f78/40a76/hero.png","srcSet":"/static/4cc29e9488098aed49e7246f0af47f78/c972b/hero.png 340w,\n/static/4cc29e9488098aed49e7246f0af47f78/27625/hero.png 680w,\n/static/4cc29e9488098aed49e7246f0af47f78/40a76/hero.png 1360w,\n/static/4cc29e9488098aed49e7246f0af47f78/ed396/hero.png 2000w","sizes":"(max-width: 1360px) 100vw, 1360px"}}}}}},"pageContext":{"slug":"/2026-08-18-product-quantization-vector-search/","previous":"blog/2026-08-19-concurrent-llm-calls-asyncio-gather/","next":"blog/2026-08-22-hyde-hypothetical-document-embeddings-rag/"}},"staticQueryHashes":["32046230"]}