{"componentChunkName":"component---src-templates-blog-post-js","path":"/blog/2026-07-09-hnsw-vector-search-explained/","result":{"data":{"site":{"siteMetadata":{"title":"M.Hassan Ahmed","author":"Hassan11196"}},"markdownRemark":{"id":"e0b07ce6-e8fc-5bd1-8029-6f8039c2ca11","excerpt":"Every RAG (retrieval-augmented generation) post I have written so far ends up calling the same function: give me the  chunks whose embeddings are closest to…","html":"<p>Every RAG (retrieval-augmented generation) post I have written so far ends up calling the same function: give me the <code class=\"language-text\">k</code> chunks whose embeddings are closest to this query. <a href=\"/blog/2026-07-06-chunking-strategies-for-rag/\">Chunking</a> decides what goes in the index, <a href=\"/blog/2026-07-02-hybrid-search-rag-bm25-vectors/\">hybrid search</a> decides how you query it, and <a href=\"/blog/2026-07-08-cross-encoder-reranking-rag/\">reranking</a> decides what survives into the prompt. In the middle of all that is one line, <code class=\"language-text\">index.search(query_vector, k=50)</code>, and I have been treating it as free.</p>\n<p>It is not free, and the thing making it fast almost certainly has four letters: HNSW. If you use pgvector, Qdrant, Weaviate, Milvus, FAISS, or OpenSearch’s k-NN plugin, HNSW is the default index behind that search call.</p>\n<p>This post is for engineers who have vector search working and want to know what the index is actually doing when:</p>\n<ul>\n<li>recall (the share of the true nearest neighbors the search returns) drops;</li>\n<li>memory blows up;</li>\n<li>a config value copied off a blog post turns out to matter.</li>\n</ul>\n<p>I will cover the exact-search problem HNSW dodges, how its layered graph routes a query, the three parameters worth understanding, and the failure modes that bite in production.</p>\n<h2>The problem: exact nearest-neighbor search does not scale</h2>\n<p>Start with the naive version. You have a query vector and a corpus of <code class=\"language-text\">N</code> document vectors, each maybe 768 or 1536 dimensions. The exact answer is a linear scan: compute the distance from the query to every one of the <code class=\"language-text\">N</code> vectors, sort, and take the top <code class=\"language-text\">k</code>. This is called a flat index, and it is genuinely exact.</p>\n<p>It is also <code class=\"language-text\">O(N)</code> per query. For a few thousand vectors that is fine, and you should not reach for anything cleverer. But in the corpus behind a real copilot, every log line, ticket, and doc page becomes a chunk. <code class=\"language-text\">N</code> is in the millions, and a full scan on every keystroke is not a latency you can pay. Worse, the <a href=\"https://en.wikipedia.org/wiki/Curse_of_dimensionality\">curse of dimensionality</a> breaks the usual shortcuts. The tree structures that speed up low-dimensional nearest-neighbor search (kd-trees and friends) collapse back toward a full scan once you are in hundreds of dimensions.</p>\n<p><img src=\"/2af8ef444e418464d0c56cfc8b06a676/flat-vs-hnsw.svg\" alt=\"Flat search compares the query to every point in the corpus, so its cost grows with N; HNSW walks a graph and visits only a handful of points, trading a small miss rate for a cost that barely moves as the corpus grows\"></p>\n<p>So you give up exactness. Approximate nearest neighbor (ANN) search accepts that it will occasionally miss the true closest vector. In exchange, query cost grows like <code class=\"language-text\">log(N)</code> instead of <code class=\"language-text\">N</code>. Every ANN method answers the same question in its own way: how do you visit a handful of promising vectors instead of all of them, without a tree? HNSW answers it with a graph.</p>\n<h2>The idea: a navigable small-world graph</h2>\n<p>HNSW stands for Hierarchical Navigable Small World. It comes from <a href=\"https://arxiv.org/abs/1603.09320\">Malkov and Yashunin’s 2016 paper</a> (later in IEEE TPAMI). The name holds two ideas, and together they are the whole trick, so it is worth pulling them apart.</p>\n<p>A <strong>navigable small world</strong> graph is one where every node has a few short-range links to near neighbors and a few long-range links to far-away nodes. That mix is what makes a <a href=\"https://en.wikipedia.org/wiki/Small-world_network\">small-world network</a>. The short links keep each neighborhood connected, and the rare long links mean any two nodes are only a few hops apart. Drop a query into such a graph and keep walking to whichever neighbor is closest to the query, and you reach the query’s neighborhood quickly. That greedy walk is the search.</p>\n<p>The catch with a single flat small-world graph is that the greedy walk can be slow to cover distance. Near the start you want huge strides, but the graph only offers whatever links the current node happens to have. That is what the <strong>hierarchy</strong> fixes.</p>\n<h2>How the layers work</h2>\n<p>HNSW stacks several graphs on top of each other. When a vector is inserted, it is assigned a top layer drawn from an exponentially decaying distribution. So most vectors live only on layer 0, a smaller fraction also reach layer 1, fewer still reach layer 2, and so on. Layer 0 contains every vector, and each layer above is a sparse sample of the one below. The exponential decay is the same probabilistic idea behind a <a href=\"https://en.wikipedia.org/wiki/Skip_list\">skip list</a>, lifted from one dimension into a graph.</p>\n<p>Search runs top-down:</p>\n<ol>\n<li>Start at a single entry point in the top, sparsest layer.</li>\n<li>Greedily hop to the neighbor closest to the query until no neighbor is closer. Because this layer is sparse, each hop covers a lot of ground.</li>\n<li>Drop straight down to the same node in the next layer and repeat, now with denser links and shorter hops.</li>\n<li>At layer 0, do the same greedy walk but keep a candidate list of the best nodes seen, and return the top <code class=\"language-text\">k</code>.</li>\n</ol>\n<p>\n  <a\n    class=\"gatsby-resp-image-link\"\n    href=\"/static/63818c89c33313ff7827d0a55b08c3f7/e957c/hnsw-layers.png\"\n    style=\"display: block\"\n    target=\"_blank\"\n    rel=\"noopener\"\n  >\n  \n  <span\n    class=\"gatsby-resp-image-wrapper\"\n    style=\"position: relative; display: block; margin: 7vw 0; max-width: 1360px; margin-left: auto; margin-right: auto;\"\n  >\n    <span\n      class=\"gatsby-resp-image-background-image\"\n      style=\"padding-bottom: 62.64705882352941%; position: relative; bottom: 0; left: 0; background-image: url('data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAABQAAAANCAIAAAAmMtkJAAAACXBIWXMAAAsSAAALEgHS3X78AAABQUlEQVQoz3VSgXKDIAzt//+jZ9d2XtuJoIAQSMKiTjfb+S7HJXBJ3nvHibksQKTbranrj6qqL5fPpnmcz9frtamqc11fJLnfn2WP05Yxs9b9V6taZTozdF3ftlp1Rko5ldJ9b1+bvQvOhrW/eAMMrmAsheRiCsay0Xtpdm4cBr/VmJkxE0aKA4GETcNz1G4Zfkj7GPPyldp+s0/Wwy/tWIimRIIwMSXM4H1ifl8smn00etjqhIV4CuRitaZoMNqUcLqh4sJuwjFtJhG/Z18yHmtGJlplZeEszRP/BeVd9gmJJJZizJBw8hqDsV2dgpEHGQdMPsPsA9sUZMdPszLm0bZ/91sYUwYXNOgK5Of0X9EqEt+YIuU97f9sFACEoJ4zT5psZ5Yf4exorfd+PDSM52EEANrwqkgQQvQ+yDmOcbn5Bpmp/FKkQH/fAAAAAElFTkSuQmCC'); background-size: cover; display: block;\"\n    >\n      <picture>\n        <source\n          srcset=\"/static/63818c89c33313ff7827d0a55b08c3f7/51a8e/hnsw-layers.webp 340w,\n/static/63818c89c33313ff7827d0a55b08c3f7/713b7/hnsw-layers.webp 680w,\n/static/63818c89c33313ff7827d0a55b08c3f7/54376/hnsw-layers.webp 1360w,\n/static/63818c89c33313ff7827d0a55b08c3f7/e0aab/hnsw-layers.webp 1920w\"\n          sizes=\"(max-width: 1360px) 100vw, 1360px\"\n          type=\"image/webp\"\n        />\n        <source\n          srcset=\"/static/63818c89c33313ff7827d0a55b08c3f7/ad208/hnsw-layers.png 340w,\n/static/63818c89c33313ff7827d0a55b08c3f7/a5a26/hnsw-layers.png 680w,\n/static/63818c89c33313ff7827d0a55b08c3f7/60356/hnsw-layers.png 1360w,\n/static/63818c89c33313ff7827d0a55b08c3f7/e957c/hnsw-layers.png 1920w\"\n          sizes=\"(max-width: 1360px) 100vw, 1360px\"\n          type=\"image/png\"\n        />\n        <img\n          class=\"gatsby-resp-image-image\"\n          style=\"width: 100%; height: 100%; margin: 0; vertical-align: middle; position: absolute; top: 0; left: 0; box-shadow: inset 0px 0px 0px 400px white;\"\n          src=\"/static/63818c89c33313ff7827d0a55b08c3f7/60356/hnsw-layers.png\"\n          alt=\"HNSW routes a search from a single entry point in the top sparse layer, hopping greedily toward the query and dropping one layer at a time until it reaches the dense bottom layer where every vector lives\"\n          title=\"\"\n          src=\"/static/63818c89c33313ff7827d0a55b08c3f7/60356/hnsw-layers.png\"\n        />\n      </picture>\n      </span>\n  </span>\n  \n  </a>\n    </p>\n<p>The top layers are a coarse approach that gets you into the right region in a few long hops. The bottom layer is the fine-grained search that finds the actual neighbors. That is why the cost scales with <code class=\"language-text\">log(N)</code>: you are descending a hierarchy, not scanning a list.</p>\n<h2>The three parameters that matter</h2>\n<p>Almost every HNSW implementation exposes the same three knobs under slightly different names, and understanding them is most of what you need to tune an index. The values below are <a href=\"https://github.com/pgvector/pgvector#hnsw\">pgvector’s</a> defaults, but the meanings carry across FAISS, OpenSearch, and the rest.</p>\n<p><strong><code class=\"language-text\">m</code>: links per node (build time).</strong> This sets how many neighbors each node keeps on the graph. Higher <code class=\"language-text\">m</code> means a denser, better-connected graph with higher recall. The cost is more memory (every link is stored) and slower builds. pgvector defaults to 16; the useful range is roughly 5 to 48. This is the parameter that most directly sets your memory footprint, because the graph edges live in RAM alongside the vectors.</p>\n<p><strong><code class=\"language-text\">ef_construction</code>: search width while building (build time).</strong> To insert a node, HNSW runs a search to find its neighbors, and <code class=\"language-text\">ef_construction</code> sets how wide that search is. Bigger means better neighbor choices and a higher-quality graph, at the cost of slower inserts. pgvector defaults to 64. You pay this once, at build time, so it is usually worth setting generously.</p>\n<p><strong><code class=\"language-text\">ef_search</code>: search width at query time (query time).</strong> This is the size of the candidate list the greedy walk keeps at layer 0. It is the runtime accuracy-versus-latency dial, and the one you will actually reach for. Raise it and the walk explores more of the graph: recall climbs toward the exact answer, and latency climbs with it.</p>\n<p><img src=\"/bc679203bd2a28924852e63cb7cd9fc3/recall-latency.svg\" alt=\"As ef_search rises, recall climbs and saturates toward the exact answer while latency keeps rising roughly linearly, so there is a point where extra candidates buy accuracy you can no longer measure but latency you can\"></p>\n<p>The shape of that trade is the important part. Recall saturates: past some <code class=\"language-text\">ef_search</code>, you pay real latency for accuracy gains too small to notice. Latency does not saturate. So the job is to find the elbow of the curve. The only honest way to find it is to measure recall against an exact flat search on your own data, which comes up again in the failure modes below.</p>\n<h2>In practice: pgvector and FAISS</h2>\n<p>If your vectors already live in Postgres, pgvector is the least-friction option, and the parameters map straight onto the section above. In the example, the build-time knobs go on the index, <code class=\"language-text\">ef_search</code> is a session setting, and the query orders results by cosine distance:</p>\n<div class=\"gatsby-highlight\" data-language=\"sql\"><pre class=\"language-sql\"><code class=\"language-sql\"><span class=\"token comment\">-- build-time knobs live on the index</span>\n<span class=\"token keyword\">CREATE</span> <span class=\"token keyword\">INDEX</span> <span class=\"token keyword\">ON</span> chunks\n  <span class=\"token keyword\">USING</span> hnsw <span class=\"token punctuation\">(</span>embedding vector_cosine_ops<span class=\"token punctuation\">)</span>\n  <span class=\"token keyword\">WITH</span> <span class=\"token punctuation\">(</span>m <span class=\"token operator\">=</span> <span class=\"token number\">16</span><span class=\"token punctuation\">,</span> ef_construction <span class=\"token operator\">=</span> <span class=\"token number\">64</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">;</span>\n\n<span class=\"token comment\">-- query-time knob is a session GUC</span>\n<span class=\"token keyword\">SET</span> hnsw<span class=\"token punctuation\">.</span>ef_search <span class=\"token operator\">=</span> <span class=\"token number\">100</span><span class=\"token punctuation\">;</span>\n\n<span class=\"token keyword\">SELECT</span> id<span class=\"token punctuation\">,</span> content\n<span class=\"token keyword\">FROM</span> chunks\n<span class=\"token keyword\">ORDER</span> <span class=\"token keyword\">BY</span> embedding <span class=\"token operator\">&lt;=></span> :query_vector   <span class=\"token comment\">-- &lt;=> is cosine distance</span>\n<span class=\"token keyword\">LIMIT</span> <span class=\"token number\">50</span><span class=\"token punctuation\">;</span></code></pre></div>\n<p>One detail deserves attention: the operator class, here <code class=\"language-text\">vector_cosine_ops</code>, has to match the distance your embeddings were made for. Suppose your embedding model produces vectors meant for cosine similarity, and you build the index with L2 distance (<code class=\"language-text\">vector_l2_ops</code>). The index is “working,” but it is quietly ranking by the wrong metric. That mismatch stays invisible until you measure retrieval quality.</p>\n<p>For a standalone index, FAISS exposes the same graph directly. Note that the positional <code class=\"language-text\">16</code> passed to the constructor is <code class=\"language-text\">m</code>:</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\"><span class=\"token keyword\">import</span> faiss\n\nd <span class=\"token operator\">=</span> <span class=\"token number\">768</span>                          <span class=\"token comment\"># embedding dimension</span>\nindex <span class=\"token operator\">=</span> faiss<span class=\"token punctuation\">.</span>IndexHNSWFlat<span class=\"token punctuation\">(</span>d<span class=\"token punctuation\">,</span> <span class=\"token number\">16</span><span class=\"token punctuation\">)</span>      <span class=\"token comment\"># 16 == m</span>\nindex<span class=\"token punctuation\">.</span>hnsw<span class=\"token punctuation\">.</span>efConstruction <span class=\"token operator\">=</span> <span class=\"token number\">64</span>\nindex<span class=\"token punctuation\">.</span>add<span class=\"token punctuation\">(</span>doc_vectors<span class=\"token punctuation\">)</span>                  <span class=\"token comment\"># np.float32, shape (N, d)</span>\n\nindex<span class=\"token punctuation\">.</span>hnsw<span class=\"token punctuation\">.</span>efSearch <span class=\"token operator\">=</span> <span class=\"token number\">100</span>               <span class=\"token comment\"># the query-time dial</span>\ndistances<span class=\"token punctuation\">,</span> ids <span class=\"token operator\">=</span> index<span class=\"token punctuation\">.</span>search<span class=\"token punctuation\">(</span>query_vectors<span class=\"token punctuation\">,</span> k<span class=\"token operator\">=</span><span class=\"token number\">50</span><span class=\"token punctuation\">)</span></code></pre></div>\n<p>Same three numbers, same meanings. FAISS’s own <a href=\"https://github.com/facebookresearch/faiss/wiki/Guidelines-to-choose-an-index\">guidance on choosing an index</a> is a good second read once you are deciding between HNSW and the quantized variants for a large corpus.</p>\n<h2>Failure modes</h2>\n<p><strong>Deletes do not really delete.</strong> HNSW is built for insert-and-search, not churn. Most implementations handle a delete by tombstoning the node (marking it deleted while leaving it in the graph) rather than removing it. Unpicking a well-connected node and rewiring its neighbors is expensive. In a corpus that updates often, tombstones accumulate, the graph degrades, and recall drifts down. I hit exactly this tension when I wrote about <a href=\"/blog/2026-06-29-incremental-rag-indexing/\">keeping a RAG index fresh</a>. The clean answer is often a periodic rebuild rather than an endless stream of in-place deletes.</p>\n<p><strong>The index is a memory cost, not just a disk cost.</strong> HNSW keeps the graph, and usually the vectors, in RAM to hit its latency numbers. A denser graph (higher <code class=\"language-text\">m</code>) takes more memory. It is easy to size a box for the raw vectors, forget the graph edges on top, and then watch the process get killed under load. Budget for both before you pick <code class=\"language-text\">m</code>.</p>\n<p><strong>“Recall” you never measured.</strong> ANN is approximate by definition. The amount of approximation is a config value you chose, sometimes by accident, by taking a library default. The only way to know your real recall is to run a sample of queries through both the HNSW index and an exact flat search, then compare the results. This is the same golden-set discipline I described in <a href=\"/blog/2026-07-04-measuring-rag-retrieval-quality/\">measuring RAG retrieval quality</a>. Your <code class=\"language-text\">ef_search</code> matters because it silently sets the recall those metrics report.</p>\n<p><strong>Tuning against the wrong stage.</strong> If the right chunk is missing from your results, the index is only one of the suspects. Bad chunking can mean the answer was never a clean vector to begin with. A metric mismatch can mean the graph is ranking by the wrong distance. Before you turn <code class=\"language-text\">ef_search</code> up, confirm the vector is even in the corpus and that a flat search finds it. An ANN index can only fail to find what an exact search would have found. If the exact search misses it too, the bug is upstream.</p>\n<h2>What I would do differently</h2>\n<p>My mistake was treating the vector index as infrastructure that either works or does not, and reaching for it as the first knob when retrieval quality dropped. The index is a dial, not a switch, and its default position was chosen by a library author who had never seen my data. Now, before the index goes near production, I:</p>\n<ol>\n<li>build a small golden set of query-to-expected-chunk pairs;</li>\n<li>measure recall against a flat search;</li>\n<li>only then decide whether the default <code class=\"language-text\">ef_search</code> is fine, or whether I am silently dropping a tenth of the right answers.</li>\n</ol>\n<p>Doing that first turns HNSW from a black box into one more measurable stage in the pipeline.</p>\n<h2>Why it matters: approximation is a number you set</h2>\n<p>HNSW is the quiet layer under every RAG post I have written: the thing that makes “find the nearest chunks” cheap enough to do on every request. It is the retrieval engine inside <a href=\"/project/archi/\">Archi</a>, the RAG copilot I worked on for CMS computing operations at CERN. It is the reason a query against millions of indexed log lines and tickets comes back in milliseconds instead of scanning the lot.</p>\n<p>The layered graph is a genuinely elegant piece of engineering. But the part that matters for a production system is smaller and less glamorous. The index is approximate, the amount of approximation is a number you set, and you do not know that number until you measure it. Once you treat the index as a stage you can score, the same way you score <a href=\"/blog/2026-07-06-chunking-strategies-for-rag/\">chunking</a> and <a href=\"/blog/2026-07-08-cross-encoder-reranking-rag/\">reranking</a>, it stops being a black box and starts being tunable.</p>\n<hr>\n<p><em>Diagrams by M. Hassan Ahmed, released under CC0. Image credit: original work by the author.</em></p>","frontmatter":{"title":"How HNSW Vector Search Actually Works","date":"2026-07-09T00:00:00.000Z","description":"Every RAG stack leans on HNSW but treats it as a black box. Here is how the layered graph index finds nearest neighbors fast, and the knobs that matter.","thumbnail":{"childImageSharp":{"fluid":{"base64":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAABQAAAALCAIAAADwazoUAAAACXBIWXMAAAsSAAALEgHS3X78AAAB60lEQVQoz1WR2W6bQBSGuWlihlmYzWyGAduYMWCb2E4cF2ibSr1JojR5/5fpIbmq9OnXz4Gz4gjb86qX22EzvNnxvezfVtcX8MvrK/hqfAe/ur7aHx/bnx/V+FfVI7e9+MRBNEYk9lhid9f2MDT7vt5d6/Zx2z7a5gK6O/Rlddo2l3b3fWPPHokQmoN6NHE8YQDEU9dPkZ8hnlG9JFFN4wYgUfPlJwMatyw9CDvQyHo8c/LymC47qgrMU89PMc+C4k7VT2IziE0vyn4y1SCrSafI+irsSOZrjy0cKnPIZKpA/gKxBVVrokqi1kyXPKyZ3lC5+g9eYGGwLBCMzVRORDaDTVjySQzN/ahm4RaUx41snmT7G3pKO05Ug9r/YYsWkcjJVl263PMAqhqXRFBVxDs/rAERtz6QdgxWBaIGoFGD1colIVza+Yb0jRdAQ3hwKSTnMBILrB81MtlJc+Jp58+tH1geNRBkwZbqksglfAyd94mpPX8aGOq5OIB3KjvPi4s2Dyq/hHfPYfes7S9h7qd4ftH5A5a5SwLY2fg6Z9J4kIzncDBl7nV2VtlJGdCzNvfBZgyPL6Z90ulRpkf4STOsp+QbpIBbT99iPcPB182+jufSmOkiXnfJqsMULhoneYN5cuupGQlmeP4PTIRMIYsQN60AAAAASUVORK5CYII=","aspectRatio":1.899441340782123,"src":"/static/19f8cb970a43134f4a3f984c45f292f7/40a76/hero.png","srcSet":"/static/19f8cb970a43134f4a3f984c45f292f7/c972b/hero.png 340w,\n/static/19f8cb970a43134f4a3f984c45f292f7/27625/hero.png 680w,\n/static/19f8cb970a43134f4a3f984c45f292f7/40a76/hero.png 1360w,\n/static/19f8cb970a43134f4a3f984c45f292f7/ed396/hero.png 2000w","sizes":"(max-width: 1360px) 100vw, 1360px"}}}}}},"pageContext":{"slug":"/2026-07-09-hnsw-vector-search-explained/","previous":"blog/2026-07-10-fastapi-blocking-event-loop/","next":"blog/2026-07-12-llm-kv-cache-gpu-memory/"}},"staticQueryHashes":["32046230"]}