{"componentChunkName":"component---src-templates-blog-post-js","path":"/blog/2026-09-11-cosine-dot-product-euclidean-vector-search/","result":{"data":{"site":{"siteMetadata":{"title":"M.Hassan Ahmed","author":"Hassan11196"}},"markdownRemark":{"id":"f283d1b6-cb95-5855-a273-a2b67a3c9b2e","excerpt":"Every vector database asks you to pick a distance metric when you create the index. Most tutorials wave past the choice with “use cosine,” and that works until…","html":"<p>Every vector database asks you to pick a distance metric when you create the index. Most tutorials wave past the choice with “use cosine,” and that works until the day it quietly doesn’t. You swap in a new embedding model, keep the same index settings, and retrieval gets subtly worse. Long documents start crowding the top results. Or two indexes built from the same corpus return different neighbors, and nobody can say why.</p>\n<p>The metric matters, but not in the way the question is usually framed. Cosine similarity, dot product (also called inner product), and Euclidean distance are not three competing ideas of “close.” On normalized vectors they produce the <em>same ranking</em>. The real decision is about what your embeddings look like and what each engine costs you at query time.</p>\n<p>This post is for anyone wiring up a RAG (retrieval-augmented generation) retriever or a k-NN (k-nearest-neighbor) index who wants to choose deliberately instead of copying a default. I led the retrieval side of <a href=\"/project/archi/\">Archi</a>, a RAG copilot for CMS operations at CERN, and this is the mental model I used when someone asked which <code class=\"language-text\">space_type</code> to set. It covers what each metric measures, why normalization makes them agree, how to match the metric to the model, and how pgvector, OpenSearch, and FAISS each expect you to configure it.</p>\n<h2>What each metric measures</h2>\n<p>Take a query vector <code class=\"language-text\">q</code> and a document vector <code class=\"language-text\">d</code>, each a list of, say, 768 floats.</p>\n<p><strong>Dot product</strong> is the raw inner product, <code class=\"language-text\">sum(q[i] * d[i])</code>. It grows with the angle between the vectors <em>and</em> with their lengths. A vector with a large magnitude can score highly just by being big.</p>\n<p><strong>Cosine similarity</strong> divides the dot product by both magnitudes: <code class=\"language-text\">dot(q, d) / (norm(q) * norm(d))</code>. That division cancels magnitude, so cosine measures only the angle between the vectors. Its value lies in <code class=\"language-text\">[-1, 1]</code>.</p>\n<p><strong>Euclidean distance</strong> (L2) is the straight-line gap between the tips of the two vectors, <code class=\"language-text\">sqrt(sum((q[i] - d[i])^2))</code>. Smaller means closer. That makes it a distance, not a similarity, so you sort ascending instead of descending.</p>\n<p>The difference that trips people up is magnitude. Cosine throws it away; the other two keep it. The diagram below shows where that matters.</p>\n<p><img src=\"/e73820df476a6d36788b4d5157a60236/metrics-geometry.svg\" alt=\"A vector diagram. From a shared origin, a blue query vector q points up and to the right. Two document vectors point in the same direction as each other but at a slightly wider angle from q: a short black vector labelled doc A and a long orange vector labelled doc B that is more than twice its length. Dashed grey lines mark the straight-line Euclidean gaps from the tip of q to each document tip, and a small arc at the origin marks the angle theta. A panel on the right summarizes three metrics: Cosine measures angle only, so cosine of q and A equals cosine of q and B and magnitude is ignored; Dot product is angle times magnitude, so because B is longer the dot product of q and B exceeds that of q and A, and a high-norm vector can win on length alone; Euclidean is the straight-line gap between the tips and also shifts with magnitude. A caption reads: same direction, different length, only cosine treats A and B as equal.\"></p>\n<p>Document A and document B point in exactly the same direction. Cosine sees only the angle, so it scores them as equally relevant to the query. Dot product ranks B above A purely because B is longer.</p>\n<p>Suppose your embedding model happens to give longer or more “confident” documents a larger norm. Dot product will then systematically float those documents to the top, whether or not they answer the question. That is the classic footgun: using dot product on embeddings you never normalized.</p>\n<h2>Normalization makes the three metrics agree</h2>\n<p>Now normalize every vector to unit length before you store it: divide each one by its own magnitude so <code class=\"language-text\">norm(v) == 1</code>.</p>\n<p>Once every vector sits on the unit sphere, cosine and dot product become the same operation. Cosine is <em>defined</em> as the dot product of normalized vectors, so if the vectors are already normalized, dividing by 1 does nothing. Euclidean distance falls in line too, through one identity worth memorizing:</p>\n<div class=\"gatsby-highlight\" data-language=\"text\"><pre class=\"language-text\"><code class=\"language-text\">||a - b||^2 = ||a||^2 + ||b||^2 - 2(a . b)\n            = 2 - 2(a . b)     when |a| = |b| = 1</code></pre></div>\n<p>The identity says that for unit vectors, Euclidean distance squared is just <code class=\"language-text\">2 - 2 * cosine</code>. As cosine goes up, Euclidean distance goes down, always. So sorting by Euclidean distance gives exactly the same order as sorting by cosine (<a href=\"https://en.wikipedia.org/wiki/Cosine_similarity#Properties\">Wikipedia’s cosine similarity page</a> spells out this relationship). On normalized vectors, “nearest by L2,” “most similar by cosine,” and “largest dot product” are three names for the same sorted list.</p>\n<p><img src=\"/74d804cd9fdd0d421ab46bc2f7989489/normalize-and-engines.svg\" alt=\"A unit circle on the left with two blue vectors a and b drawn from the center to points on the circle, an orange dashed chord connecting their tips labelled the norm of a minus b, and the identity below it: the norm of a minus b squared equals 2 minus 2 times a dot b when both are unit length. On the right, a table with columns Ranking, pgvector, OpenSearch, and FAISS. The Cosine row maps to the pgvector operator angle-bracket-equals, OpenSearch space_type cosinesimil, and FAISS normalize plus inner product. The Dot product row maps to the pgvector operator angle-hash, OpenSearch innerproduct, and FAISS METRIC_INNER_PRODUCT. The Euclidean row maps to the pgvector operator angle-dash, OpenSearch l2, and FAISS METRIC_L2. Below the table: practical default, normalize vectors at index time then use inner product, you get cosine ranking without the per-query normalization cost. A footnote notes that pgvector&#x27;s angle-hash operator returns the negative inner product.\"></p>\n<p>So “which metric?” is really two questions:</p>\n<ol>\n<li><strong>Are your vectors normalized?</strong> If yes, the metrics are interchangeable for ranking, and you pick on cost.</li>\n<li><strong>If not, how was the model trained?</strong> Un-normalized, the metrics genuinely differ, and you need to match the metric to the model’s training objective.</li>\n</ol>\n<h2>Use the similarity the model was trained with</h2>\n<p>Each embedding model is trained with a particular similarity function in its loss, and that is the one you should score with. Most modern sentence-embedding models are trained for cosine, and the model card says so.</p>\n<p>The <code class=\"language-text\">sentence-transformers</code> docs are explicit on this point. When a model outputs unit-length vectors, “dot-product and cosine-similarity are equivalent,” and dot product is preferred because it skips a redundant re-normalization (<a href=\"https://sbert.net/docs/sentence_transformer/usage/semantic_textual_similarity.html\">Semantic Textual Similarity</a>).</p>\n<p>A handful of retrieval models, including some dense passage retrievers, are trained for raw dot product. They deliberately encode importance into the vector’s norm. For those models, normalizing throws away signal the model meant to keep.</p>\n<p>The rule is boring and reliable: read the model card and use the similarity it was trained with.</p>\n<ul>\n<li><strong>The card says cosine:</strong> normalize, and you are done.</li>\n<li><strong>The card says dot product and the model doesn’t produce unit vectors:</strong> don’t normalize.</li>\n<li><strong>Euclidean:</strong> rarely the stated objective for text embeddings, though it is the natural default for other vector types. Once vectors are normalized, it gives the same ranking as cosine anyway.</li>\n</ul>\n<h2>How pgvector, OpenSearch, and FAISS spell the metric</h2>\n<p>The metric is an abstract choice, but every engine spells it differently, and a couple of the spellings have sharp edges.</p>\n<p><strong>pgvector</strong> exposes the metric as query operators rather than a config value:</p>\n<ul>\n<li><code class=\"language-text\">&lt;-&gt;</code> is L2 distance.</li>\n<li><code class=\"language-text\">&lt;=&gt;</code> is cosine distance.</li>\n<li><code class=\"language-text\">&lt;#&gt;</code> is the inner product, except that it returns the <em>negative</em> inner product.</li>\n</ul>\n<p>The sign flip exists because Postgres index scans only go ascending. Negating turns “largest dot product” into “smallest number” (<a href=\"https://github.com/pgvector/pgvector#querying\">pgvector README</a>). It is easy to miss, and forgetting it inverts your results. Both queries below sort ascending, and both return the most similar rows first:</p>\n<div class=\"gatsby-highlight\" data-language=\"sql\"><pre class=\"language-sql\"><code class=\"language-sql\"><span class=\"token comment\">-- cosine distance: smaller is more similar</span>\n<span class=\"token keyword\">SELECT</span> id <span class=\"token keyword\">FROM</span> docs <span class=\"token keyword\">ORDER</span> <span class=\"token keyword\">BY</span> embedding <span class=\"token operator\">&lt;=></span> query_vec <span class=\"token keyword\">LIMIT</span> <span class=\"token number\">10</span><span class=\"token punctuation\">;</span>\n\n<span class=\"token comment\">-- inner product: &lt;#> is NEGATIVE dot product, so ascending sort</span>\n<span class=\"token comment\">-- still returns the largest-dot-product rows first</span>\n<span class=\"token keyword\">SELECT</span> id <span class=\"token keyword\">FROM</span> docs <span class=\"token keyword\">ORDER</span> <span class=\"token keyword\">BY</span> embedding <span class=\"token operator\">&lt;</span><span class=\"token comment\">#> query_vec LIMIT 10;</span></code></pre></div>\n<p><strong>OpenSearch</strong> takes a <code class=\"language-text\">space_type</code> on the k-NN field mapping. The values are <code class=\"language-text\">l2</code>, <code class=\"language-text\">cosinesimil</code>, and <code class=\"language-text\">innerproduct</code>. OpenSearch ranks on descending score, so it converts each distance into a score where higher is better. Two details matter here:</p>\n<ul>\n<li>With the Faiss engine, <code class=\"language-text\">cosinesimil</code> normalizes your vectors to unit length at index time for you.</li>\n<li>The docs themselves suggest that if your vectors are already normalized, you should use <code class=\"language-text\">innerproduct</code> instead of <code class=\"language-text\">cosinesimil</code>. You get the same ranking with explicit control and less work (<a href=\"https://docs.opensearch.org/latest/mappings/supported-field-types/knn-spaces/\">OpenSearch vector spaces</a>).</li>\n</ul>\n<p>The mapping below follows that advice: a 768-dimension vector field scored with inner product, which assumes you normalize before indexing.</p>\n<div class=\"gatsby-highlight\" data-language=\"json\"><pre class=\"language-json\"><code class=\"language-json\">PUT /docs\n<span class=\"token punctuation\">{</span>\n  <span class=\"token property\">\"settings\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">{</span> <span class=\"token property\">\"index.knn\"</span><span class=\"token operator\">:</span> <span class=\"token boolean\">true</span> <span class=\"token punctuation\">}</span><span class=\"token punctuation\">,</span>\n  <span class=\"token property\">\"mappings\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">{</span>\n    <span class=\"token property\">\"properties\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">{</span>\n      <span class=\"token property\">\"embedding\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">{</span>\n        <span class=\"token property\">\"type\"</span><span class=\"token operator\">:</span> <span class=\"token string\">\"knn_vector\"</span><span class=\"token punctuation\">,</span>\n        <span class=\"token property\">\"dimension\"</span><span class=\"token operator\">:</span> <span class=\"token number\">768</span><span class=\"token punctuation\">,</span>\n        <span class=\"token property\">\"space_type\"</span><span class=\"token operator\">:</span> <span class=\"token string\">\"innerproduct\"</span>\n      <span class=\"token punctuation\">}</span>\n    <span class=\"token punctuation\">}</span>\n  <span class=\"token punctuation\">}</span>\n<span class=\"token punctuation\">}</span></code></pre></div>\n<p><strong>FAISS</strong> offers only <code class=\"language-text\">METRIC_L2</code> and <code class=\"language-text\">METRIC_INNER_PRODUCT</code>. It has no cosine metric at all. The <a href=\"https://github.com/facebookresearch/faiss/wiki/MetricType-and-distances\">FAISS wiki</a> tells you to normalize the vectors with <code class=\"language-text\">faiss.normalize_L2</code> and then use inner product, which is exactly the equivalence from earlier. The whole trick is to normalize in place before both <code class=\"language-text\">add</code> and <code class=\"language-text\">search</code>:</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\"><span class=\"token keyword\">import</span> faiss\n\nfaiss<span class=\"token punctuation\">.</span>normalize_L2<span class=\"token punctuation\">(</span>doc_vecs<span class=\"token punctuation\">)</span>          <span class=\"token comment\"># unit length in place</span>\nindex <span class=\"token operator\">=</span> faiss<span class=\"token punctuation\">.</span>IndexFlatIP<span class=\"token punctuation\">(</span><span class=\"token number\">768</span><span class=\"token punctuation\">)</span>        <span class=\"token comment\"># inner product</span>\nindex<span class=\"token punctuation\">.</span>add<span class=\"token punctuation\">(</span>doc_vecs<span class=\"token punctuation\">)</span>\n\nfaiss<span class=\"token punctuation\">.</span>normalize_L2<span class=\"token punctuation\">(</span>query_vecs<span class=\"token punctuation\">)</span>        <span class=\"token comment\"># normalize the query too</span>\nscores<span class=\"token punctuation\">,</span> ids <span class=\"token operator\">=</span> index<span class=\"token punctuation\">.</span>search<span class=\"token punctuation\">(</span>query_vecs<span class=\"token punctuation\">,</span> k<span class=\"token operator\">=</span><span class=\"token number\">10</span><span class=\"token punctuation\">)</span></code></pre></div>\n<p>The line people forget is normalizing the <em>query</em>. If you normalize the corpus but skip the query, every score gets scaled by the query’s arbitrary magnitude. The ranking within a single query survives that. Cross-query score thresholds, and anything that fuses scores across retrievers, stop meaning anything.</p>\n<h2>Failure modes I’ve actually hit</h2>\n<p><strong>Dot product on un-normalized vectors.</strong> This is the most common one. Retrieval looks fine on a demo corpus, then long documents dominate in production because their embeddings carry a larger norm. Switch to cosine, or normalize and keep dot product, and the length bias disappears.</p>\n<p><strong>Half-normalized indexes.</strong> You normalize at index time but not at query time, or you normalize new documents while an older segment of the index went in raw. The index returns plausible-looking neighbors that are wrong, and no error tells you. Normalization has to be a property of the whole pipeline: the same function, on ingest and on query, applied to every vector.</p>\n<p><strong>Comparing scores across metrics.</strong> Cosine sits in <code class=\"language-text\">[-1, 1]</code>, raw dot product is unbounded, and L2 is a distance where zero is perfect. A “relevance > 0.8” cutoff tuned against a cosine index is meaningless on an L2 index. Thresholds are per-metric, and honestly per-model too.</p>\n<p>This bites hardest in <a href=\"/blog/2026-09-02-reciprocal-rank-fusion-hybrid-search/\">hybrid search</a>. Naively adding a cosine score to a BM25 keyword score combines two numbers on incompatible scales. That is one reason rank-based fusion behaves better than adding raw scores together.</p>\n<p><strong>Assuming the metric changes recall.</strong> On normalized vectors it doesn’t: the ranking is identical. What <em>does</em> change recall is the approximate index structure underneath, like the graph in <a href=\"/blog/2026-07-09-hnsw-vector-search-explained/\">HNSW</a>. The metric and the ANN (approximate nearest neighbor) algorithm are separate knobs, so swapping cosine for L2 to chase a recall regression is looking in the wrong place.</p>\n<h2>Tradeoffs and my default</h2>\n<p>If the metrics rank identically once vectors are normalized, the tiebreakers are cost and clarity. Dot product on pre-normalized vectors is the cheapest to compute at query time: there is no per-query square root or division, because you did that work once at index time. That is why the engines nudge you toward it, and why I default there:</p>\n<p><strong>Normalize at index time, store unit vectors, query with inner product.</strong></p>\n<p>You get cosine’s ranking, dot product’s speed, and one invariant that is easy to test: every stored vector has norm 1.</p>\n<p>There are two exceptions. If a model was explicitly trained for raw dot product, the norm carries meaning, so leave the vectors alone. And if you are on pgvector and would rather not think about the <code class=\"language-text\">&lt;#&gt;</code> sign convention, <code class=\"language-text\">&lt;=&gt;</code> cosine distance is perfectly fine. On a small index, the performance gap is not worth a subtle bug.</p>\n<p>Looking back at early retrieval work, here is what I would do differently: decide normalization once, at the boundary where embeddings enter the system, and assert it. A single <code class=\"language-text\">assert abs(np.linalg.norm(v) - 1.0) &lt; 1e-6</code> on ingest would have caught more retrieval weirdness than any amount of metric-swapping. The metric is a small decision. Whether your vectors are actually normalized, everywhere, is what silently decides whether search works.</p>\n<p>The metric question sits downstream of picking the embedding model itself, which I covered in <a href=\"/blog/2026-07-28-choosing-embedding-model-rag/\">choosing an embedding model for RAG</a>. Get both settings right before you index a few million documents and have to redo it.</p>\n<hr>\n<p><em>Diagrams by M. Hassan Ahmed, released under CC0. No external image was used for this post; the figures are original work by the author.</em></p>","frontmatter":{"title":"Cosine, Dot Product, or Euclidean for Embeddings?","date":"2026-09-11T00:00:00.000Z","description":"Cosine, dot product, and Euclidean rank vector search results the same on normalized embeddings, differently otherwise. How to pick and configure one.","thumbnail":{"childImageSharp":{"fluid":{"base64":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAABQAAAALCAIAAADwazoUAAAACXBIWXMAAAsSAAALEgHS3X78AAAByElEQVQozz2R3XKbMBCFuQsgrYQACRAIgUAY/+DYvohjJ+1Fmk7f/426Np3OfLNzNNpztAsBnc7whPoT1vHHH/v+5T5/u4/v7vbd37+7+68Hty+8wiPfXtZONAaENf/gTVb4ze42bt6W08/98dPPb6i3hzuyWz7mw93P11xN/y0BSVuS2li0CM06mlrjTpDaEGqStHW/2PEMeR/yRjbbYXu1/gJ5t7oCkhiEZR2XLmY1l305vEI5EtkT5WS/6OnMEhODFnlfmYPUMxXt6gpi3iBEtIkaUfBikMd3cAv0B+j2dHwVy5W6PTFzJLtYuTh59j/MDZprJGL6mddEUBXmyFILiYGk4WmfyJFKB3qi3Y7Pl0R77AlpiTWImV4BXAODmFZmD4/lW4xjueNyoLiXeMQJOcasSrVv5kvWTEHEKiSEkgicR6NOK9zqMRUCmWWyx1wUGCcKjz2QtVnlWW6DCHCAknDNMrvqVHmKHwKqiJaUG3w2JAUeMQsHec5cvBCFFc1FSBUVNWQGRZxoIlsiDdOO10NSD8J4BAVVLRT4k2q0rM8EaEBeiERCoiKhadkhotvkwy7r59zv1XTI3ExLSytL8ibEzqfrL2TMQD9kW0U8AAAAAElFTkSuQmCC","aspectRatio":1.899441340782123,"src":"/static/253f7d1c56e9e7b3fd4601ef40351ea0/40a76/hero.png","srcSet":"/static/253f7d1c56e9e7b3fd4601ef40351ea0/c972b/hero.png 340w,\n/static/253f7d1c56e9e7b3fd4601ef40351ea0/27625/hero.png 680w,\n/static/253f7d1c56e9e7b3fd4601ef40351ea0/40a76/hero.png 1360w,\n/static/253f7d1c56e9e7b3fd4601ef40351ea0/ed396/hero.png 2000w","sizes":"(max-width: 1360px) 100vw, 1360px"}}}}}},"pageContext":{"slug":"/2026-09-11-cosine-dot-product-euclidean-vector-search/","previous":"blog/2026-09-07-parsing-pdfs-for-rag/","next":"blog/2026-09-10-maximal-marginal-relevance-rag/"}},"staticQueryHashes":["32046230"]}