{"componentChunkName":"component---src-templates-blog-post-js","path":"/blog/2026-08-03-pgvector-rag-postgres/","result":{"data":{"site":{"siteMetadata":{"title":"M.Hassan Ahmed","author":"Hassan11196"}},"markdownRemark":{"id":"db6e6998-df8f-5abb-9920-5537638f1d93","excerpt":"Most RAG (retrieval-augmented generation) tutorials reach for a dedicated vector database on the first page: Pinecone, Qdrant, Weaviate, Milvus. You pick one…","html":"<p>Most RAG (retrieval-augmented generation) tutorials reach for a dedicated vector database on the first page: Pinecone, Qdrant, Weaviate, Milvus. You pick one, run it next to your app, sync your data into it, and now you have two stores to keep consistent. That’s a reasonable choice at large scale. It’s also a lot of infrastructure to stand up before you’ve proven the retrieval is any good.</p>\n<p>If you already run Postgres (and for the backends I build, I almost always do), you can skip that second system entirely. <a href=\"https://github.com/pgvector/pgvector\">pgvector</a> is a Postgres extension that adds a <code class=\"language-text\">vector</code> column type and approximate nearest-neighbour indexes. Your embeddings live in the same table as the chunk text and the metadata you filter on. The same <code class=\"language-text\">pg_dump</code> backs them up, and the same <code class=\"language-text\">BEGIN</code>/<code class=\"language-text\">COMMIT</code> makes them transactional. Retrieval becomes one SQL query.</p>\n<p>This post is for engineers building a RAG pipeline who want to know whether Postgres is enough before they commit to a separate vector store. It covers the schema, how the distance operators behave, when to use an HNSW index versus IVFFlat, how metadata filtering interacts with the index, and the failure modes that bit me. The examples use FastAPI, because that’s the backend I use in <a href=\"/project/cloud-canvas-ai/\">CloudCanvasAI</a> and <a href=\"/project/archi/\">Archi</a>, the CMS operations copilot I worked on.</p>\n<h2>Where pgvector sits in a RAG pipeline</h2>\n<p>RAG has two halves:</p>\n<ul>\n<li><strong>Indexing</strong> happens once, ahead of time. <a href=\"/blog/2026-06-29-incremental-rag-indexing/\">Chunking and indexing</a> turn your documents into embeddings.</li>\n<li><strong>Retrieval</strong> happens at request time. It embeds the user’s question and finds the chunks whose vectors sit closest to the question’s vector.</li>\n</ul>\n<p>pgvector owns the storage and the retrieval half. The diagram below shows the request path. Nothing in it talks to a vector service: the “vector store” is a column and an index inside the database you already operate.</p>\n<p><img src=\"/bdf651bc713a6b32e6c778b4f9cadd76/query-flow.svg\" alt=\"A user question is turned into a query vector by an embedding model, then passed into PostgreSQL. Inside Postgres, a chunks table holds id, content, metadata, and an embedding vector(1536) column, with an HNSW index that orders rows by cosine distance and returns the top five. Those top-k chunks and their metadata go into an LLM prompt to produce a grounded answer.\"></p>\n<h2>Setup and schema</h2>\n<p>Install the extension, then create it once per database:</p>\n<div class=\"gatsby-highlight\" data-language=\"sql\"><pre class=\"language-sql\"><code class=\"language-sql\"><span class=\"token keyword\">CREATE</span> EXTENSION <span class=\"token keyword\">IF</span> <span class=\"token operator\">NOT</span> <span class=\"token keyword\">EXISTS</span> vector<span class=\"token punctuation\">;</span></code></pre></div>\n<p>Next, create a table. The dimension in <code class=\"language-text\">vector(1536)</code> has to match your embedding model exactly. 1536 is the width of OpenAI’s <code class=\"language-text\">text-embedding-3-small</code>. If you switch models, you re-embed everything, so pin this down before you ingest at scale.</p>\n<div class=\"gatsby-highlight\" data-language=\"sql\"><pre class=\"language-sql\"><code class=\"language-sql\"><span class=\"token keyword\">CREATE</span> <span class=\"token keyword\">TABLE</span> chunks <span class=\"token punctuation\">(</span>\n    id         <span class=\"token keyword\">BIGINT</span> GENERATED ALWAYS <span class=\"token keyword\">AS</span> <span class=\"token keyword\">IDENTITY</span> <span class=\"token keyword\">PRIMARY</span> <span class=\"token keyword\">KEY</span><span class=\"token punctuation\">,</span>\n    document   <span class=\"token keyword\">TEXT</span>        <span class=\"token operator\">NOT</span> <span class=\"token boolean\">NULL</span><span class=\"token punctuation\">,</span>   <span class=\"token comment\">-- source doc id, for citations</span>\n    content    <span class=\"token keyword\">TEXT</span>        <span class=\"token operator\">NOT</span> <span class=\"token boolean\">NULL</span><span class=\"token punctuation\">,</span>   <span class=\"token comment\">-- the chunk itself</span>\n    metadata   JSONB       <span class=\"token operator\">NOT</span> <span class=\"token boolean\">NULL</span> <span class=\"token keyword\">DEFAULT</span> <span class=\"token string\">'{}'</span><span class=\"token punctuation\">,</span>\n    embedding  VECTOR<span class=\"token punctuation\">(</span><span class=\"token number\">1536</span><span class=\"token punctuation\">)</span>\n<span class=\"token punctuation\">)</span><span class=\"token punctuation\">;</span></code></pre></div>\n<p>Keeping <code class=\"language-text\">content</code> and <code class=\"language-text\">metadata</code> in the same row as <code class=\"language-text\">embedding</code> is the whole point. When a query returns, you already have the text to feed the model and the source id to cite. There is no second lookup keyed on an external vector-store id that you have to keep in sync.</p>\n<h2>Writing embeddings from FastAPI</h2>\n<p>pgvector’s Python bindings register the <code class=\"language-text\">vector</code> type with your database driver, so you can pass a plain Python list. With <code class=\"language-text\">asyncpg</code> and the <a href=\"https://github.com/pgvector/pgvector-python\"><code class=\"language-text\">pgvector</code></a> package, an insert helper looks like this:</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\"><span class=\"token keyword\">import</span> asyncpg\n<span class=\"token keyword\">from</span> pgvector<span class=\"token punctuation\">.</span>asyncpg <span class=\"token keyword\">import</span> register_vector\n\n<span class=\"token keyword\">async</span> <span class=\"token keyword\">def</span> <span class=\"token function\">store_chunks</span><span class=\"token punctuation\">(</span>pool<span class=\"token punctuation\">,</span> doc_id<span class=\"token punctuation\">:</span> <span class=\"token builtin\">str</span><span class=\"token punctuation\">,</span> chunks<span class=\"token punctuation\">:</span> <span class=\"token builtin\">list</span><span class=\"token punctuation\">[</span><span class=\"token builtin\">str</span><span class=\"token punctuation\">]</span><span class=\"token punctuation\">,</span> vectors<span class=\"token punctuation\">:</span> <span class=\"token builtin\">list</span><span class=\"token punctuation\">[</span><span class=\"token builtin\">list</span><span class=\"token punctuation\">[</span><span class=\"token builtin\">float</span><span class=\"token punctuation\">]</span><span class=\"token punctuation\">]</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n    <span class=\"token keyword\">async</span> <span class=\"token keyword\">with</span> pool<span class=\"token punctuation\">.</span>acquire<span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span> <span class=\"token keyword\">as</span> conn<span class=\"token punctuation\">:</span>\n        <span class=\"token keyword\">await</span> register_vector<span class=\"token punctuation\">(</span>conn<span class=\"token punctuation\">)</span>\n        <span class=\"token keyword\">await</span> conn<span class=\"token punctuation\">.</span>executemany<span class=\"token punctuation\">(</span>\n            <span class=\"token string\">\"INSERT INTO chunks (document, content, embedding) VALUES ($1, $2, $3)\"</span><span class=\"token punctuation\">,</span>\n            <span class=\"token punctuation\">[</span><span class=\"token punctuation\">(</span>doc_id<span class=\"token punctuation\">,</span> text<span class=\"token punctuation\">,</span> vec<span class=\"token punctuation\">)</span> <span class=\"token keyword\">for</span> text<span class=\"token punctuation\">,</span> vec <span class=\"token keyword\">in</span> <span class=\"token builtin\">zip</span><span class=\"token punctuation\">(</span>chunks<span class=\"token punctuation\">,</span> vectors<span class=\"token punctuation\">)</span><span class=\"token punctuation\">]</span><span class=\"token punctuation\">,</span>\n        <span class=\"token punctuation\">)</span></code></pre></div>\n<p>Batch the inserts. The embedding API calls dominate ingest latency, not the database writes. So the pattern that matters is: embed a batch of chunks in one API call, then <code class=\"language-text\">executemany</code> them in one round trip.</p>\n<h2>Querying with the distance operators</h2>\n<p>Retrieval is an <code class=\"language-text\">ORDER BY ... LIMIT k</code>. pgvector exposes distance as operators, and the operator you pick has to match how your embeddings were trained:</p>\n<ul>\n<li><code class=\"language-text\">&lt;=&gt;</code> cosine distance</li>\n<li><code class=\"language-text\">&lt;-&gt;</code> L2 (Euclidean) distance</li>\n<li><code class=\"language-text\">&lt;#&gt;</code> negative inner product</li>\n<li><code class=\"language-text\">&lt;+&gt;</code> L1 (taxicab) distance</li>\n</ul>\n<p>Most embedding models, including OpenAI’s and the common open-weight ones, produce vectors normalized to unit length and are meant to be compared by cosine similarity. So <code class=\"language-text\">&lt;=&gt;</code> is the usual answer.</p>\n<p>One detail trips people up: pgvector returns cosine <em>distance</em>, which is <code class=\"language-text\">1 - cosine_similarity</code>. Smaller means closer, so you sort ascending. If you want a similarity score to threshold on, compute <code class=\"language-text\">1 - (embedding &lt;=&gt; $1)</code> yourself, as the search function below does:</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\"><span class=\"token keyword\">async</span> <span class=\"token keyword\">def</span> <span class=\"token function\">search</span><span class=\"token punctuation\">(</span>pool<span class=\"token punctuation\">,</span> query_vec<span class=\"token punctuation\">:</span> <span class=\"token builtin\">list</span><span class=\"token punctuation\">[</span><span class=\"token builtin\">float</span><span class=\"token punctuation\">]</span><span class=\"token punctuation\">,</span> k<span class=\"token punctuation\">:</span> <span class=\"token builtin\">int</span> <span class=\"token operator\">=</span> <span class=\"token number\">5</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n    <span class=\"token keyword\">async</span> <span class=\"token keyword\">with</span> pool<span class=\"token punctuation\">.</span>acquire<span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span> <span class=\"token keyword\">as</span> conn<span class=\"token punctuation\">:</span>\n        <span class=\"token keyword\">await</span> register_vector<span class=\"token punctuation\">(</span>conn<span class=\"token punctuation\">)</span>\n        rows <span class=\"token operator\">=</span> <span class=\"token keyword\">await</span> conn<span class=\"token punctuation\">.</span>fetch<span class=\"token punctuation\">(</span>\n            <span class=\"token triple-quoted-string string\">\"\"\"\n            SELECT document, content,\n                   1 - (embedding &lt;=> $1) AS similarity\n            FROM chunks\n            ORDER BY embedding &lt;=> $1\n            LIMIT $2\n            \"\"\"</span><span class=\"token punctuation\">,</span>\n            query_vec<span class=\"token punctuation\">,</span> k<span class=\"token punctuation\">,</span>\n        <span class=\"token punctuation\">)</span>\n        <span class=\"token keyword\">return</span> <span class=\"token punctuation\">[</span><span class=\"token builtin\">dict</span><span class=\"token punctuation\">(</span>r<span class=\"token punctuation\">)</span> <span class=\"token keyword\">for</span> r <span class=\"token keyword\">in</span> rows<span class=\"token punctuation\">]</span></code></pre></div>\n<p>You pass the query vector once and reference it as <code class=\"language-text\">$1</code> in both the <code class=\"language-text\">SELECT</code> and the <code class=\"language-text\">ORDER BY</code>. The planner is smart enough not to compute the distance twice.</p>\n<h2>Indexing: without one, every query is a full scan</h2>\n<p>The query above works on day one. By the time you have a few hundred thousand rows, it has quietly become your bottleneck. With no index, Postgres computes the distance to <em>every</em> row on every query. That is an exact, linear scan: correct, and slow.</p>\n<p>pgvector offers two approximate indexes. Both skip most of the rows, trading a small amount of recall (the share of true nearest neighbours you get back) for a large speedup.</p>\n<h3>HNSW, the default</h3>\n<p><strong>HNSW</strong> (Hierarchical Navigable Small World) builds a multi-layer graph that the search navigates greedily toward the nearest neighbours. It’s the one I default to: it gives better recall at a given speed, and it doesn’t need training data to exist first. The costs are a slower build and more memory. This is the same <a href=\"/blog/2026-07-09-hnsw-vector-search-explained/\">HNSW algorithm I broke down in an earlier post</a>, now as a Postgres index.</p>\n<div class=\"gatsby-highlight\" data-language=\"sql\"><pre class=\"language-sql\"><code class=\"language-sql\"><span class=\"token keyword\">CREATE</span> <span class=\"token keyword\">INDEX</span> <span class=\"token keyword\">ON</span> chunks\n<span class=\"token keyword\">USING</span> hnsw <span class=\"token punctuation\">(</span>embedding vector_cosine_ops<span class=\"token punctuation\">)</span>\n<span class=\"token keyword\">WITH</span> <span class=\"token punctuation\">(</span>m <span class=\"token operator\">=</span> <span class=\"token number\">16</span><span class=\"token punctuation\">,</span> ef_construction <span class=\"token operator\">=</span> <span class=\"token number\">64</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">;</span></code></pre></div>\n<p>Two things to get right in that statement:</p>\n<ul>\n<li><strong>The operator class must match your query operator.</strong> <code class=\"language-text\">vector_cosine_ops</code> goes with <code class=\"language-text\">&lt;=&gt;</code>, <code class=\"language-text\">vector_l2_ops</code> with <code class=\"language-text\">&lt;-&gt;</code>, and <code class=\"language-text\">vector_ip_ops</code> with <code class=\"language-text\">&lt;#&gt;</code>. Mismatch them and the planner quietly ignores the index and falls back to a full scan.</li>\n<li><strong>The build parameters.</strong> <code class=\"language-text\">m</code> (16 by default) is the number of connections per node. <code class=\"language-text\">ef_construction</code> (64 by default) is how hard the builder searches while inserting. Higher values mean better recall and a slower, heavier build.</li>\n</ul>\n<h3>IVFFlat, and its sharp edge</h3>\n<p><strong>IVFFlat</strong> partitions vectors into <code class=\"language-text\">lists</code> clusters and searches only the nearest few. It builds faster and uses less memory, but its recall is more sensitive to tuning.</p>\n<p>Here is the sharp edge: IVFFlat learns its clusters from the rows present at build time. So you have to build it <em>after</em> you have a representative sample of data. Load data first, then index.</p>\n<h3>Tuning recall at query time</h3>\n<p>At query time, HNSW has a knob that controls the recall/latency tradeoff for the session:</p>\n<div class=\"gatsby-highlight\" data-language=\"sql\"><pre class=\"language-sql\"><code class=\"language-sql\"><span class=\"token keyword\">SET</span> hnsw<span class=\"token punctuation\">.</span>ef_search <span class=\"token operator\">=</span> <span class=\"token number\">100</span><span class=\"token punctuation\">;</span>  <span class=\"token comment\">-- default is 40; higher = better recall, slower</span></code></pre></div>\n<p>Raise it when you’re missing relevant chunks. Lower it when p99 latency matters more than the last few points of recall. Measure this against a labelled set rather than eyeballing it. I wrote about <a href=\"/blog/2026-07-04-measuring-rag-retrieval-quality/\">how to measure retrieval quality</a> precisely because “it feels better” is not a number you can regress against.</p>\n<h2>Metadata filtering, and the trap inside it</h2>\n<p>Real retrieval is rarely “search all chunks.” You want <em>this user’s</em> documents, or only pages from the current release, or entries after some date. That’s a <code class=\"language-text\">WHERE</code> clause on the <code class=\"language-text\">jsonb</code> column:</p>\n<div class=\"gatsby-highlight\" data-language=\"sql\"><pre class=\"language-sql\"><code class=\"language-sql\"><span class=\"token keyword\">SELECT</span> document<span class=\"token punctuation\">,</span> content\n<span class=\"token keyword\">FROM</span> chunks\n<span class=\"token keyword\">WHERE</span> metadata<span class=\"token operator\">-</span><span class=\"token operator\">>></span><span class=\"token string\">'project'</span> <span class=\"token operator\">=</span> <span class=\"token string\">'wmcore'</span>\n<span class=\"token keyword\">ORDER</span> <span class=\"token keyword\">BY</span> embedding <span class=\"token operator\">&lt;=></span> $<span class=\"token number\">1</span>\n<span class=\"token keyword\">LIMIT</span> <span class=\"token number\">5</span><span class=\"token punctuation\">;</span></code></pre></div>\n<p>Here’s the part that surprises people: the HNSW index knows nothing about your <code class=\"language-text\">WHERE</code> clause. Postgres walks the vector index in nearest-first order and <em>then</em> discards rows that fail the filter.</p>\n<p>Suppose your filter is selective and keeps only 2% of rows. The index can hand back its whole candidate list before it finds five rows that pass. You then quietly get fewer than five results, or a slow query as the scan reaches deeper.</p>\n<p>pgvector’s answer is <em>iterative index scans</em> (added in 0.8.0). They let the scan go back and pull more candidates when the filter is strict:</p>\n<div class=\"gatsby-highlight\" data-language=\"sql\"><pre class=\"language-sql\"><code class=\"language-sql\"><span class=\"token keyword\">SET</span> hnsw<span class=\"token punctuation\">.</span>iterative_scan <span class=\"token operator\">=</span> strict_order<span class=\"token punctuation\">;</span></code></pre></div>\n<p>Even so, if you always filter on the same high-cardinality field, partial indexes or partitioning by that field will beat a single global index. Push the selective, structured constraint into something Postgres can index conventionally, and let the vector index do the semantic part on the smaller set.</p>\n<p>Dedicated vector databases wrap this same tradeoff in a “metadata filter” feature. In Postgres you see the mechanism directly, which is either a burden or a gift depending on your mood that day.</p>\n<h2>Hybrid search comes almost for free</h2>\n<p>Vector search misses exact matches. Ask for error code <code class=\"language-text\">HTTP 431</code> or a specific dataset name, and semantic similarity will happily return chunks that are <em>about</em> errors without containing the token you need.</p>\n<p>Postgres already ships full-text search. So you can run keyword and vector retrieval in the same database and fuse the rankings, with no second service or query engine. The keyword half is a standard full-text query:</p>\n<div class=\"gatsby-highlight\" data-language=\"sql\"><pre class=\"language-sql\"><code class=\"language-sql\"><span class=\"token keyword\">SELECT</span> id<span class=\"token punctuation\">,</span> content\n<span class=\"token keyword\">FROM</span> chunks\n<span class=\"token keyword\">WHERE</span> to_tsvector<span class=\"token punctuation\">(</span><span class=\"token string\">'english'</span><span class=\"token punctuation\">,</span> content<span class=\"token punctuation\">)</span> @@ plainto_tsquery<span class=\"token punctuation\">(</span><span class=\"token string\">'english'</span><span class=\"token punctuation\">,</span> $<span class=\"token number\">1</span><span class=\"token punctuation\">)</span>\n<span class=\"token keyword\">LIMIT</span> <span class=\"token number\">20</span><span class=\"token punctuation\">;</span></code></pre></div>\n<p>Combining that lexical result with the vector result, usually with Reciprocal Rank Fusion, is the <a href=\"/blog/2026-07-02-hybrid-search-rag-bm25-vectors/\">hybrid search pattern I covered in its own post</a>. The point here is that pgvector doesn’t force you to choose: both retrievers read the same table.</p>\n<h2>Failure modes I’ve actually hit</h2>\n<ul>\n<li><strong>The index that isn’t used.</strong> The operator class doesn’t match the query operator, or you wrapped the column in a function, and the planner drops to a sequential scan. Run <code class=\"language-text\">EXPLAIN ANALYZE</code> and look for <code class=\"language-text\">Index Scan using ..._hnsw</code>. If you see <code class=\"language-text\">Seq Scan</code>, the index is decorative.</li>\n<li><strong>Dimension drift.</strong> You changed embedding models, and the new vectors are 3072-dim against a <code class=\"language-text\">vector(1536)</code> column. The insert errors. The subtler version is comparing vectors from two different models: the spaces don’t align, so you get confidently wrong neighbours.</li>\n<li><strong>Distance sign confusion.</strong> <code class=\"language-text\">&lt;#&gt;</code> returns the <em>negative</em> inner product so that “smaller is closer” holds for the index. If you treat it as a raw similarity, you’ll rank everything backwards.</li>\n<li><strong>Memory during build.</strong> HNSW builds hold the graph in <code class=\"language-text\">maintenance_work_mem</code>. On a large table with the default setting, the build spills and crawls. Raise <code class=\"language-text\">maintenance_work_mem</code> for the session before <code class=\"language-text\">CREATE INDEX</code>, then set it back.</li>\n<li><strong>Recall you never measured.</strong> Approximate indexes are <em>approximate</em>, and the full scan is your ground truth. Sample a few hundred queries, compare the index results against the exact ones, and know your recall number before you tune <code class=\"language-text\">ef_search</code> in the dark.</li>\n</ul>\n<h2>What I’d do differently, and where the line is</h2>\n<p>For a first RAG system, I’d start in Postgres every time. I’d move to a dedicated vector store only when a concrete number forces it:</p>\n<ul>\n<li>index builds that take longer than your ingest window;</li>\n<li>memory that won’t fit the box;</li>\n<li>query volume that needs horizontal sharding a single Postgres node can’t give you.</li>\n</ul>\n<p>Those are real limits, and pgvector doesn’t pretend otherwise. But most projects reach “good enough retrieval” long before they hit those walls. Every week spent operating a second datastore is a week not spent improving chunking, <a href=\"/blog/2026-07-08-cross-encoder-reranking-rag/\">reranking</a>, and evaluation, which is where retrieval quality actually comes from.</p>\n<p>The one thing I’d do earlier than I used to: raise <code class=\"language-text\">maintenance_work_mem</code>, and set <code class=\"language-text\">hnsw.ef_search</code> deliberately from a measured recall target, not by feel. Both are one line, and together they are the difference between “the demo works” and “retrieval holds up under a real query load.”</p>\n<p>For the projects behind this post (the retrieval layer in <a href=\"/project/archi/\">Archi</a> and the document backends in <a href=\"/project/cloud-canvas-ai/\">CloudCanvasAI</a>), keeping vectors in Postgres has meant one system to reason about, back up, and transact against. The best infrastructure decision is often the one that leaves you fewer moving parts to break.</p>\n<h2>References</h2>\n<ul>\n<li><a href=\"https://github.com/pgvector/pgvector\">pgvector</a>: the extension, README, and changelog</li>\n<li><a href=\"https://github.com/pgvector/pgvector-python\">pgvector-python</a>: bindings for psycopg, asyncpg, SQLAlchemy, and more</li>\n<li><a href=\"https://www.postgresql.org/docs/current/textsearch.html\">PostgreSQL full-text search</a>: the lexical half of hybrid retrieval</li>\n<li>Malkov &#x26; Yashunin, <a href=\"https://arxiv.org/abs/1603.09320\"><em>Efficient and robust approximate nearest neighbor search using HNSW graphs</em></a>: the algorithm behind the index</li>\n</ul>","frontmatter":{"title":"RAG Vector Search in Postgres with pgvector","date":"2026-08-03T00:00:00.000Z","description":"Store and query RAG embeddings inside Postgres with pgvector: HNSW indexing, distance operators, metadata filtering, hybrid search, and the tradeoffs I hit.","thumbnail":{"childImageSharp":{"fluid":{"base64":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAABQAAAALCAIAAADwazoUAAAACXBIWXMAAAsSAAALEgHS3X78AAAB4klEQVQoz02QaW+bQBCG+dAkro25Dy/HLhgMC+wB2AkGx457qZH6oe3//zMdnKaq9AjNDPvwzqI0JGuSrMZbGpGnWrzsh+thOHeH6+Nwkv2lf3xu99B+Gaa+qKs4nc+/o9xp0b0W6z7VUeURkbFTIS9vT0IH2r5AvROXqvvkJq3qlcAHLb7T8IOOFWtTOEFpeJnubtcWWRrxysCqOReqhW9tvDJhEhs2Md10xiFQqyZWPNxFxQnTsxW2m+0AtYv3gEcOTtyHu4nQ8yY7cgqbX8ZunPrTUY51xpZ6pMDCdixRNhioceM2KiaUHQE/PZiI2ZF0cWtFkmb8s2ivXJ4ZqDLBdDXLXmEF3E8frZDboQhmcwDgQ7pfmUFjBcxArN4214Z95fx3L8aKJ3G50kNILu1IgAxHnUjcYgdoXdyZqIHwWQ447Pldsitjr5I/1zwFWUOQXFohOMOcHMkwn4J8BPz0CS5iweYhB5nl7JsQl4ZD+LEWOCqXs+zu7IDNF5tl7hP4Vf0NSK5MVFuo1jxKk+KHqH611U9ZvfKqwPlC3SgOKsO0RYQ7AY3zPkxlnHUIN4gIHzPNzXU3X6wRsv2xJMcyGQoylWQXbO5VX1moHvBx7f8r3liofyf/4S/e3z7cJn8A4m1V7m1z4lwAAAAASUVORK5CYII=","aspectRatio":1.899441340782123,"src":"/static/74e573a27101a4894857b92b08201f6b/40a76/hero.png","srcSet":"/static/74e573a27101a4894857b92b08201f6b/c972b/hero.png 340w,\n/static/74e573a27101a4894857b92b08201f6b/27625/hero.png 680w,\n/static/74e573a27101a4894857b92b08201f6b/40a76/hero.png 1360w,\n/static/74e573a27101a4894857b92b08201f6b/ed396/hero.png 2000w","sizes":"(max-width: 1360px) 100vw, 1360px"}}}}}},"pageContext":{"slug":"/2026-08-03-pgvector-rag-postgres/","previous":"blog/2026-07-30-managing-context-window-llm-agents/","next":"blog/2026-08-02-contextual-retrieval-for-rag/"}},"staticQueryHashes":["32046230"]}