OpenSearch as a Vector Store for RAG
Run RAG retrieval inside the OpenSearch cluster you already operate. Set up a knn_vector index, pick faiss vs lucene, and avoid the memory traps.
Most RAG (retrieval-augmented generation) tutorials open a new account somewhere the moment they need vector search. They spin up Pinecone, stand up a Qdrant pod, or add pgvector to a Postgres you now have to babysit. If you are starting from zero, that is fine. But many teams are not starting from zero. They already run a search cluster for logs and full-text search, and now someone wants a retrieval copilot on top of the same documents.
That was the situation on Archi, the retrieval copilot I worked on for CMS computing operations at CERN. The logbooks, JIRA tickets, and wiki pages we wanted to retrieve over were already flowing into OpenSearch, because that is where the operations telemetry lives. A second datastore just for embeddings would have been one more thing to secure through CERN SSO, back up, monitor, and page someone about at 2am. OpenSearch has had a k-NN (k-nearest-neighbor) plugin for years, so the obvious move was to keep the vectors in the cluster we already ran.
This post is for engineers who operate OpenSearch (or Elasticsearch, which has a close equivalent) and are adding vector retrieval, rather than choosing a vector database from a blank slate. It covers:
- how to turn a normal index into a vector index;
- how to choose between the engines OpenSearch ships with;
- how a k-NN query reads;
- the operational edges that bit us.
The memory model matters most. It is the one thing that never shows up in a tutorial and always shows up in an on-call rotation.
Why keep vectors in the cluster you already run
A dedicated vector database makes sense when you have a billion vectors, need aggressive quantization (compressing vectors to save memory), or want a managed service to own the problem. Most RAG systems are nowhere near that. Archi indexes on the order of hundreds of thousands of chunks. At that size, the deciding factor is operations, not raw vector throughput.
Keeping retrieval in OpenSearch buys three concrete things:
- One security surface. The embeddings inherit the authentication, snapshots, and index lifecycle you already configured.
- Hybrid search in one query. Both halves of retrieval live in one index, so a single query can run BM25 keyword search and vector search together and fuse the scores. That is the whole point of hybrid search.
- Existing monitoring. The dashboards you built for monitoring already watch the cluster the retrieval runs on, so a slow query shows up next to everything else.
The tradeoff is that you now run vector search on infrastructure that was sized for logs. That is manageable, but only if you understand where the vector index spends memory. Most of this post builds toward that.
Turn an index into a vector index
Two things make an index a vector index: the setting index.knn: true, and at least one field of type knn_vector. The field declares its dimension and how the approximate-nearest-neighbor (ANN) index is built. The mapping below creates an index with a text field, a keyword field for filtering, and a 1024-dimensional vector field:
PUT /archi-chunks
{
"settings": {
"index.knn": true
},
"mappings": {
"properties": {
"text": { "type": "text" },
"site": { "type": "keyword" },
"embedding": {
"type": "knn_vector",
"dimension": 1024,
"method": {
"name": "hnsw",
"engine": "faiss",
"space_type": "cosinesimil",
"parameters": { "ef_construction": 128, "m": 16 }
}
}
}
}
}A few of these values matter more than they look.
dimension must exactly match your embedding model. If the model outputs 1024-dimensional vectors and the field says 768, indexing fails; no silent truncation rescues you. This sounds trivial until you swap embedding models and forget the field is frozen. You cannot change dimension on a live field, so a model change means a new index and a reindex.
space_type is the distance metric, and it must agree with how your model was trained. The common ones are cosine similarity (cosinesimil), inner product (innerproduct), and Euclidean (l2) (OpenSearch supported field types). Most sentence-embedding models are tuned for cosine. Pick the wrong metric and search still runs; it just quietly returns worse neighbors. That is the hardest kind of bug to notice, because nothing errors.
ef_construction and m are the HNSW build parameters. HNSW is the graph structure the ANN search walks.
mis the number of graph links per node. The docs recommend a range of roughly 8 to 64.ef_constructioncontrols how hard the builder searches while inserting each vector.
Higher values raise recall and cost more RAM and build time. I wrote up what these knobs do to the graph in HNSW, explained. The short version: m between 16 and 32 covers most RAG workloads, and it is rarely the first thing you need to touch.
Pick the engine before you pick the parameters
engine is the k-NN library that builds and searches the graph, and it is the choice with the largest operational consequences. OpenSearch ships three: faiss, lucene, and nmslib. As of recent versions, nmslib is deprecated in favor of faiss and lucene (methods and engines). So the real decision is faiss versus lucene.
The difference that matters on an operations rotation is where the graph lives in memory.
Faiss keeps the graph off-heap. Faiss (Meta’s similarity-search library) builds a native index that OpenSearch loads into off-heap memory, meaning memory outside the JVM heap. A dedicated k-NN circuit breaker manages that memory. Faiss has the most features, including product quantization for compressing vectors, and it behaves better at larger scale. The cost is a second memory budget to reason about. The graph needs RAM that the JVM heap does not account for. If you size the heap as if it were the whole story, the node runs out of memory in a way that is confusing the first time you see it.
Lucene keeps the graph inside the segment. The Lucene engine stores the graph inside the Lucene segment itself. There is no separate native library and no off-heap budget to size, and it plays cleanly with filtering. It is the simpler thing to operate. For a cluster that mostly does full-text search and logs, with a modest vector field bolted on, that simplicity is worth a lot.
Archi started on faiss because we wanted the option of quantization later. But for a team adding its first vector field, I would default to Lucene and move to faiss only when a real need shows up.
How a k-NN query reads
An approximate k-NN query is a knn clause that names the field, the query vector, and k, the number of neighbors to retrieve (approximate k-NN search). This one asks for the five nearest chunks:
GET /archi-chunks/_search
{
"size": 5,
"query": {
"knn": {
"embedding": {
"vector": [0.12, -0.04, 0.88, "...1024 floats..."],
"k": 5
}
}
}
}You embed the user’s question once, on the application side, and pass the resulting array as vector. OpenSearch walks the HNSW graph and returns the closest chunks by the field’s space_type, scored so that a closer vector sorts higher.
Two subtleties trip people up:
kandsizeare different knobs.k, inside theknnclause, is how many candidates the graph search collects per segment. The top-levelsizeis how many hits you get back. When they disagree, you can get fewer results than you expected.- You can also search exactly. A
knn_vectorfield can be searched without the graph, using a scoring script that brute-forces the distance. That is slow on a large index but exact. It is the right tool for small indexes, or for checking how much recall your approximate search gives up.
Filtering deserves its own mention. If you need “only chunks from this site, from this year,” you want the filter applied during the graph traversal, not after it; otherwise recall quietly collapses. OpenSearch treats this as a first-class part of the k-NN query. I went through why post-filtering hurts, and how the engines differ, in metadata filtering for vector search.
The failure modes that actually happened
None of the following showed up in a tutorial. All of them showed up in practice.
Off-heap memory is the real capacity limit. With faiss, the graphs live outside the JVM heap, under the k-NN circuit breaker. When the vectors plus their graph overhead exceed the budget, the breaker trips and k-NN searches start failing. The node survives, but retrieval is down. On a cluster shared with logs, it is easy to size the heap carefully and forget native memory entirely. The fix is boring: compute vector count × dimension × bytes per value, add the HNSW overhead, and make sure the circuit breaker limit and the machine’s RAM leave room for it next to everything else the cluster does.
Every segment is its own graph. OpenSearch builds a separate HNSW graph per Lucene segment, so a query fans out across all of them and merges the results. Many small segments mean many small graphs and more work per query. A force merge down to a few segments, on an index you are done writing to, reduces that overhead and improves both latency and recall. It is the same operational lever that matters for full-text indexes, now paying off twice.
Method parameters are frozen at field creation. You cannot raise m or switch engine on a live field; changing either means a new index and a reindex. That makes the initial choice worth a little thought. It also makes zero-downtime reindexing a skill you will use, because eventually you will want to change one of them.
A wrong space_type fails silently. I am saying it twice because it cost real debugging time. If the metric does not match how the model was trained, nothing errors; results just get worse. Write down the metric your embedding model expects next to the mapping, and test retrieval quality on a handful of known-good questions before you trust it.
When I would move to a dedicated vector database
Running vectors in OpenSearch has been the right call for Archi. It still has edges, and it is worth being honest about where it stops being the obvious choice.
I would reach for a dedicated vector database in two cases:
- the vector count climbs into the tens or hundreds of millions, and quantization becomes mandatory rather than optional;
- the vector workload is heavy enough to deserve its own nodes, sized for it, rather than sharing with logs.
Below that line, which is where many internal RAG tools actually sit, the operational savings of one cluster outweigh the extra features. Hybrid search in a single query is a genuine advantage, and bolting on a second store gives it up.
If you already run OpenSearch, the honest first experiment is not “which vector database should we adopt.” It is turning index.knn on, adding one knn_vector field, and seeing how far the cluster you already operate carries you. For Archi, running on the same OpenSearch that backs CMS workflow operations, it carried us most of the way. And on the day retrieval gets slow, it shows up on the same dashboard as everything else.