OpenSearch as a Vector Store for RAG

Run RAG retrieval inside the OpenSearch cluster you already operate. Set up a knn_vector index, pick faiss vs lucene, and avoid the memory traps.

Most RAG (retrieval-augmented generation) tutorials open a new account somewhere the moment they need vector search. They spin up Pinecone, stand up a Qdrant pod, or add pgvector to a Postgres you now have to babysit. If you are starting from zero, that is fine. But many teams are not starting from zero. They already run a search cluster for logs and full-text search, and now someone wants a retrieval copilot on top of the same documents.

That was the situation on Archi, the retrieval copilot I worked on for CMS computing operations at CERN. The logbooks, JIRA tickets, and wiki pages we wanted to retrieve over were already flowing into OpenSearch, because that is where the operations telemetry lives. A second datastore just for embeddings would have been one more thing to secure through CERN SSO, back up, monitor, and page someone about at 2am. OpenSearch has had a k-NN (k-nearest-neighbor) plugin for years, so the obvious move was to keep the vectors in the cluster we already ran.

This post is for engineers who operate OpenSearch (or Elasticsearch, which has a close equivalent) and are adding vector retrieval, rather than choosing a vector database from a blank slate. It covers:

  • how to turn a normal index into a vector index;
  • how to choose between the engines OpenSearch ships with;
  • how a k-NN query reads;
  • the operational edges that bit us.

The memory model matters most. It is the one thing that never shows up in a tutorial and always shows up in an on-call rotation.

Why keep vectors in the cluster you already run

A dedicated vector database makes sense when you have a billion vectors, need aggressive quantization (compressing vectors to save memory), or want a managed service to own the problem. Most RAG systems are nowhere near that. Archi indexes on the order of hundreds of thousands of chunks. At that size, the deciding factor is operations, not raw vector throughput.

Keeping retrieval in OpenSearch buys three concrete things:

  • One security surface. The embeddings inherit the authentication, snapshots, and index lifecycle you already configured.
  • Hybrid search in one query. Both halves of retrieval live in one index, so a single query can run BM25 keyword search and vector search together and fuse the scores. That is the whole point of hybrid search.
  • Existing monitoring. The dashboards you built for monitoring already watch the cluster the retrieval runs on, so a slow query shows up next to everything else.

One OpenSearch cluster serving two retrieval paths. On the left, documents such as logs, tickets, and wiki pages are chunked and embedded into a float array per chunk, then indexed. In the middle sits one OpenSearch index with index.knn set to true, holding a text field backed by a BM25 inverted index and an embedding field of type knn_vector backed by an HNSW graph per segment. On the right, a user question is embedded once and sent as a single request combining a knn clause and a match clause; the ranked passages come back and feed an LLM answer that uses the top-k as context. A caption notes there is no second datastore to run, back up, or secure, and that retrieval reuses the auth, snapshots, and dashboards already on the cluster.

The tradeoff is that you now run vector search on infrastructure that was sized for logs. That is manageable, but only if you understand where the vector index spends memory. Most of this post builds toward that.

Turn an index into a vector index

Two things make an index a vector index: the setting index.knn: true, and at least one field of type knn_vector. The field declares its dimension and how the approximate-nearest-neighbor (ANN) index is built. The mapping below creates an index with a text field, a keyword field for filtering, and a 1024-dimensional vector field:

PUT /archi-chunks
{
  "settings": {
    "index.knn": true
  },
  "mappings": {
    "properties": {
      "text": { "type": "text" },
      "site": { "type": "keyword" },
      "embedding": {
        "type": "knn_vector",
        "dimension": 1024,
        "method": {
          "name": "hnsw",
          "engine": "faiss",
          "space_type": "cosinesimil",
          "parameters": { "ef_construction": 128, "m": 16 }
        }
      }
    }
  }
}

A few of these values matter more than they look.

dimension must exactly match your embedding model. If the model outputs 1024-dimensional vectors and the field says 768, indexing fails; no silent truncation rescues you. This sounds trivial until you swap embedding models and forget the field is frozen. You cannot change dimension on a live field, so a model change means a new index and a reindex.

space_type is the distance metric, and it must agree with how your model was trained. The common ones are cosine similarity (cosinesimil), inner product (innerproduct), and Euclidean (l2) (OpenSearch supported field types). Most sentence-embedding models are tuned for cosine. Pick the wrong metric and search still runs; it just quietly returns worse neighbors. That is the hardest kind of bug to notice, because nothing errors.

ef_construction and m are the HNSW build parameters. HNSW is the graph structure the ANN search walks.

  • m is the number of graph links per node. The docs recommend a range of roughly 8 to 64.
  • ef_construction controls how hard the builder searches while inserting each vector.

Higher values raise recall and cost more RAM and build time. I wrote up what these knobs do to the graph in HNSW, explained. The short version: m between 16 and 32 covers most RAG workloads, and it is rarely the first thing you need to touch.

Pick the engine before you pick the parameters

engine is the k-NN library that builds and searches the graph, and it is the choice with the largest operational consequences. OpenSearch ships three: faiss, lucene, and nmslib. As of recent versions, nmslib is deprecated in favor of faiss and lucene (methods and engines). So the real decision is faiss versus lucene.

The difference that matters on an operations rotation is where the graph lives in memory.

Where the vector graph lives on a data node, and how that decides the failure mode. A data node contains a JVM heap on the left, used for query coordination, Lucene, and aggregations; inside it, the lucene engine stores its HNSW graph in the Lucene segment, memory-mapped, with no separate memory budget, the simplest ops and fewest tuning knobs. On the right sits off-heap native memory, governed by the k-NN circuit breaker, where the faiss and nmslib engines load their graphs outside the heap; these offer the most features such as product quantization and larger scale, but when the circuit breaker trips, searches fail rather than the node crashing. A caption warns that sizing the off-heap budget alongside logs and metrics is the trap in a shared cluster, because the graph needs RAM the heap does not see. A footnote repeats that nmslib is deprecated in favor of faiss and lucene.

Faiss keeps the graph off-heap. Faiss (Meta’s similarity-search library) builds a native index that OpenSearch loads into off-heap memory, meaning memory outside the JVM heap. A dedicated k-NN circuit breaker manages that memory. Faiss has the most features, including product quantization for compressing vectors, and it behaves better at larger scale. The cost is a second memory budget to reason about. The graph needs RAM that the JVM heap does not account for. If you size the heap as if it were the whole story, the node runs out of memory in a way that is confusing the first time you see it.

Lucene keeps the graph inside the segment. The Lucene engine stores the graph inside the Lucene segment itself. There is no separate native library and no off-heap budget to size, and it plays cleanly with filtering. It is the simpler thing to operate. For a cluster that mostly does full-text search and logs, with a modest vector field bolted on, that simplicity is worth a lot.

Archi started on faiss because we wanted the option of quantization later. But for a team adding its first vector field, I would default to Lucene and move to faiss only when a real need shows up.

How a k-NN query reads

An approximate k-NN query is a knn clause that names the field, the query vector, and k, the number of neighbors to retrieve (approximate k-NN search). This one asks for the five nearest chunks:

GET /archi-chunks/_search
{
  "size": 5,
  "query": {
    "knn": {
      "embedding": {
        "vector": [0.12, -0.04, 0.88, "...1024 floats..."],
        "k": 5
      }
    }
  }
}

You embed the user’s question once, on the application side, and pass the resulting array as vector. OpenSearch walks the HNSW graph and returns the closest chunks by the field’s space_type, scored so that a closer vector sorts higher.

Two subtleties trip people up:

  • k and size are different knobs. k, inside the knn clause, is how many candidates the graph search collects per segment. The top-level size is how many hits you get back. When they disagree, you can get fewer results than you expected.
  • You can also search exactly. A knn_vector field can be searched without the graph, using a scoring script that brute-forces the distance. That is slow on a large index but exact. It is the right tool for small indexes, or for checking how much recall your approximate search gives up.

Filtering deserves its own mention. If you need “only chunks from this site, from this year,” you want the filter applied during the graph traversal, not after it; otherwise recall quietly collapses. OpenSearch treats this as a first-class part of the k-NN query. I went through why post-filtering hurts, and how the engines differ, in metadata filtering for vector search.

The failure modes that actually happened

None of the following showed up in a tutorial. All of them showed up in practice.

Off-heap memory is the real capacity limit. With faiss, the graphs live outside the JVM heap, under the k-NN circuit breaker. When the vectors plus their graph overhead exceed the budget, the breaker trips and k-NN searches start failing. The node survives, but retrieval is down. On a cluster shared with logs, it is easy to size the heap carefully and forget native memory entirely. The fix is boring: compute vector count × dimension × bytes per value, add the HNSW overhead, and make sure the circuit breaker limit and the machine’s RAM leave room for it next to everything else the cluster does.

Every segment is its own graph. OpenSearch builds a separate HNSW graph per Lucene segment, so a query fans out across all of them and merges the results. Many small segments mean many small graphs and more work per query. A force merge down to a few segments, on an index you are done writing to, reduces that overhead and improves both latency and recall. It is the same operational lever that matters for full-text indexes, now paying off twice.

Method parameters are frozen at field creation. You cannot raise m or switch engine on a live field; changing either means a new index and a reindex. That makes the initial choice worth a little thought. It also makes zero-downtime reindexing a skill you will use, because eventually you will want to change one of them.

A wrong space_type fails silently. I am saying it twice because it cost real debugging time. If the metric does not match how the model was trained, nothing errors; results just get worse. Write down the metric your embedding model expects next to the mapping, and test retrieval quality on a handful of known-good questions before you trust it.

When I would move to a dedicated vector database

Running vectors in OpenSearch has been the right call for Archi. It still has edges, and it is worth being honest about where it stops being the obvious choice.

I would reach for a dedicated vector database in two cases:

  • the vector count climbs into the tens or hundreds of millions, and quantization becomes mandatory rather than optional;
  • the vector workload is heavy enough to deserve its own nodes, sized for it, rather than sharing with logs.

Below that line, which is where many internal RAG tools actually sit, the operational savings of one cluster outweigh the extra features. Hybrid search in a single query is a genuine advantage, and bolting on a second store gives it up.

If you already run OpenSearch, the honest first experiment is not “which vector database should we adopt.” It is turning index.knn on, adding one knn_vector field, and seeing how far the cluster you already operate carries you. For Archi, running on the same OpenSearch that backs CMS workflow operations, it carried us most of the way. And on the day retrieval gets slow, it shows up on the same dashboard as everything else.