Incremental Indexing for RAG Pipelines
How to keep a RAG index fresh without full rebuilds: detect changed files with checksums, re-embed only what changed, and handle deletions safely.
Most RAG (retrieval-augmented generation) tutorials stop at “embed your documents, store the vectors, query them.” That works until the documents change. Then you have two bad options: rebuild the whole index on every edit, or let the answers go stale. Neither is acceptable once real users depend on the system.
This post is for engineers who already have a working retrieval pipeline and need it to stay current. It covers detecting changed files with content hashes, re-embedding only those files, watching a folder, and the failure modes that show up later.
The pattern comes from ARGRAG, a RAG system my teammate and I built for the Argusa AI Challenge. Its corpus was a live folder of enterprise documents that could change while the app was running. The same approach drives the ingestion side of Archi, the LLM copilot I worked on for CMS computing operations at CERN, where logbooks and tickets are updated constantly.
Why full reindexing breaks down
A full reindex re-reads every file, re-chunks it, re-embeds every chunk, and rewrites the vector store. For a small demo, that takes a few seconds. For a corpus of a few thousand mixed PDFs, Office files, and source code, it costs minutes of compute and a load spike every time a single file changes.
Compute is only part of the cost. During a rebuild you either serve stale results or block queries, and re-embedding unchanged content wastes money if you pay per token. The fix is to treat indexing the way build systems treat compilation: only touch what actually changed.
Detect changes with content hashes
Modification timestamps lie. A file can be rewritten with identical content, copied, or restored from backup, and its mtime (last-modified time) jumps even though the bytes did not change. A content hash is the honest signal. Hash the file, compare the result with the last hash you stored, and skip anything that matches.
The function below streams a file through SHA-256 in fixed-size blocks, so large files never have to fit in memory.
import hashlib
from pathlib import Path
def file_checksum(path: Path, chunk_size: int = 65536) -> str:
"""SHA-256 of file contents, streamed so large files don't blow up memory."""
h = hashlib.sha256()
with path.open("rb") as f:
for block in iter(lambda: f.read(chunk_size), b""):
h.update(block)
return h.hexdigest()Store each checksum next to the file’s identity in a relational table. In ARGRAG, the vectors lived in ChromaDB, while the checksums and document metadata lived in PostgreSQL. Keeping the two stores separate matters. “What changed?” is a relational query, and “what is similar?” is a vector query. Forcing both into one store makes one of them awkward.
classify compares the stored checksums with the files on disk and returns three groups: new (never seen), changed (hash differs), and deleted (stored but gone from disk).
def classify(paths, stored):
"""Split the current files into new, changed, and unchanged sets."""
current = {p: file_checksum(p) for p in paths}
new = {p: c for p, c in current.items() if p not in stored}
changed = {p: c for p, c in current.items()
if p in stored and stored[p] != c}
deleted = [p for p in stored if p not in current]
return new, changed, deletedRe-embed only the files that changed
Once you know which files changed, the update is mechanical. For each new or changed file, delete its old chunks from the vector store and insert fresh ones. ChromaDB makes the delete easy if you tag every chunk with its source path in the metadata.
In the function below, notice the source (file path) tag in each chunk’s metadata. It is what lets one delete call remove all of that file’s old chunks.
def reindex_file(collection, embedder, path, checksum):
# Drop any existing vectors for this file before re-inserting.
collection.delete(where={"source": str(path)})
chunks = chunk_document(path) # your splitter of choice
embeddings = embedder.encode([c.text for c in chunks])
collection.add(
ids=[f"{path}:{i}" for i in range(len(chunks))],
embeddings=embeddings.tolist(),
documents=[c.text for c in chunks],
metadatas=[{"source": str(path), "checksum": checksum} for c in chunks],
)The delete-then-add order is deliberate. If you add first, a query that lands between the two calls can return both the old and the new chunks for the same file. That hands the model duplicated or contradictory context. So delete first, accept a brief window where that file is missing, then add.
We used SentenceTransformers for embeddings because it runs locally with no API cost. That kept each per-file re-embed cheap enough to run on every change. ChromaDB handled the vector storage and the metadata-filtered deletes. Read the docs for both before you commit to a chunk-ID scheme, because the ID format is hard to change later.
Watch the folder and debounce events
To make the index update on its own, watch the corpus directory. ARGRAG polled for changes every second, which is simple and predictable. If you prefer event-driven watching, the watchdog library hooks directly into the operating system’s file events.
Whichever you pick, debounce: wait for a burst of events to settle before acting. Editors and sync tools fire several events for one logical save: a write, a rename of a temp file, a permission change. React to each one and you re-embed the same file several times in a row. Instead, collect events for a short window, then run one pass over the settled set.
The sketch below scans the folder every interval seconds and hands a file to run_incremental_index only once it has sat in pending for settle seconds.
import time
def watch(folder, interval=1.0, settle=2.0):
pending = {}
while True:
for p in Path(folder).rglob("*"):
if p.is_file():
pending[p] = time.monotonic()
ready = [p for p, t in pending.items()
if time.monotonic() - t >= settle]
if ready:
run_incremental_index(ready)
for p in ready:
pending.pop(p, None)
time.sleep(interval)Failure modes that bite later
These are easy to miss in the first version and painful to debug in production.
Deleted files, the one people forget. New and changed files announce themselves, but a deleted file simply stops showing up in the scan. If you never remove its vectors, the index keeps returning chunks for documents that no longer exist, and the model cites sources the user cannot open. Handle the deleted set explicitly with collection.delete(where={"source": ...}).
Embedding model upgrades. Stored vectors are only comparable if they all came from the same model. After an upgrade, the content hash still says each file is unchanged, so an incremental pass skips it, and the index quietly mixes vectors from two models that do not share a space. Store the model name and version next to the checksum, and treat a model change as its own reason to re-embed, even when the bytes match.
Chunk-boundary drift. This shows up after you touch the splitter: once you chunk differently, the old chunk IDs no longer line up with the new ones. The per-file delete keyed on source saves you here. Because it deletes by stable per-file metadata instead of exact chunk IDs, a re-chunk replaces that file’s chunks cleanly.
Partial writes. If you index a file while another process is still writing it, you hash and embed half a document. The watcher’s settle window helps. A stricter version checks that the checksum is stable across two reads before it indexes anything.
Tradeoffs: a second source of truth
Incremental indexing trades simplicity for freshness and cost. You now carry a second source of truth, the checksum table, and it must stay consistent with the vector store. If the two drift, a file can end up marked as indexed with no chunks, for example when the process dies after deleting vectors but before recording the new checksum.
A periodic full reconcile, run during quiet hours, is cheap insurance. List every file, compare it against the stored state, and repair any mismatches.
For a static corpus that changes once a month, none of this is worth it. Rebuild on a schedule and move on. The pattern earns its keep when the corpus is live and rebuilds are expensive. That is exactly the case for operational tooling like the dashboards and pipelines I worked on for CMS workflow operations.
What I would do differently
In a hackathon you optimize for working software in two days, so ARGRAG kept the bookkeeping minimal. With more time, I would change two things:
- Version the chunks instead of hard-deleting them, so a bad re-index can roll back.
- Record the embedding model version from day one, instead of bolting it on after the first model upgrade forces a full rebuild.
Both are cheap to add early and annoying to retrofit.
If you want the surrounding context, the full project writeups are on the portfolio: ARGRAG for the RAG system this pattern came from, and Archi for the same ingestion ideas applied to CERN operations data.
Image credit: incremental indexing diagram by M. Hassan Ahmed, created for this post, released under CC0 (public domain).