#RAG(30)

September 2026
#AI #RAG #OpenSearch #Search #Python

Reciprocal Rank Fusion for Hybrid Search

Reciprocal Rank Fusion merges two ranked lists without tuning score weights. The RRF formula, why the k constant matters, and how to run it in OpenSearch.

Read more →
August 2026
#AI #LLM #RAG #Embeddings #Search

Contextual Retrieval for RAG, Step by Step

Vector search misses chunks that lost their document context. Contextual Retrieval prepends a short LLM-written summary to each chunk before you embed it.

Read more →
August 2026
#AI #LLM #RAG #Retrieval #Python

HyDE: Hypothetical Document Embeddings for RAG

Vector search fails when a short question looks nothing like its answer. HyDE has an LLM draft a fake answer, embeds that, and retrieves against it instead.

Read more →
August 2026
#AI #LLM #RAG #Security #Prompt Injection

Prompt Injection Defense for RAG Systems

Retrieved documents are untrusted input. A practical guide to defending a RAG copilot against direct and indirect prompt injection, and why filters alone fail.

Read more →
August 2026
#AI #LLM #RAG #Evaluation #Python

LLM-as-a-Judge: Scoring RAG Answer Quality

Retrieval metrics say the right docs came back, not that the answer is right. Build an LLM-as-a-judge to score RAG answers for faithfulness and quality.

Read more →
August 2026
#Python #AsyncIO #FastAPI #LLM #RAG

Fan Out Concurrent LLM Calls with asyncio.gather

Awaiting retrieval and LLM calls one by one wastes seconds per request. Here's how to fan them out with asyncio.gather, bound it, and handle partial failures.

Read more →
August 2026
#AI #RAG #VectorSearch #Embeddings #FAISS

Product Quantization for Vector Search

HNSW makes vector search fast, but the embeddings still fill your RAM. Product quantization compresses them ~32x with a small recall hit. Here is how it works.

Read more →
August 2026
#AI #LLM #RAG #FastAPI #Python

Build a Semantic Cache for LLM Apps

Semantic caching for LLM apps: cache answers by embedding similarity in FastAPI, tune the cutoff, and avoid false cache hits. Working code and failure modes.

Read more →
August 2026
#AI #RAG #Embeddings #Vector Search #LLM

Matryoshka Embeddings: Smaller Vectors for RAG

Matryoshka embeddings front-load meaning into a vector's first dimensions, so you can truncate them for cheaper, faster RAG search without losing much recall.

Read more →
August 2026
#AI #LLM #RAG #VectorSearch #Python

Late Interaction Retrieval for RAG with ColBERT

A bi-encoder averages token detail away; a cross-encoder is too slow to rank a corpus. Late interaction with ColBERT sits between them. Here is how it works.

Read more →
August 2026
#AI #RAG #pgvector #Postgres #FastAPI

RAG Vector Search in Postgres with pgvector

Store and query RAG embeddings inside Postgres with pgvector: HNSW indexing, distance operators, metadata filtering, hybrid search, and the tradeoffs I hit.

Read more →
August 2026
#AI #LLM #RAG #Retrieval #Search

Contextual Retrieval for RAG Pipelines

Chunking a document for RAG strips each piece of its context. Contextual retrieval adds an LLM-written note to every chunk before you index it.

Read more →
August 2026
#AI #LLM #RAG #VectorSearch #Python

Metadata Filtering in Vector Search for RAG

Adding a metadata filter to a vector search can silently return fewer results or wreck recall. How post-filter, pre-filter, and filterable HNSW actually differ.

Read more →
July 2026
#AI #LLM #Agents #Python #RAG

Managing the Context Window in Long Agent Runs

An LLM agent that runs long enough fills its context window and starts to slow or fail. How to prune, compact, and offload context so agents keep going.

Read more →
July 2026
#AI #LLM #Sampling #RAG

LLM Sampling: Temperature, Top-p, and Top-k

Temperature, top-p, and top-k are the three knobs that shape how an LLM picks each token. How each one works, when to reach for it, and how they interact.

Read more →
July 2026
#AI #LLM #RAG #Embeddings #Search

How to Choose an Embedding Model for RAG

The embedding model sets the ceiling on RAG retrieval quality. How to choose one by task fit, sequence length, dimensions, and domain, plus the silent bugs.

Read more →
July 2026
#AI #LLM #RAG #Python #FastAPI

Query Rewriting for Better RAG Retrieval

Short, vague, follow-up questions don't match how your docs are written. How query rewriting, multi-query expansion, and HyDE fix retrieval before it runs.

Read more →
July 2026
#AI #LLM #Evaluation #RAG #Python

LLM-as-a-Judge: Evaluating LLM Output Quality

Human review does not scale for grading LLM answers. How to use an LLM as a judge: write a rubric, score with structured output, and control the biases.

Read more →
July 2026
#AI #LLM #RAG #Security #Agents

Indirect Prompt Injection in RAG Systems

A RAG copilot reads tickets and logs, so whoever writes them can plant instructions in the prompt. How indirect prompt injection works and how to contain it.

Read more →
July 2026
#AI #LLM #RAG #VectorSearch #Python

How HNSW Vector Search Actually Works

Every RAG stack leans on HNSW but treats it as a black box. Here is how the layered graph index finds nearest neighbors fast, and the knobs that matter.

Read more →
July 2026
#AI #LLM #RAG #Python #FastAPI

Cross-Encoder Reranking for RAG Pipelines

Retrieval puts the right chunk at rank 8, but the generator only reads the top few. How a cross-encoder reranker reorders RAG candidates, and where it fails.

Read more →
July 2026
#AI #LLM #RAG #Python #FastAPI

Chunking Strategies for RAG Pipelines

How to chunk documents for a RAG pipeline: why fixed-size splitting fails, structure-aware splitting, size and overlap tradeoffs, and the failure modes.

Read more →
July 2026
#AI #LLM #RAG #Evaluation #Python

Measuring RAG Retrieval Quality: recall@k, MRR

Changed your embeddings or added a reranker? Measure it. Build a golden set and score retrieval with recall@k, MRR, and nDCG before you trust the change.

Read more →
July 2026
#AI #LLM #RAG #OpenSearch #Python

Hybrid Search for RAG: BM25 + Vectors

Vector search alone misses exact IDs and error codes. Here's how to combine BM25 keyword search with dense retrieval, fuse the rankings with RRF, and rerank.

Read more →
June 2026
#AI #LLM #RAG #FastAPI #Python

Incremental Indexing for RAG Pipelines

How to keep a RAG index fresh without full rebuilds: detect changed files with checksums, re-embed only what changed, and handle deletions safely.

Read more →