Maximal Marginal Relevance for RAG
Top-k vector search keeps returning near-duplicate chunks that fill the context window. How Maximal Marginal Relevance reranks for relevant, diverse results.
Top-k vector search keeps returning near-duplicate chunks that fill the context window. How Maximal Marginal Relevance reranks for relevant, diverse results.
Small chunks search precisely but read like fragments. Parent document retrieval matches on child chunks, then feeds the whole parent to the model.
Vector search fails when a short question looks nothing like its answer. HyDE has an LLM draft a fake answer, embeds that, and retrieves against it instead.
Chunking a document for RAG strips each piece of its context. Contextual retrieval adds an LLM-written note to every chunk before you index it.