HyperLogLog: Count Distinct at Scale in 12KB
Counting unique items exactly costs memory that grows with your data. HyperLogLog estimates the cardinality of billions in about 12KB. Here's how it works.
Counting unique items exactly costs memory that grows with your data. HyperLogLog estimates the cardinality of billions in about 12KB. Here's how it works.
Roaring bitmaps store integer sets in three container types so filters stay small and fast. How they work, why OpenSearch leans on them, and the tradeoffs.
One async request scatters a dozen log lines through concurrent traffic. Thread a correlation ID through FastAPI with contextvars and trace it in OpenSearch.
Float32 embeddings make a vector index expensive to keep in RAM. Binary quantization cuts them 32x; Hamming search plus rescoring keeps recall high.
Run RAG retrieval inside the OpenSearch cluster you already operate. Set up a knn_vector index, pick faiss vs lucene, and avoid the memory traps.
Cosine, dot product, and Euclidean rank vector search results the same on normalized embeddings, differently otherwise. How to pick and configure one.
Log indexes grow until a data node runs out of disk. How OpenSearch ISM rolls indexes over by size, ages them through hot and warm, then deletes on schedule.
Reciprocal Rank Fusion merges two ranked lists without tuning score weights. The RRF formula, why the k constant matters, and how to run it in OpenSearch.
Post-filtering vector search silently drops good results. Here is pre-filter vs filtered ANN for RAG, why filtered HNSW is hard, and how to choose.
Change an OpenSearch mapping without dropping writes or serving stale data. A step-by-step reindex with aliases, the Reindex API, and the failure modes.
Vector search alone misses exact IDs and error codes. Here's how to combine BM25 keyword search with dense retrieval, fuse the rankings with RRF, and rerank.
Build an OpenSearch dashboard that watches a job pipeline: structured logs, an index mapping, the queries behind each panel, and alerts that fire early.