{"componentChunkName":"component---src-templates-blog-post-js","path":"/blog/2026-09-26-hyperloglog-count-distinct-at-scale/","result":{"data":{"site":{"siteMetadata":{"title":"M.Hassan Ahmed","author":"Hassan11196"}},"markdownRemark":{"id":"5bf2fb34-7429-5537-b5a9-516a5dfe0ae6","excerpt":"How many distinct things are in this stream? Distinct users today, distinct error signatures this hour, distinct workflow IDs that touched a broken storage…","html":"<p>How many <em>distinct</em> things are in this stream? Distinct users today, distinct error signatures this hour, distinct workflow IDs that touched a broken storage endpoint. The question sounds trivial until you try to answer it fast. The exact answer requires remembering everything you have already seen, so that you never count it twice, and “remember everything” is the part that does not scale.</p>\n<p>I hit this on the telemetry side of <a href=\"/project/cms-workflow-operations/\">CMS workflow operations</a> at CERN. The logs from WMCore and its surrounding databases are large enough that a plain <code class=\"language-text\">COUNT(DISTINCT)</code> over a wide time range either times out or quietly eats a node’s heap. HyperLogLog is the algorithm that sidesteps this. It answers “how many unique?” for billions of items using a fixed slab of memory measured in kilobytes, and its summaries can be combined across shards. The catch is that the answer is an estimate.</p>\n<p>This post is for engineers who need distinct counts at a scale where exact counting is no longer an option, and who want to know what they are trading away. It covers why exact counting breaks, how HyperLogLog works, why it merges so well, how to use it in Redis and OpenSearch, and where it misleads you.</p>\n<h2>Why exact counting stops scaling</h2>\n<p>The obvious way to count uniques is a set. Walk the stream, drop every item into a hash set, and the size of the set is your answer. It is exact and simple. Its memory cost is the problem: the set must hold one entry per unique value, so it grows with the number of <em>distinct</em> items.</p>\n<p>Ten million unique 20-byte workflow IDs already make a couple hundred megabytes of set, before overhead. Do that for fifty different metrics across a retention window, and you are budgeting real RAM for a question nobody needs a byte-perfect answer to.</p>\n<p>Distribution makes it worse. When your data lives on many shards, each shard can build its own set. But to get the global distinct count you have to merge those sets, which means shipping every ID to one place and de-duplicating there. You have moved the whole dataset across the network just to count it.</p>\n<p>That is the wall. What you want instead is a small, fixed-size summary of a set that you can merge cheaply. That is exactly what a HyperLogLog sketch is. (A sketch is a compact data structure that summarizes a large dataset and answers approximate queries about it.)</p>\n<h2>The core idea: rare bit patterns hint at large counts</h2>\n<p>The whole family of algorithms rests on a bet about randomness. Run each item through a good hash function, so its output looks like uniform random bits. Then look at how many zero bits the hash starts with:</p>\n<ul>\n<li>About half of random hashes start with a <code class=\"language-text\">1</code>, so no leading zeros.</li>\n<li>A quarter start with <code class=\"language-text\">01</code>.</li>\n<li>An eighth start with <code class=\"language-text\">001</code>.</li>\n</ul>\n<p>A hash that begins with <em>ten</em> leading zeros is a one-in-a-thousand event. If you have seen one, you have probably looked at something on the order of a thousand distinct values.</p>\n<p>That is the entire intuition, and it goes back to <a href=\"https://www.sciencedirect.com/science/article/pii/0022000085900418\">Flajolet and Martin’s 1985 probabilistic counting work</a>. Track the maximum number of leading zeros you have seen across all items, call it <em>ρ</em>, and <code class=\"language-text\">2^ρ</code> is a rough guess at the cardinality (the number of distinct items).</p>\n<p>“Rough” is the key word. A single unlucky hash with a long run of zeros throws the estimate off by a factor of two, because the whole count rests on one extreme observation. HyperLogLog’s real contribution is not the leading-zeros idea. It is how it beats down that variance while storing only a few kilobytes.</p>\n<h2>Many registers, combined with a harmonic mean</h2>\n<p>Instead of one running maximum, HyperLogLog keeps many. It splits each hash into two parts:</p>\n<ul>\n<li>The first <em>p</em> bits select one of <code class=\"language-text\">m = 2^p</code> <strong>registers</strong> (small counters).</li>\n<li>The remaining bits are where it counts leading zeros.</li>\n</ul>\n<p>Each register stores the largest leading-zero count (plus one) that any item routed to it has produced. You are now running <code class=\"language-text\">m</code> independent little experiments in parallel, and each register is one noisy estimator. The diagram walks through a single update.</p>\n<p><img src=\"/c503aa51a02e13fc56f89304fc9eea26/hll-update.svg\" alt=\"How one item updates a HyperLogLog. The item is hashed to a 64-bit value. The first p equals 4 bits, 0110, select register index 6. The remaining bits, starting with three zeros, give a rho of 4. Register 6 is updated to the maximum of its old value 1 and the new value 4, becoming 4. A row of 16 registers is shown holding small integers. A note explains a register holding rho hints it has seen roughly 2 to the rho distinct items, that one register is noisy, and the estimate combines all m registers with a harmonic mean that tames outliers.\"></p>\n<h3>Combining the registers</h3>\n<p>The way HyperLogLog combines the registers is where it earns its name. You might expect an ordinary average, but an arithmetic mean is still wrecked by the occasional huge register. Instead, HyperLogLog takes the <a href=\"https://en.wikipedia.org/wiki/Harmonic_mean\">harmonic mean</a> of <code class=\"language-text\">2</code> raised to each register value, scaled by a correction constant. The harmonic mean is dominated by the <em>small</em> values, so it shrugs off large outliers. The published estimator is:</p>\n<div class=\"gatsby-highlight\" data-language=\"text\"><pre class=\"language-text\"><code class=\"language-text\">E = α_m · m² / Σ_j 2^(−M[j])</code></pre></div>\n<p>Here <code class=\"language-text\">M[j]</code> is the value in register <em>j</em>, and <code class=\"language-text\">α_m</code> is a constant near <code class=\"language-text\">0.7213</code> that corrects the systematic bias of the raw formula. You do not need to memorize the formula.</p>\n<h3>The one tuning knob: error versus memory</h3>\n<p>The property that matters falls out of the formula: with <code class=\"language-text\">m</code> registers, the <a href=\"https://algo.inria.fr/flajolet/Publications/FlFuGaMe07.pdf\">standard error is about <code class=\"language-text\">1.04 / √m</code></a>. Quadruple the registers and you halve the error. That relationship is the entire tuning knob, and it explains the numbers you see in real systems. Here is how it plays out in Redis:</p>\n<ul>\n<li>Redis uses <code class=\"language-text\">m = 16384</code> registers.</li>\n<li><code class=\"language-text\">1.04 / √16384 = 1.04 / 128</code>, which is <code class=\"language-text\">0.0081</code>: the 0.81% error <a href=\"https://redis.io/docs/latest/develop/data-types/probabilistic/hyperloglogs/\">its docs quote</a>.</li>\n<li>Each register only needs to hold a number up to about 64, so six bits is plenty.</li>\n<li><code class=\"language-text\">16384 × 6 bits</code> is <code class=\"language-text\">12288</code> bytes: the famous 12KB.</li>\n</ul>\n<p>That memory is fixed whether you feed the sketch a thousand items or a trillion.</p>\n<h3>The add path</h3>\n<p>Written out, adding an item is unremarkable, which is the point. Hash the item, use the top bits to pick a register, count leading zeros in the rest, and keep the maximum:</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\"><span class=\"token keyword\">def</span> <span class=\"token function\">add</span><span class=\"token punctuation\">(</span>self<span class=\"token punctuation\">,</span> item<span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n    h <span class=\"token operator\">=</span> hash64<span class=\"token punctuation\">(</span>item<span class=\"token punctuation\">)</span>                    <span class=\"token comment\"># a good 64-bit hash</span>\n    j <span class=\"token operator\">=</span> h <span class=\"token operator\">>></span> <span class=\"token punctuation\">(</span><span class=\"token number\">64</span> <span class=\"token operator\">-</span> self<span class=\"token punctuation\">.</span>p<span class=\"token punctuation\">)</span>              <span class=\"token comment\"># first p bits -> register index</span>\n    w <span class=\"token operator\">=</span> <span class=\"token punctuation\">(</span>h <span class=\"token operator\">&lt;&lt;</span> self<span class=\"token punctuation\">.</span>p<span class=\"token punctuation\">)</span> <span class=\"token operator\">&amp;</span> MASK_64        <span class=\"token comment\"># the rest of the bits</span>\n    rho <span class=\"token operator\">=</span> leading_zeros<span class=\"token punctuation\">(</span>w<span class=\"token punctuation\">)</span> <span class=\"token operator\">+</span> <span class=\"token number\">1</span>         <span class=\"token comment\"># position of the leftmost 1</span>\n    self<span class=\"token punctuation\">.</span>registers<span class=\"token punctuation\">[</span>j<span class=\"token punctuation\">]</span> <span class=\"token operator\">=</span> <span class=\"token builtin\">max</span><span class=\"token punctuation\">(</span>self<span class=\"token punctuation\">.</span>registers<span class=\"token punctuation\">[</span>j<span class=\"token punctuation\">]</span><span class=\"token punctuation\">,</span> rho<span class=\"token punctuation\">)</span></code></pre></div>\n<p>Every add is constant time and touches one register. There is no growth, no rehashing, no compaction.</p>\n<h2>Why sketches merge, and why that is the real win</h2>\n<p>The register update is a <code class=\"language-text\">max</code>, and that one detail is what makes HyperLogLog usable in a distributed system. You can merge two sketches built over two different slices of data by taking the larger of the two values at each register position. The result is bit-for-bit the sketch you would have gotten by feeding both slices into one HyperLogLog from the start.</p>\n<p><img src=\"/5178d5518aa208e606764d44d9348296/hll-merge.svg\" alt=\"Merging two sketches is a register-wise max. Shard A&#x27;s registers 2, 0, 5, 1, 3 and Shard B&#x27;s registers 1, 4, 2, 5, 3 combine into a merged sketch 2, 4, 5, 5, 3 by taking the maximum at each position, which then yields an estimate of the size of the union of A and B. Notes explain the max is associative and commutative so shard order and grouping never change the result, double-counting an item across shards is harmless because seeing the same hash twice can only leave a register where it was, and this is exactly what Redis PFMERGE and an OpenSearch coordinator do.\"></p>\n<p>Two properties of <code class=\"language-text\">max</code> make this safe:</p>\n<ul>\n<li><strong>Order does not matter.</strong> <code class=\"language-text\">max</code> is associative and commutative, so neither the order you merge shards in nor how you group them can change the answer.</li>\n<li><strong>Duplicates are harmless.</strong> An item that shows up on two shards is not double-counted. Its hash produces the same <code class=\"language-text\">ρ</code> in the same register on both, and <code class=\"language-text\">max(ρ, ρ)</code> is <code class=\"language-text\">ρ</code>.</li>\n</ul>\n<p>This is why a query engine can have each shard compute a tiny local sketch, ship only the sketch to a coordinator, and combine them there. You send twelve kilobytes per shard instead of the whole dataset. Redis exposes this directly as <a href=\"https://redis.io/docs/latest/develop/data-types/probabilistic/hyperloglogs/\"><code class=\"language-text\">PFMERGE</code></a>. OpenSearch does it internally every time you run a distributed cardinality aggregation.</p>\n<h2>Using it in Redis and OpenSearch</h2>\n<p>You rarely implement HyperLogLog yourself, because the systems you already run ship it.</p>\n<h3>Redis</h3>\n<p>In Redis, HyperLogLog is three commands, available <a href=\"https://redis.io/docs/latest/develop/data-types/probabilistic/hyperloglogs/\">since 2.8.9</a>. The example adds three IDs (one repeated) to a daily key, reads the count, and rolls two days into a weekly key:</p>\n<div class=\"gatsby-highlight\" data-language=\"text\"><pre class=\"language-text\"><code class=\"language-text\">PFADD workflows:2026-09-26 wf-4487 wf-4488 wf-4487\nPFCOUNT workflows:2026-09-26          # -&gt; 2, wf-4487 counted once\nPFMERGE workflows:week workflows:2026-09-26 workflows:2026-09-25</code></pre></div>\n<ul>\n<li><code class=\"language-text\">PFADD</code> folds items into the sketch.</li>\n<li><code class=\"language-text\">PFCOUNT</code> reads the estimate.</li>\n<li><code class=\"language-text\">PFMERGE</code> unions sketches into a new key. That is how you roll daily counts up into a weekly one without ever storing the IDs.</li>\n</ul>\n<h3>OpenSearch</h3>\n<p>In OpenSearch, the same machinery sits behind the <a href=\"https://docs.opensearch.org/latest/aggregations/metric/cardinality/\"><code class=\"language-text\">cardinality</code> aggregation</a>. It uses <a href=\"https://research.google/pubs/hyperloglog-in-practice-algorithmic-engineering-of-a-state-of-the-art-cardinality-estimation-algorithm/\">HyperLogLog++</a>, Google’s refinement of the original with 64-bit hashes and a sparse layout for small counts. This query counts distinct workflow IDs across the log indexes:</p>\n<div class=\"gatsby-highlight\" data-language=\"json\"><pre class=\"language-json\"><code class=\"language-json\">GET logs-*/_search\n<span class=\"token punctuation\">{</span>\n  <span class=\"token property\">\"size\"</span><span class=\"token operator\">:</span> <span class=\"token number\">0</span><span class=\"token punctuation\">,</span>\n  <span class=\"token property\">\"aggs\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">{</span>\n    <span class=\"token property\">\"distinct_workflows\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">{</span>\n      <span class=\"token property\">\"cardinality\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">{</span>\n        <span class=\"token property\">\"field\"</span><span class=\"token operator\">:</span> <span class=\"token string\">\"workflow_id\"</span><span class=\"token punctuation\">,</span>\n        <span class=\"token property\">\"precision_threshold\"</span><span class=\"token operator\">:</span> <span class=\"token number\">3000</span>\n      <span class=\"token punctuation\">}</span>\n    <span class=\"token punctuation\">}</span>\n  <span class=\"token punctuation\">}</span>\n<span class=\"token punctuation\">}</span></code></pre></div>\n<p>The <code class=\"language-text\">precision_threshold</code> is the same error-versus-memory knob from earlier, under a different name. Counts at or below the threshold are near-exact. Above it, you get the HyperLogLog estimate, which OpenSearch documents as <a href=\"https://docs.opensearch.org/latest/aggregations/metric/cardinality/\">typically within about 6% of the true value</a>. The default is <code class=\"language-text\">3000</code> and the maximum is <code class=\"language-text\">40000</code>. Raising it buys accuracy with memory, in the same <code class=\"language-text\">1.04/√m</code> trade.</p>\n<p>For a dashboard panel showing “distinct failing workflows over the last 24 hours,” the default is usually fine. Being off by a few percent on a trending number is not the kind of error that misleads anyone.</p>\n<h2>Where it bites</h2>\n<p>The estimate is the honest tradeoff, but a few sharper edges catch people.</p>\n<p><strong>The error is relative, not absolute.</strong> A 1% standard error on a true count of 50 million is a swing of half a million. That is fine for a trend line and wrong for anything you would put in a billing row or a compliance report. Distinct counts that feed money or audits need exact counting, full stop.</p>\n<p><strong>A sketch cannot list its members.</strong> It counts; it does not enumerate. If the next question after “how many distinct workflows failed?” is “which ones?”, a HyperLogLog has nothing to give you, because it never stored them. Use it for the count, and go back to the source when you need the list.</p>\n<p><strong>Intersections are a trap.</strong> Union is exact and cheap, but there is no direct way to intersect two sketches. The usual workaround is inclusion-exclusion, <code class=\"language-text\">|A ∩ B| = |A| + |B| − |A ∪ B|</code>. It combines the errors of three separate estimates, so it degrades badly when the sets are similar in size and the overlap is small. If you genuinely need set intersections at scale, HyperLogLog is the wrong tool; something like a <a href=\"https://en.wikipedia.org/wiki/MinHash\">MinHash</a> sketch fits better.</p>\n<p><strong>Small cardinalities were the original weak spot.</strong> The raw estimator is biased low when the true count is small relative to the number of registers. That is precisely why <a href=\"https://research.google/pubs/hyperloglog-in-practice-algorithmic-engineering-of-a-state-of-the-art-cardinality-estimation-algorithm/\">HyperLogLog++</a> added empirical bias correction and a sparse representation for that range. Redis and OpenSearch already include those fixes. If you hand-roll from the 2007 paper, you will see the low-end skew yourself.</p>\n<p><strong>Everything rests on the hash.</strong> HyperLogLog assumes the hash spreads inputs uniformly across the bit space. A weak hash, or one with structure in the high bits you use to pick registers, breaks that assumption, and the estimate drifts. Use a hash designed for even distribution: not a cryptographic one you picked for other reasons, and not something homegrown.</p>\n<h2>What I use it for</h2>\n<p>HyperLogLog earns its keep on the dashboard question that would otherwise be a full scan: distinct workflows, distinct users, or distinct error fingerprints over a time range wide enough that exact counting is off the table. It fits naturally into an observability setup. I have written about <a href=\"/blog/2026-07-01-opensearch-workflow-monitoring/\">building OpenSearch dashboards to watch workflow health</a> and about <a href=\"/blog/2026-09-03-opensearch-ism-index-lifecycle/\">keeping the indexes behind them from filling a disk</a>. Cardinality panels are the natural next thing to put on those dashboards. Now you know the sketch doing the counting and, more usefully, when to stop trusting it.</p>\n<p>The pattern generalizes past this one algorithm. A fixed-size, mergeable summary that answers a question approximately is often worth far more than an exact answer you cannot afford to compute. Distinct counts are just the cleanest example.</p>\n<hr>\n<p><em>Diagrams by M. Hassan Ahmed, released under CC0. No external image was used for this post; the figures are original work by the author.</em></p>","frontmatter":{"title":"HyperLogLog: Count Distinct at Scale in 12KB","date":"2026-09-26T00:00:00.000Z","description":"Counting unique items exactly costs memory that grows with your data. HyperLogLog estimates the cardinality of billions in about 12KB. Here's how it works.","thumbnail":{"childImageSharp":{"fluid":{"base64":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAABQAAAALCAIAAADwazoUAAAACXBIWXMAAAsSAAALEgHS3X78AAAB7ElEQVQoz12RyZKbMBCGOWbACCSBEBZm38xuYxvHYI+XmZrjpHLP+79GGjunVP3V1VL3V7+6JeG8BqG0cPpjfn1E4yV/fyTTNTpdvMMpe7/DMThO2flW3D79YURJAf36k5Jk1ZFVgbDvRduyOVfdpWwveXlK8yHODutqLJpzlO7zaoRqEO90GsoLAZSCVtKCBsiMNStRjUghoUJDhQSIxZhniCWzzBhKMgnedA+qGoNjCJRCXAmxlIoG26Vm5U+tEcssr3WzwXQ7w2ksbwMNzGncaC/8rWakmplrLFN0B+BE52sqqhcMOQh41YT7wnRbnuyIqKxm8O8f7vVuTRMtd8hIZA1gM8F2QZbVi4QIJCTM39jpniVbErYaz42fJ/b7y/r1Jf58s+mMcCTrYoZ1u6CifjnDE0Q6pP3N7yYr7TW3xH6t2Wte7OPr52o38e6IoxYZsYyWkmrGeFnSZU3tigdbHvdG2FG/NsOO57OzlfV4VbFVG9XTMtjq88A5WL4hPjuDIS8Oznhjh5PRH58a+HH07g/3evMeH0a2tb2NE+/ndViZBtGM31TrNXPJvE4Xpb4sDbclqwbegkVNIRE1cWoi6n+LhNFgL7yYnVUmKdhVaaASXyXBf9JZQniGwQq+mvpzGw1fcYHdHwv2F7XYUxz50t48AAAAAElFTkSuQmCC","aspectRatio":1.899441340782123,"src":"/static/d8c408d0f784bb15fc0ce20a2f9339e1/40a76/hero.png","srcSet":"/static/d8c408d0f784bb15fc0ce20a2f9339e1/c972b/hero.png 340w,\n/static/d8c408d0f784bb15fc0ce20a2f9339e1/27625/hero.png 680w,\n/static/d8c408d0f784bb15fc0ce20a2f9339e1/40a76/hero.png 1360w,\n/static/d8c408d0f784bb15fc0ce20a2f9339e1/ed396/hero.png 2000w","sizes":"(max-width: 1360px) 100vw, 1360px"}}}}}},"pageContext":{"slug":"/2026-09-26-hyperloglog-count-distinct-at-scale/","previous":"blog/2026-09-27-rope-position-embeddings-context-length/","next":"blog/2026-09-25-gpu-memory-serve-llm/"}},"staticQueryHashes":["32046230"]}