{"componentChunkName":"component---src-templates-blog-post-js","path":"/blog/2026-07-25-keda-autoscale-kubernetes-queue-depth/","result":{"data":{"site":{"siteMetadata":{"title":"M.Hassan Ahmed","author":"Hassan11196"}},"markdownRemark":{"id":"8393e977-7768-556b-9aef-153b2b7900d9","excerpt":"A worker that pulls jobs off a queue is the most common shape of backend code I run on Kubernetes. For a long time it was also the thing I autoscaled worst. The…","html":"<p>A worker that pulls jobs off a queue is the most common shape of backend code I run on Kubernetes. For a long time it was also the thing I autoscaled worst. The default tool is the <a href=\"https://kubernetes.io/docs/tasks/run-application/horizontal-pod-autoscale/\">HorizontalPodAutoscaler</a> (HPA), which adds or removes pods based on a metric, and out of the box that metric is CPU. That works for a request-serving API, where more traffic means more CPU. It falls apart for a queue worker, because the whole point of a queue is that work can pile up before anyone does anything with it.</p>\n<p>Picture a worker that fetches a URL, waits on the response, and writes a row. It spends most of its wall-clock time waiting on the network. So a single pod might sit at 30% CPU while a thousand messages stack up behind it. The HPA sees 30%, decides everything is calm, and adds no pods. The backlog grows, latency climbs, and the signal you are scaling on never moves. You were measuring the wrong thing.</p>\n<p>This post is for people running queue-driven workers on Kubernetes who have watched a backlog build while the HPA sat still. It covers:</p>\n<ul>\n<li>why CPU is the wrong signal;</li>\n<li>what <a href=\"https://keda.sh/\">KEDA</a> actually does under the hood;</li>\n<li>a <code class=\"language-text\">ScaledObject</code> you can copy, and the arithmetic behind it;</li>\n<li>how scale-to-zero works;</li>\n<li>the failure modes that cost real time.</li>\n</ul>\n<p>My running example is the <a href=\"/project/cms-workflow-operations/\">CMS workflow operations</a> stack I maintained at CERN. There, Monte Carlo production and reconstruction requests are queued and processed across grid sites, and “how many workers should be running right now” has a moving answer.</p>\n<h2>Queue depth measures the work; CPU only approximates it</h2>\n<p>The HPA scales on a metric because it assumes the metric tracks the work. For a CPU-bound service, that holds. For an I/O-bound worker, it does not, and the backlog hides in the gap between the proxy (CPU) and the real thing (waiting work).</p>\n<p><img src=\"/78b719b8692327d6337b39b00398a74d/queue-vs-cpu.svg\" alt=\"A time chart contrasting queue depth, which spikes when a burst of jobs arrives, against worker CPU, which stays near 30% and never crosses the HPA target, so a CPU-based autoscaler adds no pods while the backlog grows\"></p>\n<p>The number that actually describes the work sits in the message broker: how many messages are waiting. Scale on that, and the feedback loop closes. More messages bring more pods, the pods drain the queue, the count falls, and the pods go away again.</p>\n<p>You could wire this up yourself with the HPA’s <a href=\"https://kubernetes.io/docs/tasks/run-application/horizontal-pod-autoscale/#scaling-on-custom-metrics\">external metrics API</a> and a custom metrics adapter. But you would be writing and running a metrics server for each source. KEDA is that adapter, already written for a long list of sources.</p>\n<h2>What KEDA does: it feeds the HPA instead of replacing it</h2>\n<p>KEDA (Kubernetes Event-Driven Autoscaling) is a <a href=\"https://www.cncf.io/projects/keda/\">CNCF</a> project that graduated in 2023. It does not replace the HPA; it feeds it. What made it click for me is that KEDA and the HPA split the job:</p>\n<ul>\n<li><strong>KEDA owns the 0-to-1 transition.</strong> A plain HPA cannot scale a Deployment to zero, and it cannot bring one back from zero, because with no pods there are no pod metrics to read. KEDA watches the event source directly. So it can start the first pod when work appears and remove the last one when the queue is empty.</li>\n<li><strong>The HPA owns 1-to-N.</strong> Once at least one pod is running, KEDA’s metrics adapter serves the queue depth as an external metric, and a normal HPA does the scaling math on it. KEDA generates that HPA object for you from your config.</li>\n</ul>\n<p>So you write one custom resource, a <code class=\"language-text\">ScaledObject</code>, and KEDA turns it into a metrics adapter plus a managed HPA.</p>\n<p><img src=\"/3c12cf3247b504da71b2b5afd380afc3/keda-scaling-flow.svg\" alt=\"KEDA polls the message queue and serves its depth as an external metric; the built-in HorizontalPodAutoscaler divides that depth by the per-pod target to pick a replica count, while KEDA handles the activation from zero and the scale back to zero after the cooldown\"></p>\n<h2>A ScaledObject for a queue worker</h2>\n<p>Here is a <code class=\"language-text\">ScaledObject</code> for a worker that reads a <a href=\"https://keda.sh/docs/latest/scalers/rabbitmq-queue/\">RabbitMQ</a> queue. It targets the existing Deployment by name and points a trigger at the broker:</p>\n<div class=\"gatsby-highlight\" data-language=\"yaml\"><pre class=\"language-yaml\"><code class=\"language-yaml\"><span class=\"token key atrule\">apiVersion</span><span class=\"token punctuation\">:</span> keda.sh/v1alpha1\n<span class=\"token key atrule\">kind</span><span class=\"token punctuation\">:</span> ScaledObject\n<span class=\"token key atrule\">metadata</span><span class=\"token punctuation\">:</span>\n  <span class=\"token key atrule\">name</span><span class=\"token punctuation\">:</span> reco<span class=\"token punctuation\">-</span>worker\n<span class=\"token key atrule\">spec</span><span class=\"token punctuation\">:</span>\n  <span class=\"token key atrule\">scaleTargetRef</span><span class=\"token punctuation\">:</span>\n    <span class=\"token key atrule\">name</span><span class=\"token punctuation\">:</span> reco<span class=\"token punctuation\">-</span>worker          <span class=\"token comment\"># the Deployment to scale</span>\n  <span class=\"token key atrule\">minReplicaCount</span><span class=\"token punctuation\">:</span> <span class=\"token number\">0</span>            <span class=\"token comment\"># allow scale to zero when idle</span>\n  <span class=\"token key atrule\">maxReplicaCount</span><span class=\"token punctuation\">:</span> <span class=\"token number\">40</span>           <span class=\"token comment\"># ceiling, so a bad producer can't melt the cluster</span>\n  <span class=\"token key atrule\">pollingInterval</span><span class=\"token punctuation\">:</span> <span class=\"token number\">20</span>           <span class=\"token comment\"># seconds between queue checks (default 30)</span>\n  <span class=\"token key atrule\">cooldownPeriod</span><span class=\"token punctuation\">:</span> <span class=\"token number\">300</span>           <span class=\"token comment\"># wait 5 min of no activity before going to zero</span>\n  <span class=\"token key atrule\">triggers</span><span class=\"token punctuation\">:</span>\n    <span class=\"token punctuation\">-</span> <span class=\"token key atrule\">type</span><span class=\"token punctuation\">:</span> rabbitmq\n      <span class=\"token key atrule\">metadata</span><span class=\"token punctuation\">:</span>\n        <span class=\"token key atrule\">protocol</span><span class=\"token punctuation\">:</span> auto\n        <span class=\"token key atrule\">queueName</span><span class=\"token punctuation\">:</span> reconstruction\n        <span class=\"token key atrule\">mode</span><span class=\"token punctuation\">:</span> QueueLength\n        <span class=\"token key atrule\">value</span><span class=\"token punctuation\">:</span> <span class=\"token string\">\"50\"</span>             <span class=\"token comment\"># target messages per pod</span>\n        <span class=\"token key atrule\">activationValue</span><span class=\"token punctuation\">:</span> <span class=\"token string\">\"5\"</span>    <span class=\"token comment\"># need >5 waiting before starting the first pod</span>\n      <span class=\"token key atrule\">authenticationRef</span><span class=\"token punctuation\">:</span>\n        <span class=\"token key atrule\">name</span><span class=\"token punctuation\">:</span> rabbitmq<span class=\"token punctuation\">-</span>conn</code></pre></div>\n<p>Four numbers control the behavior, and each maps to a decision you actually care about:</p>\n<ul>\n<li><code class=\"language-text\">value: &quot;50&quot;</code> is the target backlog per pod. It is not a maximum concurrency; it is the queue depth per pod that the HPA aims to hold. A lower value means more pods for the same queue.</li>\n<li><code class=\"language-text\">activationValue: &quot;5&quot;</code> gates the jump from zero. Below six waiting messages, the worker stays at zero and costs nothing. It is deliberately separate from <code class=\"language-text\">value</code>, for reasons I come back to below.</li>\n<li><code class=\"language-text\">maxReplicaCount: 40</code> limits the blast radius. When you autoscale on an external number, a misbehaving producer can ask for unbounded pods, so this ceiling is not optional.</li>\n<li><code class=\"language-text\">cooldownPeriod: 300</code> is how long the queue must stay quiet before KEDA removes the last pod. It applies only to the final step down to zero.</li>\n</ul>\n<h2>The replica math is the standard HPA formula</h2>\n<p>KEDA picks the replica count with no magic. It is the standard <a href=\"https://kubernetes.io/docs/tasks/run-application/horizontal-pod-autoscale/#algorithm-details\">HPA algorithm</a>, with queue depth as the metric:</p>\n<div class=\"gatsby-highlight\" data-language=\"text\"><pre class=\"language-text\"><code class=\"language-text\">desiredReplicas = ceil(currentQueueDepth / value)</code></pre></div>\n<p>With <code class=\"language-text\">value: &quot;50&quot;</code> and 512 messages waiting, that is <code class=\"language-text\">ceil(512 / 50) = 11</code> pods. Drain the queue to 90, and the HPA settles toward <code class=\"language-text\">ceil(90 / 50) = 2</code>.</p>\n<p>The target is a ratio, not a rate, and that trips people up. It does not say “50 messages per second per pod.” It says “keep roughly 50 messages of backlog behind each pod.” If your jobs are slow, 50 messages of backlog might be minutes of work. So tune <code class=\"language-text\">value</code> against how long one message takes, not against an abstract throughput number.</p>\n<p>Because the HPA does the 1-to-N math, its scale-down <a href=\"https://kubernetes.io/docs/tasks/run-application/horizontal-pod-autoscale/#stabilization-window\">stabilization window</a> still applies (five minutes by default). The window smooths out flapping between, say, 4 and 5 pods. The <code class=\"language-text\">cooldownPeriod</code> is a different setting that only governs the very last step to zero. If you confuse the two, you will tune the wrong knob when scale-down feels sluggish.</p>\n<h2>Scale to zero, and why activation is a separate number</h2>\n<p>Scale-to-zero is the feature people come for. The <code class=\"language-text\">activationValue</code> is the part they skip and then get burned by.</p>\n<p><code class=\"language-text\">value</code> decides how many pods you run once you are running. <code class=\"language-text\">activationValue</code> decides whether you run at all. They are separate because the HPA math breaks down at zero: with zero replicas, there is no current metric to divide. So KEDA has to make the 0-to-1 decision itself, based on a threshold you set.</p>\n<p>Suppose you set only <code class=\"language-text\">value: &quot;50&quot;</code> and leave activation at its default of 0. Then a single stray message wakes the whole Deployment. For a worker whose image takes 40 seconds to pull and warm up, bouncing between 0 and 1 on one-off messages is worse than running one pod all the time. Set <code class=\"language-text\">activationValue</code> above the noise floor, so you only cold-start when there is real work.</p>\n<p>The other half of the tradeoff is the cold start itself. Coming back from zero costs a scheduling delay, an image pull, and whatever your app does before it reads the first message. If tail latency matters more than the cost of an idle pod, do not scale to zero. Set <code class=\"language-text\">minReplicaCount: 1</code> and keep one pod warm.</p>\n<h2>Where it breaks</h2>\n<p>Most of the sharp edges are not in KEDA itself. They come from the fact that your pod count now depends on a number a producer controls.</p>\n<p><strong>Long jobs versus a shrinking queue.</strong> Queue depth counts messages that are waiting. With most brokers, a message being processed has already left the ready count. If each job takes ten minutes, the depth can read low while every pod is busy, so the HPA tries to scale <em>down</em> mid-job. Three things help:</p>\n<ul>\n<li>For RabbitMQ, count unacknowledged plus ready messages rather than ready alone.</li>\n<li>Keep <code class=\"language-text\">terminationGracePeriodSeconds</code> long enough for a pod to finish its in-flight work.</li>\n<li>Rely on the stabilization window, so a brief dip does not evict a busy worker.</li>\n</ul>\n<p><strong>Poison messages scale you to the ceiling.</strong> A poison message is one that always fails, gets requeued, and fails again, which keeps the depth high forever. KEDA does the only thing it can and adds pods. You sit pinned at <code class=\"language-text\">maxReplicaCount</code>, paying for work that will never succeed. Queue-depth autoscaling assumes messages eventually leave the queue. Pair it with a dead-letter queue and a retry cap, so failures exit instead of recirculating.</p>\n<p><strong>Polling is not instant.</strong> KEDA checks the source every <code class=\"language-text\">pollingInterval</code> seconds (30 by default), so a burst that arrives and clears within one interval may never be seen. A tighter interval helps, but it adds load on the broker’s management API, which for RabbitMQ is not free. For genuinely spiky traffic, a shorter interval plus a small <code class=\"language-text\">minReplicaCount</code> beats polling every two seconds.</p>\n<p><strong>The bottleneck moves downstream.</strong> Autoscaling the workers is easy. The shared database they all write to is not elastic. Scale from 2 workers to 40, and you may have just pointed 40 concurrent writers at a connection pool sized for 5. I have traced a “scaling problem” that was really Postgres refusing connections. Cap <code class=\"language-text\">maxReplicaCount</code> at what the <em>slowest</em> downstream dependency can absorb, not at what the queue wants.</p>\n<p><strong>Non-idempotent work punishes scale events.</strong> Every scale-down can kill a pod mid-message. If your handler is not idempotent (safe to run twice on the same message), a redelivered message writes twice. This is a property of your consumer, not of KEDA. But autoscaling makes pod churn routine, so it exposes bugs that a fixed replica count hid. Make the handler idempotent before you make the fleet elastic.</p>\n<h2>What I would set up first</h2>\n<p>If I were adding this to a queue worker today, I would go in this order:</p>\n<ol>\n<li>Make the worker idempotent and have it dead-letter poison messages. <em>Then</em> add the <code class=\"language-text\">ScaledObject</code>.</li>\n<li>Start with <code class=\"language-text\">minReplicaCount: 1</code>, and skip scale-to-zero until I have watched the metric behave for a week. A cold start hiding behind a p99 latency spike is annoying to diagnose after the fact.</li>\n<li>Set <code class=\"language-text\">maxReplicaCount</code> from the downstream limit, not from the queue.</li>\n<li>Pick <code class=\"language-text\">value</code> by timing one message, not by guessing a throughput number.</li>\n</ol>\n<p>Queue-depth scaling fits the <a href=\"/project/cms-workflow-operations/\">CMS workflow tooling</a> because the work there is bursty and genuinely queued. Production and reconstruction requests arrive in waves. The useful signal is always the backlog, never the CPU of a worker that spends its life waiting on grid I/O.</p>\n<p>The same reasoning is behind treating <a href=\"/blog/2026-07-15-kubernetes-oomkilled-requests-vs-limits/\">requests and limits as a reliability setting</a> rather than a formality, and behind <a href=\"/blog/2026-07-03-argocd-sync-waves-ordering-rollout/\">ordering a rollout with ArgoCD sync waves</a>. On Kubernetes, the defaults are a reasonable starting point and a poor finish line. Measure the thing that describes the work, then scale on it.</p>\n<p><em>Diagrams by the author, released under <a href=\"https://creativecommons.org/publicdomain/zero/1.0/\">CC0</a>. No external image was used for the thumbnail.</em></p>","frontmatter":{"title":"Scale Kubernetes Workers on Queue Depth with KEDA","date":"2026-07-25T00:00:00.000Z","description":"A CPU-based HPA can't see a backed-up queue, so workers fall behind. How KEDA autoscales Kubernetes workers on queue depth, and where it breaks.","thumbnail":{"childImageSharp":{"fluid":{"base64":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAABQAAAALCAIAAADwazoUAAAACXBIWXMAAAsSAAALEgHS3X78AAACAUlEQVQoz02R6YvTQBiH86HadpOZzGQyyeRqmrSbzdl021xt0rIriKAoon73qJ9FVvBf8NjFA/9hJ7sihYeX37w87xyMwKyOY/S1pawJ48fL9YsofbI4f57lz1bFqyR7ykl5Ll/yPne4eet3whiZnBE0IPWQPvfDYlFcBOkmyDbx+S4vL5PVLlxseSfMWzdYcQeQ6RCysWwKInYlxREVHagm0pwTrA8kfB+Qe5LCGYh4ICp3y4GkDKEKVRsQC1EfEE8YAd31L9rt56b+sN1c1eXH8Ox1Eh6i8A0njt5H4ds4OsTRu74fH+ryU1NdzcMHin4qDAF1Jo/2zZ+m+LItr/fN767+vt/8aqvrturDrr7pa/Ozrb519U1X/9iWXzWrEpEjjKEBsG/Zl5StdaNhRkv1ghkbnlV9bZgd1UuNVczsiLY0zB3PlJW285BPCSfQkJBLtTUmiUIzRV0AFMo4kZUEkVjGKSIJUmJMMoACQpeKmmA1VbWVJE/6k8eQjYDGHz+UNAmbjpe482wySyGxvSCXVUdlc2oElpuJyOQON7nPp4QxYLfwLYyRpMtkcpbUp3HFK3PiOG8tL7O9pR/UQbSTkDO6k3v6Yf0IDVLf9As37LCRUiendub42XSe8es4Xoz12UjS/vsCXxwjIpv/IdJmkEwBcaHqQvIPRKciso7lv5mmWZvH4n2kAAAAAElFTkSuQmCC","aspectRatio":1.899441340782123,"src":"/static/087e118db1fb141ff439e029fa956a73/40a76/hero.png","srcSet":"/static/087e118db1fb141ff439e029fa956a73/c972b/hero.png 340w,\n/static/087e118db1fb141ff439e029fa956a73/27625/hero.png 680w,\n/static/087e118db1fb141ff439e029fa956a73/40a76/hero.png 1360w,\n/static/087e118db1fb141ff439e029fa956a73/ed396/hero.png 2000w","sizes":"(max-width: 1360px) 100vw, 1360px"}}}}}},"pageContext":{"slug":"/2026-07-25-keda-autoscale-kubernetes-queue-depth/","previous":"blog/2026-07-26-llm-quantization-int8-gptq-awq/","next":"blog/2026-07-24-speculative-decoding-llm-inference/"}},"staticQueryHashes":["32046230"]}