{"componentChunkName":"component---src-templates-blog-post-js","path":"/blog/2026-07-15-kubernetes-oomkilled-requests-vs-limits/","result":{"data":{"site":{"siteMetadata":{"title":"M.Hassan Ahmed","author":"Hassan11196"}},"markdownRemark":{"id":"2cbf563b-f059-5403-b205-cf7a8cb3f986","excerpt":"A service that ran fine all week starts dying.  shows it restarting, and  has one part that matters: Nothing in the code changed, and the image is the same one…","html":"<p>A service that ran fine all week starts dying. <code class=\"language-text\">kubectl get pods</code> shows it restarting, and <code class=\"language-text\">kubectl describe</code> has one part that matters:</p>\n<div class=\"gatsby-highlight\" data-language=\"text\"><pre class=\"language-text\"><code class=\"language-text\">    State:          Running\n    Last State:     Terminated\n      Reason:       OOMKilled\n      Exit Code:    137</code></pre></div>\n<p>Nothing in the code changed, and the image is the same one that shipped on Monday. What changed is how much memory the container needed. It crossed a line you set weeks ago and forgot about.</p>\n<p>This post is for engineers who run services on Kubernetes and have hit <code class=\"language-text\">OOMKilled</code> without a clear picture of what is going on. It covers what memory requests and limits actually do, why the two kills that mention memory are different events, and which settings actually stop the kills. I ran the <a href=\"/project/cms-workflow-operations/\">Kubernetes side of the CMS workflow operations stack</a> at CERN. There, a WMCore service quietly growing past its memory limit is the difference between a slow afternoon and a paged night.</p>\n<h2>Exit code 137 means “killed by a signal”</h2>\n<p>Start by reading the exit code correctly. <code class=\"language-text\">137</code> is <code class=\"language-text\">128 + 9</code>. The <code class=\"language-text\">128 +</code> is the shell convention for “terminated by a signal”, and <code class=\"language-text\">9</code> is <code class=\"language-text\">SIGKILL</code>.</p>\n<p>So the container did not throw an exception, did not log a stack trace, and did not get a chance to clean up. Something outside the process sent it a <code class=\"language-text\">SIGKILL</code>, which a process cannot catch or handle, and the process stopped mid-instruction. Pair that with <code class=\"language-text\">Reason: OOMKilled</code> and you know who pulled the trigger: the Linux kernel’s out-of-memory (OOM) killer. It acted because the container’s <a href=\"https://www.kernel.org/doc/html/latest/admin-guide/cgroup-v2.html\">cgroup</a> (the kernel mechanism that caps a group of processes’ resources) hit its memory ceiling.</p>\n<p>That detail rules things out. A <code class=\"language-text\">SIGKILL</code> from the OOM killer means no graceful shutdown ran, so any “on exit, flush the buffer” logic did not fire. If you are debugging lost writes alongside an OOMKill, that is why. The process did not get to say goodbye.</p>\n<h2>Why memory kills and CPU only slows down</h2>\n<p>CPU and memory look like the same kind of resource in a pod spec. Under pressure they behave nothing alike.</p>\n<p>CPU is <em>compressible</em>. When a container wants more CPU than its limit, the kernel <a href=\"https://kubernetes.io/docs/concepts/configuration/manage-resources-containers/#how-pods-with-resource-limits-are-run\">throttles</a> it. The container runs slower, and everyone survives.</p>\n<p>Memory is <em>incompressible</em>. You cannot run a process on “less memory than it is asking for right now.” Once a container’s allocations reach its memory limit, there is nothing to throttle. The kernel’s only lever is to kill something in the cgroup and reclaim the pages.</p>\n<p>So the mental model is simple: <strong>a CPU limit slows you down; a memory limit kills you.</strong> Every OOMKill follows from that asymmetry.</p>\n<h2>Requests and limits are read by different components</h2>\n<p>The pod spec asks for two memory numbers per container. Two different parts of Kubernetes read them, at two different times. The comments in this snippet say which part reads which:</p>\n<div class=\"gatsby-highlight\" data-language=\"yaml\"><pre class=\"language-yaml\"><code class=\"language-yaml\"><span class=\"token key atrule\">resources</span><span class=\"token punctuation\">:</span>\n  <span class=\"token key atrule\">requests</span><span class=\"token punctuation\">:</span>\n    <span class=\"token key atrule\">memory</span><span class=\"token punctuation\">:</span> <span class=\"token string\">\"512Mi\"</span>   <span class=\"token comment\"># what the SCHEDULER reserves</span>\n  <span class=\"token key atrule\">limits</span><span class=\"token punctuation\">:</span>\n    <span class=\"token key atrule\">memory</span><span class=\"token punctuation\">:</span> <span class=\"token string\">\"1Gi\"</span>     <span class=\"token comment\"># what the KERNEL enforces</span></code></pre></div>\n<p>The <strong>request</strong> is a scheduling input. The <a href=\"https://kubernetes.io/docs/concepts/scheduling-eviction/kube-scheduler/\">kube-scheduler</a> places your pod on a node only if the node’s existing memory requests plus yours still fit in the node’s allocatable memory. The request is a reservation. Once the pod is scheduled, the request has mostly done its job, and the running container is not held to it.</p>\n<p>The <strong>limit</strong> is a runtime input. It becomes the memory ceiling on the container’s cgroup, and the kernel enforces it directly. Cross it and the OOM killer runs.</p>\n<p>The scheduler never looks at limits when it places a pod. That causes the most dangerous surprise: <strong>limits can add up to more than the node has.</strong> The scheduler packed the node by requests and cheerfully let the limits overcommit it.</p>\n<p><img src=\"/77a080e6e9f9496080ace41f97f53229/requests-vs-limits.svg\" alt=\"Two memory bars for one node. The top bar shows three pods packed by their requests, which the scheduler guarantees fit inside node capacity. The bottom bar shows the same pods sized by their limits, which sum past the node edge, so if every pod uses its full limit at once the node runs out.\"></p>\n<p>Keep the gap between those two bars in mind, because both failure modes live there:</p>\n<ul>\n<li>Set one container’s limit too low, and that container OOMKills itself.</li>\n<li>Let the limits overcommit the node too far, and the whole node runs out of memory.</li>\n</ul>\n<h2>Two different kills that both say “memory”</h2>\n<p>When people say “the pod got OOMKilled,” they usually mean one specific thing. But Kubernetes has two distinct memory kills, and if you confuse them you end up debugging the wrong layer.</p>\n<p><img src=\"/01a4f7b116c1b6e850683c952f623c97/memory-kill-triggers.svg\" alt=\"Two panels. On the left, a single container&#x27;s working set climbs past its limit line and the kernel OOM killer terminates it with exit 137. On the right, a node&#x27;s memory fills past the kubelet eviction threshold, and the kubelet evicts whole pods worst-QoS-first: BestEffort, then Burstable over its request, then Guaranteed.\"></p>\n<p><strong>Container-limit OOMKill.</strong> One container reaches <em>its own</em> limit. The cgroup OOM killer terminates the process in that container, the kubelet reports <code class=\"language-text\">Reason: OOMKilled</code>, and the pod restarts according to its <code class=\"language-text\">restartPolicy</code>. This kill is local. The node might have gigabytes free; this one container simply asked for more than its own cap allowed. This is the kill you see in <code class=\"language-text\">kubectl describe pod</code>.</p>\n<p><strong>Node-pressure eviction.</strong> The <em>node</em> runs low on memory, whatever any single container’s limit is. The kubelet (the agent on each node that runs pods) <a href=\"https://kubernetes.io/docs/concepts/scheduling-eviction/node-pressure-eviction/\">monitors node memory</a>. When free memory drops below its eviction threshold, it proactively evicts whole pods to reclaim memory. This shows up as <code class=\"language-text\">Reason: Evicted</code>, not <code class=\"language-text\">OOMKilled</code>, often with a message like <code class=\"language-text\">The node was low on resource: memory</code>. The evicted pod may have been well within its own limit. It was chosen because the node as a whole was in trouble.</p>\n<p>To tell which one you hit, check the reason field and where it appears:</p>\n<ul>\n<li><code class=\"language-text\">OOMKilled</code> on the container in <code class=\"language-text\">describe pod</code> means a limit problem for that container.</li>\n<li><code class=\"language-text\">Evicted</code> on the pod, plus <code class=\"language-text\">MemoryPressure</code> on the node in <code class=\"language-text\">kubectl describe node</code>, means a node-packing problem.</li>\n</ul>\n<p>The two have different fixes, so read the field before you change anything.</p>\n<h2>QoS class decides who gets evicted first</h2>\n<p>When the kubelet evicts pods under node pressure, it does not choose at random. It ranks pods by their <a href=\"https://kubernetes.io/docs/concepts/workloads/pods/pod-qos/\">Quality of Service class</a> (QoS class). Kubernetes assigns the class based on how you set requests and limits:</p>\n<ul>\n<li><strong>Guaranteed</strong>: every container sets a memory (and CPU) limit <em>equal</em> to its request. Evicted last. Use this class for anything that must not be sacrificed to save the node.</li>\n<li><strong>Burstable</strong>: at least one container has a request or limit set, but the pod does not match the Guaranteed pattern. This is the common case. Under pressure, a Burstable pod using more than its request is a prime eviction target.</li>\n<li><strong>BestEffort</strong>: no requests or limits at all. Evicted first, always. Fine for a throwaway batch job, wrong for anything that holds state.</li>\n</ul>\n<p>In practice, this means the request is not only a scheduling number; it also sets your place in the eviction order. A pod running under its request is a poor eviction candidate. A pod running well over its request is the first Burstable pod to go. A request that reflects real steady-state usage keeps a healthy pod from becoming collateral damage when a noisy neighbor fills the node.</p>\n<h2>Diagnosing an OOMKill without guessing</h2>\n<p>Start with the pod, then widen out. The exit reason tells you which of the two kills you are dealing with. These three commands show the last exit reason per container, the full pod status and events, and whether the node itself is under memory pressure:</p>\n<div class=\"gatsby-highlight\" data-language=\"bash\"><pre class=\"language-bash\"><code class=\"language-bash\"><span class=\"token comment\"># Which containers were OOMKilled, and their last exit reason</span>\nkubectl get pod <span class=\"token operator\">&lt;</span>pod<span class=\"token operator\">></span> -o <span class=\"token assign-left variable\">jsonpath</span><span class=\"token operator\">=</span><span class=\"token string\">'{range .status.containerStatuses[*]}{.name}{\": \"}{.lastState.terminated.reason}{\" (\"}{.lastState.terminated.exitCode}{\")\"}{\"<span class=\"token entity\" title=\"\\n\">\\n</span>\"}{end}'</span>\n\n<span class=\"token comment\"># Full picture: events, restart count, the Terminated block</span>\nkubectl describe pod <span class=\"token operator\">&lt;</span>pod<span class=\"token operator\">></span>\n\n<span class=\"token comment\"># Is the NODE under memory pressure? (eviction vs container OOM)</span>\nkubectl describe node <span class=\"token operator\">&lt;</span>node<span class=\"token operator\">></span> <span class=\"token operator\">|</span> <span class=\"token function\">grep</span> -A5 Conditions</code></pre></div>\n<p>If it is a container OOMKill, the next question is how close the container runs to its limit in normal operation. Compare the limit against actual usage, not against the request. If you have metrics-server or Prometheus, the working set is the number to watch:</p>\n<div class=\"gatsby-highlight\" data-language=\"bash\"><pre class=\"language-bash\"><code class=\"language-bash\">kubectl <span class=\"token function\">top</span> pod <span class=\"token operator\">&lt;</span>pod<span class=\"token operator\">></span> --containers</code></pre></div>\n<p>That command gives you a single reading. Look at <code class=\"language-text\">container_memory_working_set_bytes</code> over time instead. The working set is roughly the memory the kernel cannot reclaim under pressure, and that is what the OOM decision is made against.</p>\n<p>For example, a container that idles at 300Mi and spikes to 1.1Gi on a large request needs a limit above that spike. A single snapshot from <code class=\"language-text\">kubectl top</code> will not show you the spike. Graph it.</p>\n<h2>Setting the numbers so the kills stop</h2>\n<p>There is no universal right value, but there is a reliable procedure.</p>\n<ol>\n<li><strong>Measure the working set under real load, including the peaks.</strong> Run the service through a representative workload and watch <code class=\"language-text\">container_memory_working_set_bytes</code>. You want the p99 of that curve, not the average, because the OOM killer does not care about your average.</li>\n<li><strong>Set the request at the steady working set.</strong> This is what the scheduler reserves, and it anchors your QoS ranking. Too low, and the node gets overcommitted and your pod becomes an eviction target. Too high, and you waste reserved memory the pod never uses and fit fewer pods per node.</li>\n<li><strong>Set the limit at the peak plus honest headroom.</strong> The limit must sit above the worst legitimate spike you measured, with margin for allocations you did not think to test. A limit equal to the average guarantees an OOMKill the first time real traffic arrives.</li>\n<li><strong>For anything stateful or latency-critical, make the request equal the limit.</strong> That gets the pod the Guaranteed QoS class, so the kubelet evicts it last. It also removes the overcommit gap for that pod entirely. You pay in reserved memory that sits idle, and for a service that must not die, that is the right trade.</li>\n</ol>\n<p>This is what step 4 looks like in the spec:</p>\n<div class=\"gatsby-highlight\" data-language=\"yaml\"><pre class=\"language-yaml\"><code class=\"language-yaml\"><span class=\"token key atrule\">resources</span><span class=\"token punctuation\">:</span>\n  <span class=\"token key atrule\">requests</span><span class=\"token punctuation\">:</span>\n    <span class=\"token key atrule\">memory</span><span class=\"token punctuation\">:</span> <span class=\"token string\">\"1Gi\"</span>\n  <span class=\"token key atrule\">limits</span><span class=\"token punctuation\">:</span>\n    <span class=\"token key atrule\">memory</span><span class=\"token punctuation\">:</span> <span class=\"token string\">\"1Gi\"</span>   <span class=\"token comment\"># request == limit  ->  Guaranteed QoS</span></code></pre></div>\n<h2>Failure modes that send people in circles</h2>\n<p><strong>The runtime does not see the cgroup limit.</strong> This is the classic one. A JVM from before container awareness sizes its heap from the <em>host’s</em> total memory, not the container limit. It happily grows a heap far past the cgroup cap, then gets OOMKilled while it still thinks it has plenty of “free” heap. Runtimes handle this differently:</p>\n<ul>\n<li>Modern JVMs honor the limit via <a href=\"https://docs.oracle.com/en/java/javase/17/docs/specs/man/java.html\"><code class=\"language-text\">UseContainerSupport</code></a> (on by default) plus <code class=\"language-text\">-XX:MaxRAMPercentage</code>.</li>\n<li>Go added <a href=\"https://pkg.go.dev/runtime#hdr-Environment_Variables\"><code class=\"language-text\">GOMEMLIMIT</code></a> for the same reason.</li>\n<li>A plain Python or Node process has no such knob and will run right up to the cgroup wall and die.</li>\n</ul>\n<p>If your OOMKilled container hosts a VM-based runtime, check that the runtime knows about the limit before you touch the limit itself.</p>\n<p><strong>CrashLoopBackOff that is really a repeating OOMKill.</strong> A container that OOMKills on startup, restarts, and OOMKills again lands in <code class=\"language-text\">CrashLoopBackOff</code>. The crash-loop label hides the memory cause. Always read <code class=\"language-text\">lastState.terminated.reason</code>. If it says <code class=\"language-text\">OOMKilled</code>, the fix is memory, not the backoff timing or the readiness probe.</p>\n<p><strong>Blaming the limit when the node was the problem.</strong> If the reason is <code class=\"language-text\">Evicted</code> and the node shows <code class=\"language-text\">MemoryPressure</code>, raising the pod’s limit does nothing, because the pod was not over its limit. The node was overcommitted. Fix it at the packing layer: reduce overcommit by setting honest requests across the workloads on that node, or add capacity. I covered the safe-rollout side of this in <a href=\"/blog/2026-07-03-argocd-sync-waves-ordering-rollout/\">ArgoCD sync waves</a>. Getting requests right keeps a GitOps deploy from quietly overpacking a node over successive releases.</p>\n<p><strong>A memory leak disguised as a bad limit.</strong> If a container’s working set climbs steadily and never levels off, raising the limit only spaces out the same OOMKills. A leak needs fixing, not a bigger cap. The tell is the shape of the curve: a genuine peak-and-return workload plateaus, while a leak trends up forever. I push exactly this into the dashboards described in <a href=\"/blog/2026-07-01-opensearch-workflow-monitoring/\">OpenSearch dashboards for workflow monitoring</a>. “Is this a spike or a leak?” is a question a graph answers in seconds and <code class=\"language-text\">kubectl describe</code> never will.</p>\n<p><strong>A liveness probe restarting an OOM victim into a loop.</strong> If a container is being OOMKilled and <em>also</em> has an aggressive liveness probe, you can lose track of which one restarted it. Keep the two concerns separate; I wrote about not overloading health checks in <a href=\"/blog/2026-07-07-kubernetes-liveness-readiness-startup-probes/\">Kubernetes probes</a>. An OOMKill is the kubelet reacting to the kernel, not to a failed probe, and the exit reason tells you which happened.</p>\n<h2>What I would do differently</h2>\n<p>Early on, I set memory limits by copying whatever number the last service used. That is how you end up with a fleet of limits unrelated to what anything actually needs. Some were far too tight and OOMKilled under normal peaks. Others were so loose that a leak ran for hours before anyone noticed.</p>\n<p>The habit that fixed it was boring:</p>\n<ul>\n<li>Measure each service’s working set under real load <em>before</em> setting the numbers.</li>\n<li>Set the request at the steady state and the limit above the measured peak.</li>\n<li>Promote anything that must not die to Guaranteed QoS by making the request equal the limit.</li>\n</ul>\n<p>Requests and limits stopped being magic numbers and became two measurements with two jobs.</p>\n<p>The other change was treating a restart with <code class=\"language-text\">OOMKilled</code> as a first-class alert, not something I noticed while debugging something else. A container that OOMKills is telling you its real memory need has outgrown the box you drew around it. You want to know that on the day it starts, not in the week the restarts get bad enough to page someone.</p>\n<h2>Where this runs</h2>\n<p>This is the memory discipline behind the <a href=\"/project/cms-workflow-operations/\">WMCore and Unified operations services</a>. They schedule Monte Carlo production and reconstruction for the CMS experiment across the Worldwide LHC Computing Grid (WLCG). Those services process bursty workloads of varying size, so the gap between steady state and peak is real and worth measuring rather than guessing.</p>\n<p>Getting requests and limits right is not glamorous work. But it decides whether a node packs cleanly and degrades gracefully, or instead wastes half its memory on unused reservations or falls over the first time three pods spike at once. The kernel’s OOM killer is not the problem. It is doing exactly what an expectation you never properly set told it to do.</p>\n<p><em>Diagrams by M. Hassan Ahmed, made for this post.</em></p>\n<p>Image credit: Diagrams created by M. Hassan Ahmed for this post, released under CC0 (public domain).</p>","frontmatter":{"title":"Kubernetes OOMKilled: Requests vs Limits","date":"2026-07-15T00:00:00.000Z","description":"Your pod died with OOMKilled and exit 137. What Kubernetes memory requests and limits actually do, what triggers the kill, and how to stop it happening.","thumbnail":{"childImageSharp":{"fluid":{"base64":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAABQAAAALCAIAAADwazoUAAAACXBIWXMAAAsSAAALEgHS3X78AAACJklEQVQozyVRy27TQBT1CvKo52mPZ8avafyIHduJ4zgpdVxaSlsEAipaBAiExJI9S36ADbsKVPGzjEG6ujq6M0fnnnsM4fW6mOh8ddZs366am3pzU9XX6/Z2qfvmdlG9rJs3dTMMpX/iyE64/f8yxkiOoQB26AR5lLdpdZSvuqTc6q7xoumjYpvX+/mqixZb288BOxwBMYZyglzDpDNgJ8iZQzYnIoMsndLItOIDGk1xCHGIcAhQYLGYOYllR9SaYaIACfWrAZ2UeEssSyyXZXulFidIlMStsFtCqjgJHRIyHCwXj453Tz2eaqyHFAcT5BuAJcRdQp5jN3VUi2Q2ZTGWFZYFIGpgEkWRI/mMWxEyLQo9hyqCgzH0NDlFsqDuet/9uDz/83h/v9l9scMSCU3WsopAVqUfnp/9Pj++uzj6HquKIp8STXaNwbDMz/pfTfMNiiovXnx89cmbVYAXUGtCu0zfX/X3oagxSV9vP989u7aQp5UnQBoT4sbZu777+YAIHpTzNGNeVhQt/KdsIX7R3Xl8jYAdsFlse183J3M5N6E7BsIwLcVVz8M9cRe7ts3LDeAlDwfPJhrsKbdjJFh52WWySpwYYyWsQ4zckcn12rE+LBD5LD/eHJ2ud6d2uEayQrIEyB8Oi0TqZk+Suo9WjSo5Ha5IkDcymTFFvk4V2JE2P9HZWjFkiWnNTHo4NTmGUv+zSYCwBzWgOiRPy5pAPDxgfwFwmlj+1mWLqQAAAABJRU5ErkJggg==","aspectRatio":1.899441340782123,"src":"/static/6399392464251e994ff9f039b0376535/40a76/hero.png","srcSet":"/static/6399392464251e994ff9f039b0376535/c972b/hero.png 340w,\n/static/6399392464251e994ff9f039b0376535/27625/hero.png 680w,\n/static/6399392464251e994ff9f039b0376535/40a76/hero.png 1360w,\n/static/6399392464251e994ff9f039b0376535/ed396/hero.png 2000w","sizes":"(max-width: 1360px) 100vw, 1360px"}}}}}},"pageContext":{"slug":"/2026-07-15-kubernetes-oomkilled-requests-vs-limits/","previous":"blog/2026-07-16-cancel-llm-stream-react-abortcontroller/","next":"blog/2026-07-19-indirect-prompt-injection-rag/"}},"staticQueryHashes":["32046230"]}