{"componentChunkName":"component---src-templates-blog-post-js","path":"/blog/2026-08-08-kubernetes-jobs-cronjobs-batch-workflows/","result":{"data":{"site":{"siteMetadata":{"title":"M.Hassan Ahmed","author":"Hassan11196"}},"markdownRemark":{"id":"ad3f2662-0583-52a3-b8fb-a6d92630e63c","excerpt":"Most of what you deploy to Kubernetes is meant to run forever. A Deployment keeps a web server alive: if a pod dies, the Deployment replaces it, indefinitely…","html":"<p>Most of what you deploy to Kubernetes is meant to run forever. A <a href=\"https://kubernetes.io/docs/concepts/workloads/controllers/deployment/\">Deployment</a> keeps a web server alive: if a pod dies, the Deployment replaces it, indefinitely, because a server is never supposed to be “done”. Then you need to run a database migration, reprocess yesterday’s dataset, or generate a nightly report, and the model flips. Now the point of the work is to finish, and “keep restarting forever” is exactly the wrong behavior.</p>\n<p>That is what <a href=\"https://kubernetes.io/docs/concepts/workloads/controllers/job/\">Jobs</a> and <a href=\"https://kubernetes.io/docs/concepts/workloads/controllers/cron-job/\">CronJobs</a> are for. A Job runs one or more pods until a set number of them exit successfully, then stops. A CronJob creates Jobs on a schedule. The API looks small, and your first Job will work on the first try. The problems show up later:</p>\n<ul>\n<li>a Job that retries a failure that can never succeed, a dozen times;</li>\n<li>finished pods piling up until you cannot list them;</li>\n<li>a CronJob quietly stacking overlapping runs because the previous one ran long.</li>\n</ul>\n<p>This post is for engineers who are comfortable with Deployments and now have finite or scheduled work to run. I maintained the workflow tooling for CMS production at CERN (<a href=\"/project/cms-workflow-operations/\">WMCore + Unified</a>), where batch is the whole job rather than a side task. Most of what follows is something I have either relied on or been burned by. I cover the completion model, retries and deadlines, cleanup, and then CronJobs and the ways both fail.</p>\n<h2>How a Job differs from a Deployment</h2>\n<p>The first thing that trips people up is <code class=\"language-text\">restartPolicy</code>. A Deployment’s pods use <code class=\"language-text\">Always</code>, because a server that exits should come back. A Job’s pods cannot use <code class=\"language-text\">Always</code> at all; the only legal values are <code class=\"language-text\">OnFailure</code> and <code class=\"language-text\">Never</code>. That one field captures the whole idea. A Job pod that exits with code 0 is <em>done</em>, and the controller must not treat a clean exit as something to recover from.</p>\n<p>Here is a minimal Job. It runs one reprocessing script, allows four retries, and sets memory requests and limits:</p>\n<div class=\"gatsby-highlight\" data-language=\"yaml\"><pre class=\"language-yaml\"><code class=\"language-yaml\"><span class=\"token key atrule\">apiVersion</span><span class=\"token punctuation\">:</span> batch/v1\n<span class=\"token key atrule\">kind</span><span class=\"token punctuation\">:</span> Job\n<span class=\"token key atrule\">metadata</span><span class=\"token punctuation\">:</span>\n  <span class=\"token key atrule\">name</span><span class=\"token punctuation\">:</span> dataset<span class=\"token punctuation\">-</span>reprocess\n<span class=\"token key atrule\">spec</span><span class=\"token punctuation\">:</span>\n  <span class=\"token key atrule\">backoffLimit</span><span class=\"token punctuation\">:</span> <span class=\"token number\">4</span>\n  <span class=\"token key atrule\">template</span><span class=\"token punctuation\">:</span>\n    <span class=\"token key atrule\">spec</span><span class=\"token punctuation\">:</span>\n      <span class=\"token key atrule\">restartPolicy</span><span class=\"token punctuation\">:</span> Never\n      <span class=\"token key atrule\">containers</span><span class=\"token punctuation\">:</span>\n        <span class=\"token punctuation\">-</span> <span class=\"token key atrule\">name</span><span class=\"token punctuation\">:</span> reprocess\n          <span class=\"token key atrule\">image</span><span class=\"token punctuation\">:</span> registry.example/reprocess<span class=\"token punctuation\">:</span><span class=\"token number\">2.1</span>\n          <span class=\"token key atrule\">command</span><span class=\"token punctuation\">:</span> <span class=\"token punctuation\">[</span><span class=\"token string\">'python'</span><span class=\"token punctuation\">,</span> <span class=\"token string\">'reprocess.py'</span><span class=\"token punctuation\">,</span> <span class=\"token string\">'--date'</span><span class=\"token punctuation\">,</span> <span class=\"token string\">'2026-08-07'</span><span class=\"token punctuation\">]</span>\n          <span class=\"token key atrule\">resources</span><span class=\"token punctuation\">:</span>\n            <span class=\"token key atrule\">requests</span><span class=\"token punctuation\">:</span> <span class=\"token punctuation\">{</span> <span class=\"token key atrule\">cpu</span><span class=\"token punctuation\">:</span> <span class=\"token string\">'1'</span><span class=\"token punctuation\">,</span> <span class=\"token key atrule\">memory</span><span class=\"token punctuation\">:</span> <span class=\"token string\">'2Gi'</span> <span class=\"token punctuation\">}</span>\n            <span class=\"token key atrule\">limits</span><span class=\"token punctuation\">:</span> <span class=\"token punctuation\">{</span> <span class=\"token key atrule\">memory</span><span class=\"token punctuation\">:</span> <span class=\"token string\">'2Gi'</span> <span class=\"token punctuation\">}</span></code></pre></div>\n<p>The pod runs <code class=\"language-text\">reprocess.py</code>. If it exits 0, the Job is complete. If it exits non-zero, that counts as a failure. What happens next depends on two fields most people copy without reading, <code class=\"language-text\">restartPolicy</code> and <code class=\"language-text\">backoffLimit</code>, which I cover below.</p>\n<p>Note the memory <code class=\"language-text\">requests</code> and <code class=\"language-text\">limits</code>. Batch pods get <a href=\"/blog/2026-07-15-kubernetes-oomkilled-requests-vs-limits/\">OOMKilled</a> the same way service pods do. A Job that dies 90% of the way through a four-hour run because you set memory too low is a special kind of painful.</p>\n<h2>Completions and parallelism: three shapes of batch work</h2>\n<p>A Job is not always one pod. Two fields, <code class=\"language-text\">completions</code> and <code class=\"language-text\">parallelism</code>, turn the same object into three different shapes of work. Knowing which shape you want is most of using Jobs well.</p>\n<p><img src=\"/e381c670e5e0150cbc0384d9e5893016/job-patterns.svg\" alt=\"Three ways a Job runs to completion. A single run uses completions 1 and parallelism 1 for one migration or report. A fixed count in parallel uses completions 6 and parallelism 3 so six indexed pods run three at a time, each owning a shard. A work queue uses parallelism 3 with no completions so workers pull from a shared queue until it is drained. All three end when the success condition is met, restartPolicy must be Never or OnFailure, and a Job that never reaches completions keeps retrying until backoffLimit stops it.\"></p>\n<ul>\n<li><strong>One and done.</strong> Leave both fields at their default of 1. One pod runs once. This is the migration, report, or dump case, and it covers most Jobs.</li>\n<li><strong>Fixed count, run in parallel.</strong> Set <code class=\"language-text\">completions: 6, parallelism: 3</code>, and the controller keeps three pods running until six have succeeded. This is how you shard a fixed amount of work: split your input into six slices and process them three at a time. With <a href=\"https://kubernetes.io/docs/concepts/workloads/controllers/job/#completion-mode\"><code class=\"language-text\">completionMode: Indexed</code></a> (stable since Kubernetes 1.24), each pod gets a <code class=\"language-text\">JOB_COMPLETION_INDEX</code> from 0 to 5 in its environment. Pod 3 can then deterministically pick slice 3, with no coordination between pods.</li>\n<li><strong>Work queue.</strong> Set <code class=\"language-text\">parallelism</code> but leave <code class=\"language-text\">completions</code> unset, and point every pod at a shared queue (Redis, a message broker, a database table). Each worker pulls items until the queue is empty, then exits. The Job finishes when all the pods have exited successfully. Use this pattern when you do not know the item count up front, or when items are cheap enough that static sharding would leave some pods idle while others grind.</li>\n</ul>\n<p>I use the indexed pattern most. “Here are N things, process them” describes most batch work, and the index removes the coordination problem entirely. There is no leader, no locking, and no double-processing: the pod’s index <em>is</em> its assignment.</p>\n<h2>When a Job fails: backoff, deadlines, and which failures to retry</h2>\n<p>A batch pod can fail for two very different reasons: a transient problem that a retry would fix, or a permanent one that no retry will. Kubernetes historically handled both with one blunt instrument, <code class=\"language-text\">backoffLimit</code>. It is the number of retries before the Job gives up and is marked <code class=\"language-text\">Failed</code>. It defaults to 6. Failed pods are recreated with an <a href=\"https://kubernetes.io/docs/concepts/workloads/controllers/job/#pod-backoff-failure-policy\">exponential back-off</a> that starts at 10 seconds and caps at 6 minutes.</p>\n<p>The trap is that <code class=\"language-text\">backoffLimit</code> cannot tell a flaky failure from a doomed one. A network blip during a download is worth retrying. A <code class=\"language-text\">SIGKILL</code> from the OOM killer, or an exit code that means “your input file is malformed”, will fail the same way on all six retries. You just wait through six growing back-offs to reach the failure you already had after the first pod.</p>\n<p><a href=\"https://kubernetes.io/docs/concepts/workloads/controllers/job/#pod-failure-policy\">Pod failure policy</a> (<code class=\"language-text\">podFailurePolicy</code>, stable since Kubernetes 1.31) fixes this. It lets you branch on the actual exit code, or on the reason a pod was killed. This example has two rules:</p>\n<div class=\"gatsby-highlight\" data-language=\"yaml\"><pre class=\"language-yaml\"><code class=\"language-yaml\"><span class=\"token key atrule\">spec</span><span class=\"token punctuation\">:</span>\n  <span class=\"token key atrule\">backoffLimit</span><span class=\"token punctuation\">:</span> <span class=\"token number\">4</span>\n  <span class=\"token key atrule\">podFailurePolicy</span><span class=\"token punctuation\">:</span>\n    <span class=\"token key atrule\">rules</span><span class=\"token punctuation\">:</span>\n      <span class=\"token comment\"># A specific \"bad input\" exit code is not worth retrying, so fail fast.</span>\n      <span class=\"token punctuation\">-</span> <span class=\"token key atrule\">action</span><span class=\"token punctuation\">:</span> FailJob\n        <span class=\"token key atrule\">onExitCodes</span><span class=\"token punctuation\">:</span> <span class=\"token punctuation\">{</span> <span class=\"token key atrule\">operator</span><span class=\"token punctuation\">:</span> In<span class=\"token punctuation\">,</span> <span class=\"token key atrule\">values</span><span class=\"token punctuation\">:</span> <span class=\"token punctuation\">[</span><span class=\"token number\">42</span><span class=\"token punctuation\">]</span> <span class=\"token punctuation\">}</span>\n      <span class=\"token comment\"># Pods evicted by the node (preemption, drain) shouldn't burn a retry.</span>\n      <span class=\"token punctuation\">-</span> <span class=\"token key atrule\">action</span><span class=\"token punctuation\">:</span> Ignore\n        <span class=\"token key atrule\">onPodConditions</span><span class=\"token punctuation\">:</span>\n          <span class=\"token punctuation\">-</span> <span class=\"token key atrule\">type</span><span class=\"token punctuation\">:</span> DisruptionTarget</code></pre></div>\n<p>Each rule saves something real:</p>\n<ul>\n<li><strong><code class=\"language-text\">FailJob</code> on exit code 42.</strong> Your reprocess script can <code class=\"language-text\">sys.exit(42)</code> on a corrupt input, and the Job fails immediately instead of retrying a file that will never parse.</li>\n<li><strong><code class=\"language-text\">Ignore</code> on <code class=\"language-text\">DisruptionTarget</code>.</strong> A pod the cluster killed for its own reasons (a node drain, a spot-instance reclaim) does not count against your four real retries. A maintenance window no longer uses up the budget you meant for genuine application errors.</li>\n</ul>\n<p>The other guardrail limits time rather than attempts. <code class=\"language-text\">activeDeadlineSeconds</code> caps the wall-clock lifetime of the whole Job. Once it passes, the Job is terminated and marked <code class=\"language-text\">Failed</code>, however many retries are left. Set it on anything that talks to a flaky dependency. Otherwise “retry with back-off” and “hang forever on a socket” can combine into a Job that never ends and never fails.</p>\n<h2>Cleaning up finished Jobs</h2>\n<p>This one surprises people: <strong>a Job that succeeds does not delete itself.</strong> The Job object and its completed pods stay around so you can read logs and status. Run a Job every few minutes from a CronJob and you accumulate thousands of <code class=\"language-text\">Completed</code> pods. <code class=\"language-text\">kubectl get pods</code> becomes a wall of noise, and eventually etcd and the API server feel the load.</p>\n<p>The fix is <a href=\"https://kubernetes.io/docs/concepts/workloads/controllers/ttlafterfinished/\"><code class=\"language-text\">ttlSecondsAfterFinished</code></a> (stable since Kubernetes 1.23). When it is set, a controller deletes the Job that many seconds after it finishes, whether it succeeded or failed. The Job’s pods go with it, because the Job owns them:</p>\n<div class=\"gatsby-highlight\" data-language=\"yaml\"><pre class=\"language-yaml\"><code class=\"language-yaml\"><span class=\"token key atrule\">spec</span><span class=\"token punctuation\">:</span>\n  <span class=\"token key atrule\">ttlSecondsAfterFinished</span><span class=\"token punctuation\">:</span> <span class=\"token number\">3600</span> <span class=\"token comment\"># clean up an hour after finishing</span>\n  <span class=\"token key atrule\">backoffLimit</span><span class=\"token punctuation\">:</span> <span class=\"token number\">4</span>\n  <span class=\"token key atrule\">template</span><span class=\"token punctuation\">:</span>\n    <span class=\"token comment\"># ...</span></code></pre></div>\n<p>An hour is a reasonable default. It is long enough to grab logs after something breaks, and short enough that nothing piles up.</p>\n<p>If you need the logs to outlive the pod (and for anything scheduled, you do), ship them somewhere first. This is where the <a href=\"/blog/2026-07-01-opensearch-workflow-monitoring/\">structured-logging-to-OpenSearch</a> setup earns its place. Once each run’s logs land in an index, deleting the finished pod costs you nothing, because the record already lives somewhere you can query.</p>\n<h2>CronJobs: Jobs on a schedule</h2>\n<p>A CronJob is a thin wrapper that creates a Job on a cron schedule. The spec has what you would expect, plus a few fields that matter more than they look. This one runs a report every night at 02:00 Zurich time, never overlaps runs, and limits how much history it keeps:</p>\n<div class=\"gatsby-highlight\" data-language=\"yaml\"><pre class=\"language-yaml\"><code class=\"language-yaml\"><span class=\"token key atrule\">apiVersion</span><span class=\"token punctuation\">:</span> batch/v1\n<span class=\"token key atrule\">kind</span><span class=\"token punctuation\">:</span> CronJob\n<span class=\"token key atrule\">metadata</span><span class=\"token punctuation\">:</span>\n  <span class=\"token key atrule\">name</span><span class=\"token punctuation\">:</span> nightly<span class=\"token punctuation\">-</span>report\n<span class=\"token key atrule\">spec</span><span class=\"token punctuation\">:</span>\n  <span class=\"token key atrule\">schedule</span><span class=\"token punctuation\">:</span> <span class=\"token string\">'0 2 * * *'</span> <span class=\"token comment\"># 02:00 every day</span>\n  <span class=\"token key atrule\">timeZone</span><span class=\"token punctuation\">:</span> <span class=\"token string\">'Europe/Zurich'</span> <span class=\"token comment\"># stable since 1.27; without it, runs in the controller's zone</span>\n  <span class=\"token key atrule\">concurrencyPolicy</span><span class=\"token punctuation\">:</span> Forbid\n  <span class=\"token key atrule\">startingDeadlineSeconds</span><span class=\"token punctuation\">:</span> <span class=\"token number\">300</span>\n  <span class=\"token key atrule\">successfulJobsHistoryLimit</span><span class=\"token punctuation\">:</span> <span class=\"token number\">3</span>\n  <span class=\"token key atrule\">failedJobsHistoryLimit</span><span class=\"token punctuation\">:</span> <span class=\"token number\">1</span>\n  <span class=\"token key atrule\">jobTemplate</span><span class=\"token punctuation\">:</span>\n    <span class=\"token key atrule\">spec</span><span class=\"token punctuation\">:</span>\n      <span class=\"token key atrule\">backoffLimit</span><span class=\"token punctuation\">:</span> <span class=\"token number\">2</span>\n      <span class=\"token key atrule\">ttlSecondsAfterFinished</span><span class=\"token punctuation\">:</span> <span class=\"token number\">86400</span>\n      <span class=\"token key atrule\">template</span><span class=\"token punctuation\">:</span>\n        <span class=\"token key atrule\">spec</span><span class=\"token punctuation\">:</span>\n          <span class=\"token key atrule\">restartPolicy</span><span class=\"token punctuation\">:</span> Never\n          <span class=\"token key atrule\">containers</span><span class=\"token punctuation\">:</span>\n            <span class=\"token punctuation\">-</span> <span class=\"token key atrule\">name</span><span class=\"token punctuation\">:</span> report\n              <span class=\"token key atrule\">image</span><span class=\"token punctuation\">:</span> registry.example/report<span class=\"token punctuation\">:</span><span class=\"token number\">1.4</span>\n              <span class=\"token key atrule\">command</span><span class=\"token punctuation\">:</span> <span class=\"token punctuation\">[</span><span class=\"token string\">'python'</span><span class=\"token punctuation\">,</span> <span class=\"token string\">'report.py'</span><span class=\"token punctuation\">]</span></code></pre></div>\n<p>The <code class=\"language-text\">schedule</code> is standard cron syntax. Set <code class=\"language-text\">timeZone</code> explicitly (<a href=\"https://kubernetes.io/docs/concepts/workloads/controllers/cron-job/#time-zones\">supported since Kubernetes 1.27</a>). If you leave it out, the schedule runs in whatever zone the controller manager thinks it is in. That is a great way to have a “2 AM” report fire at 2 AM UTC and confuse everyone working CERN hours.</p>\n<p>The history limits keep the last few Jobs around for inspection and garbage-collect the rest. So a CronJob cleans up after itself even without a per-Job TTL.</p>\n<h3>concurrencyPolicy: what happens when a run overruns</h3>\n<p>The field that actually decides behavior under load is <code class=\"language-text\">concurrencyPolicy</code>. It controls what the CronJob does when a new run is due but the previous one has not finished.</p>\n<p><img src=\"/cde78a0c6b5786b2ec6326153c1871e8/cronjob-concurrency.svg\" alt=\"concurrencyPolicy when a run overruns its schedule. A schedule fires at t1, t2 and t3; run A from t1 is still going when t2 arrives. Under Allow, the default, run B starts anyway and A and B overlap. Under Forbid, the t2 run is skipped and A finishes alone. Under Replace, A is killed at t2 and run B replaces it. Pick Forbid for jobs that must not run twice at once, and Replace when only the latest run matters.\"></p>\n<p>There are three options:</p>\n<ul>\n<li><strong><code class=\"language-text\">Allow</code></strong> (the default) starts a new run even if the previous one is still going. That is fine for short, independent jobs. It is a real problem for anything that takes a lock, writes to the same output, or does heavy I/O: two overlapping nightly reports can double the load or corrupt a shared file.</li>\n<li><strong><code class=\"language-text\">Forbid</code></strong> skips the new run if the previous one has not finished. Use it when a job must never run twice at once.</li>\n<li><strong><code class=\"language-text\">Replace</code></strong> kills the running one and starts fresh. Use it when only the newest run matters and a stale run in flight is just waste.</li>\n</ul>\n<h3>startingDeadlineSeconds: how late a run may start</h3>\n<p><code class=\"language-text\">startingDeadlineSeconds</code> is the other field worth setting. Sometimes the controller cannot start a scheduled run on time, because it was down or the cluster was busy. This field sets how long after the scheduled time the controller may still start the run late. Leave it unset, and a controller that misses too many scheduled times gives up in a way that surprises people. The next section explains how.</p>\n<h2>Where it breaks</h2>\n<p>The happy path is a page of YAML. The interesting engineering is in the failures.</p>\n<p><strong>Not shipping logs before the TTL deletes them.</strong> This is the most common own goal with scheduled Jobs. You set <code class=\"language-text\">ttlSecondsAfterFinished</code>, a run fails at 3 AM, and by the time you look, the pod is gone and so are its logs. TTL and log retention are the same decision. Send logs to a store on the way out, then let the pod be deleted.</p>\n<p><strong><code class=\"language-text\">concurrencyPolicy: Allow</code> plus a job that runs long.</strong> The default lets runs pile up. If run N regularly overruns into run N+1’s slot, you get overlap you never designed for. If the runs share a lock or an output path, that overlap is a correctness bug, not just extra load. Choose the policy on purpose.</p>\n<p><strong>The missed-schedule cliff.</strong> The CronJob controller counts how many scheduled times it missed since it last saw the object. If that count passes 100, it stops scheduling and logs <code class=\"language-text\">Cannot determine if job needs to be started ... too many missed start times</code>. That can happen if, say, the controller was down for a couple of hours on a <code class=\"language-text\">*/1 * * * *</code> schedule. Setting <code class=\"language-text\">startingDeadlineSeconds</code> bounds how far back the controller looks, so a short outage does not break a frequent schedule. The <a href=\"https://kubernetes.io/docs/concepts/workloads/controllers/cron-job/#cron-job-limitations\">CronJob docs call this out directly</a>, and it has surprised more than one on-call engineer.</p>\n<p><strong>Assuming exactly-once.</strong> A CronJob gives you <em>at-least-once</em> semantics, not exactly-once. Under certain conditions it can create two Jobs for one scheduled time, or none. Your job body must be idempotent (safe to run twice for the same slot), or you need your own guard, such as a row lock or an “already processed 2026-08-07” marker. Do not build a design that assumes each scheduled run happens once and only once.</p>\n<p><strong>Retrying failures that cannot succeed.</strong> I covered this above, but it is the most common waste. Without a <code class=\"language-text\">podFailurePolicy</code>, <code class=\"language-text\">backoffLimit</code> patiently retries a corrupt-input or OOM failure through every growing back-off before giving up. Branch on the exit code, and fail fast on failures that trying again cannot fix.</p>\n<h2>Tradeoffs, and when to reach for something bigger</h2>\n<p>Jobs and CronJobs are deliberately simple, and that makes them the right tool for a large slice of batch work: run this once, run this on a schedule, shard this fixed set of items. When the work is a single step or an embarrassingly parallel fan-out (independent pieces with no coordination), they are all you need. Anything heavier is over-engineering.</p>\n<p>They stop fitting when the work becomes a <em>graph</em>. A Job has no concept of “run B after A succeeds, then C and D in parallel, then E.” You can fake a small DAG (directed acyclic graph of steps) by chaining Jobs from outside the cluster. But once dependencies, fan-out/fan-in, per-step retries, and passing artifacts between steps enter the picture, you want a real workflow engine. <a href=\"https://argo-workflows.readthedocs.io/\">Argo Workflows</a> is the common Kubernetes-native choice; it runs each step as a pod and owns the ordering. It is the same lesson as ordering deploys with <a href=\"/blog/2026-07-03-argocd-sync-waves-ordering-rollout/\">ArgoCD sync waves</a>: the primitive handles one step well, and ordering many steps is a separate problem that deserves a tool built for it.</p>\n<p>In the CMS <a href=\"/project/cms-workflow-operations/\">workflow tooling</a> I worked on, the multi-stage physics pipelines run on WMCore and the grid, a purpose-built workflow system that predates all of this. But the operational glue around it is squarely Job and CronJob territory: nightly reconciliation, one-off reprocessing, scheduled health checks, and report generation. For that glue, getting the boring fields right separates batch that runs quietly from batch you babysit:</p>\n<ul>\n<li>a sane <code class=\"language-text\">backoffLimit</code>;</li>\n<li>a <code class=\"language-text\">podFailurePolicy</code> that fails fast on bad input;</li>\n<li><code class=\"language-text\">ttlSecondsAfterFinished</code> paired with shipped logs;</li>\n<li><code class=\"language-text\">concurrencyPolicy: Forbid</code> on anything that takes a lock.</li>\n</ul>\n<p>The API is small. The defaults are where the surprises live.</p>\n<hr>\n<p><em>Diagrams by M. Hassan Ahmed, released under CC0. No external image was used for this post; the figures are original work by the author.</em></p>","frontmatter":{"title":"Kubernetes Jobs and CronJobs for Batch Work","date":"2026-08-08T00:00:00.000Z","description":"Run finite and scheduled work on Kubernetes with Jobs and CronJobs: completions, parallelism, backoffLimit, cleanup, and the failure modes that bite.","thumbnail":{"childImageSharp":{"fluid":{"base64":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAABQAAAALCAIAAADwazoUAAAACXBIWXMAAAsSAAALEgHS3X78AAAB4ElEQVQozy1Q2W7bMBAUUCCVRJEUD1HUQer0IVmyXbh24gtBgT61TZAez+3//0VXjoHBYrgzyx3S4cszW5xkf+2ffzXHH7Pz6/zy2k38e/X4bXF9g2b3/LM9TRWkm/8MFeAgknk4CUU5bq+L1VM3nvr1uR9Pi/6pXe7hOOsOw/Y67w6AbjhSXoIfpgAO4hbxAuAz64d2qsx6oYGKWOHfiEcNGHDUvJPbyDTlIJBp7pMsYBbzEo4u1ojmQN4BKkhR3hfLA5EVcJ9m/s3geCTFso7yUeajyIak3pXLY8ALD4LBZJgTbokoiCipLHlcY26pKDEkpbkTqkZXOyzKgMEeWGJiu1Z2C9fDEZ5HWCplZpOsMkZr6/ri7c+/l99/P7jSifKBRjWi6WpzHj5dinbz4EmedqFqXaSInrN6L4o+6fbMDsx0ftr3m4upRjfQDlWtSFewJIxKACSESEm1m5IHsai26fYrbfZs/BLOHmnzma+f42Z4wMojieMGcajmyqxF2pOoYXoRmRGLCtYimrG4DljuIfnRZa7LXSQ9FgdJgaSBqx0Pa/gSoVtdrJVZyaxjcTMN4BikKTk3oarha0JVAZiqA5JBHwywWflEQ3VRNCGIoOsCoBNMxLtD3wmY76r6Dy2ARUHKEGsOAAAAAElFTkSuQmCC","aspectRatio":1.899441340782123,"src":"/static/cc4e4f17fbc960d6928d553d8582c8ce/40a76/hero.png","srcSet":"/static/cc4e4f17fbc960d6928d553d8582c8ce/c972b/hero.png 340w,\n/static/cc4e4f17fbc960d6928d553d8582c8ce/27625/hero.png 680w,\n/static/cc4e4f17fbc960d6928d553d8582c8ce/40a76/hero.png 1360w,\n/static/cc4e4f17fbc960d6928d553d8582c8ce/ed396/hero.png 2000w","sizes":"(max-width: 1360px) 100vw, 1360px"}}}}}},"pageContext":{"slug":"/2026-08-08-kubernetes-jobs-cronjobs-batch-workflows/","previous":"blog/2026-08-09-byte-pair-encoding-llm-tokenization/","next":"blog/2026-08-12-matryoshka-embeddings-rag/"}},"staticQueryHashes":["32046230"]}