{"componentChunkName":"component---src-templates-blog-post-js","path":"/blog/2026-07-03-argocd-sync-waves-ordering-rollout/","result":{"data":{"site":{"siteMetadata":{"title":"M.Hassan Ahmed","author":"Hassan11196"}},"markdownRemark":{"id":"ef6bfe12-ca66-5892-9ebf-7f276bba0df5","excerpt":"The first time a database migration ran at the same time as the pods that depended on it, I learned what “declarative” actually costs you. ArgoCD takes the…","html":"<p>The first time a database migration ran at the same time as the pods that depended on it, I learned what “declarative” actually costs you. <a href=\"https://argo-cd.readthedocs.io/en/stable/\">ArgoCD</a> takes the manifests in your Git repo and drives the cluster toward them. That is the whole point of <a href=\"https://opengitops.dev/\">GitOps</a>: Git holds the desired state, and a controller makes the cluster match it. But by default ArgoCD does not care what order the pieces arrive in. It computes the diff between Git and the cluster and applies everything at once. Your <code class=\"language-text\">Deployment</code>, your <code class=\"language-text\">ConfigMap</code>, and your migration <code class=\"language-text\">Job</code> all ship in the same breath.</p>\n<p>Most of the time that is fine, because Kubernetes is eventually consistent and controllers retry. A pod that starts before its ConfigMap exists will crash-loop for a few seconds and then come up. But retrying cannot fix some orderings. Suppose the schema migration has not run yet. The new pods connect to an old database and fail their readiness probe, and if you are unlucky, the old pods are already gone. Now you have downtime that no retry loop can fix, because the thing that needed to happen first never happened.</p>\n<p>This post is for people running services on Kubernetes through ArgoCD who have hit an ordering problem and learned that “it usually works” is not the same as “it works.” I will cover the two mechanisms Argo gives you (sync waves and resource hooks), how they interact, and the failure modes that cost me the most time to diagnose.</p>\n<h2>What “apply everything at once” really means</h2>\n<p>An ArgoCD sync has structure even before you add anything. Every sync runs in three <a href=\"https://argo-cd.readthedocs.io/en/stable/user-guide/sync-waves/\">phases</a>: <strong>PreSync</strong>, then <strong>Sync</strong>, then <strong>PostSync</strong>. By default all of your resources live in the Sync phase, so the phases look like a single step. Inside a phase, resources are grouped into <strong>sync waves</strong>, and every resource sits in wave <code class=\"language-text\">0</code> unless you say otherwise.</p>\n<p>So the default is one phase, one wave, everything together. You can change two things: which wave a resource is in, and which phase. Those are the only knobs, and both are just annotations.</p>\n<p><img src=\"/06dcd021349d0f4493326f5209826c13/sync-waves.svg\" alt=\"One ArgoCD sync ordered into PreSync, Sync, and PostSync phases, with resources grouped into numbered waves; Argo waits for each wave to report Healthy before starting the next\"></p>\n<h2>Sync waves: a number that orders your resources</h2>\n<p>A sync wave is an integer annotation on a resource. This one puts a <code class=\"language-text\">Deployment</code> in wave 1:</p>\n<div class=\"gatsby-highlight\" data-language=\"yaml\"><pre class=\"language-yaml\"><code class=\"language-yaml\"><span class=\"token key atrule\">apiVersion</span><span class=\"token punctuation\">:</span> apps/v1\n<span class=\"token key atrule\">kind</span><span class=\"token punctuation\">:</span> Deployment\n<span class=\"token key atrule\">metadata</span><span class=\"token punctuation\">:</span>\n  <span class=\"token key atrule\">name</span><span class=\"token punctuation\">:</span> workflow<span class=\"token punctuation\">-</span>console\n  <span class=\"token key atrule\">annotations</span><span class=\"token punctuation\">:</span>\n    <span class=\"token key atrule\">argocd.argoproj.io/sync-wave</span><span class=\"token punctuation\">:</span> <span class=\"token string\">\"1\"</span></code></pre></div>\n<p>Argo processes waves from the lowest number to the highest. Negative numbers are allowed and run before wave <code class=\"language-text\">0</code>. That lets you slot work <em>ahead of</em> every resource without an annotation, without editing all the other manifests. Resources in the same wave are applied together; the ordering only exists <em>between</em> waves.</p>\n<p>Here is what makes waves useful rather than cosmetic: <strong>Argo waits for every resource in a wave to become <a href=\"https://argo-cd.readthedocs.io/en/stable/operator-manual/health/\">Healthy</a> before it starts the next wave.</strong> Health is a real assessment, not “the API accepted the manifest.” A <code class=\"language-text\">Deployment</code> is Healthy when its pods are actually up and passing readiness. A <code class=\"language-text\">Job</code> is Healthy when it has completed. So take a rollout like this:</p>\n<ul>\n<li><strong>wave 0</strong>: the <code class=\"language-text\">ConfigMap</code> and <code class=\"language-text\">Secret</code> the app reads on boot</li>\n<li><strong>wave 1</strong>: the <code class=\"language-text\">Deployment</code></li>\n<li><strong>wave 2</strong>: the <code class=\"language-text\">Service</code> and <code class=\"language-text\">Ingress</code> that route traffic to it</li>\n</ul>\n<p>The config exists and is settled before any pod starts, and traffic is wired up only after the pods report Ready. That is exactly the order you would follow by hand, except now it is declarative and happens the same way every time.</p>\n<p>One detail will surprise you if you do not know it. Argo adds a deliberate delay between waves: two seconds by default, tunable through the <code class=\"language-text\">ARGOCD_SYNC_WAVE_DELAY</code> environment variable on the application controller. The delay gives other controllers a moment to react to what just landed. On a rollout with many waves it adds up, and it is the honest answer to “why did a sync that changed nothing still take twenty seconds?”</p>\n<h2>Hooks: for one-off work like migrations</h2>\n<p>Waves order resources that are supposed to <em>stay</em> in the cluster. A migration is different. You want it to run at a specific point, succeed, and then get out of the way. That is what <a href=\"https://argo-cd.readthedocs.io/en/stable/user-guide/resource_hooks/\">resource hooks</a> are for. A hook annotation moves a resource into the PreSync or PostSync phase. Here, the migration <code class=\"language-text\">Job</code> runs as a PreSync hook and is cleaned up once it succeeds:</p>\n<div class=\"gatsby-highlight\" data-language=\"yaml\"><pre class=\"language-yaml\"><code class=\"language-yaml\"><span class=\"token key atrule\">apiVersion</span><span class=\"token punctuation\">:</span> batch/v1\n<span class=\"token key atrule\">kind</span><span class=\"token punctuation\">:</span> Job\n<span class=\"token key atrule\">metadata</span><span class=\"token punctuation\">:</span>\n  <span class=\"token key atrule\">name</span><span class=\"token punctuation\">:</span> db<span class=\"token punctuation\">-</span>migrate\n  <span class=\"token key atrule\">annotations</span><span class=\"token punctuation\">:</span>\n    <span class=\"token key atrule\">argocd.argoproj.io/hook</span><span class=\"token punctuation\">:</span> PreSync\n    <span class=\"token key atrule\">argocd.argoproj.io/hook-delete-policy</span><span class=\"token punctuation\">:</span> HookSucceeded\n<span class=\"token key atrule\">spec</span><span class=\"token punctuation\">:</span>\n  <span class=\"token key atrule\">template</span><span class=\"token punctuation\">:</span>\n    <span class=\"token key atrule\">spec</span><span class=\"token punctuation\">:</span>\n      <span class=\"token key atrule\">restartPolicy</span><span class=\"token punctuation\">:</span> Never\n      <span class=\"token key atrule\">containers</span><span class=\"token punctuation\">:</span>\n        <span class=\"token punctuation\">-</span> <span class=\"token key atrule\">name</span><span class=\"token punctuation\">:</span> migrate\n          <span class=\"token key atrule\">image</span><span class=\"token punctuation\">:</span> registry.internal/workflow<span class=\"token punctuation\">-</span>console<span class=\"token punctuation\">:</span><span class=\"token punctuation\">{</span><span class=\"token punctuation\">{</span> .Values.tag <span class=\"token punctuation\">}</span><span class=\"token punctuation\">}</span>\n          <span class=\"token key atrule\">command</span><span class=\"token punctuation\">:</span> <span class=\"token punctuation\">[</span><span class=\"token string\">\"alembic\"</span><span class=\"token punctuation\">,</span> <span class=\"token string\">\"upgrade\"</span><span class=\"token punctuation\">,</span> <span class=\"token string\">\"head\"</span><span class=\"token punctuation\">]</span></code></pre></div>\n<p>A <code class=\"language-text\">PreSync</code> hook runs, and must succeed, before the Sync phase begins. So the migration finishes while the old code is still serving traffic, and then the new pods roll out into a schema that already matches them. A <code class=\"language-text\">PostSync</code> hook runs after every Sync-phase resource is Healthy. That is where a smoke test belongs: the check that hits a real endpoint and fails the sync if the new version is not actually serving.</p>\n<p><code class=\"language-text\">hook-delete-policy</code> is the part people forget. Without it, every sync leaves another completed <code class=\"language-text\">Job</code> object lying around, until someone wonders why the namespace is full of <code class=\"language-text\">db-migrate-xxxxx</code>. <a href=\"https://argo-cd.readthedocs.io/en/stable/user-guide/resource_hooks/\"><code class=\"language-text\">HookSucceeded</code></a> deletes the Job once it finishes cleanly. Just as usefully, it keeps the Job around when it fails, so you can read the logs.</p>\n<p>Waves work <em>inside</em> hook phases too. A <code class=\"language-text\">PreSync</code> hook with <code class=\"language-text\">sync-wave: &quot;-1&quot;</code> runs before a <code class=\"language-text\">PreSync</code> hook with wave <code class=\"language-text\">0</code>. That is how you order two setup steps that both need to finish before anything else starts.</p>\n<h2>Where it breaks</h2>\n<p><strong>A wave that never goes Healthy stalls the whole sync.</strong> This is the failure mode that reads as a hang. Say a wave-1 <code class=\"language-text\">Deployment</code> has a bad readiness probe, or an image that will not pull. It never becomes Healthy, so Argo never advances to wave 2, and the sync sits in Progressing until it times out. The application looks stuck for no visible reason. The reason is always in the specific resource Argo is waiting on. So the first move is to find which wave is in flight and look at that resource’s health, not the Application’s top-level status.</p>\n<p><strong>A hook with no clear success condition.</strong> ArgoCD decides a <code class=\"language-text\">Job</code> hook is done from the Job’s own completion status. If the migration container exits <code class=\"language-text\">0</code> before the migration has actually committed, Argo reads that as success and rolls the pods forward anyway. The hook is only as trustworthy as the exit code of the thing inside it. Test this deliberately: a migration that half-applies and returns zero makes for a much worse day than one that fails loudly.</p>\n<p><strong>Putting a <code class=\"language-text\">Namespace</code> or CRD in the wrong wave.</strong> A custom resource needs its <code class=\"language-text\">CustomResourceDefinition</code> (CRD) to exist before the API server will accept it. If both land in wave 0, the sync can fail on the first apply, because the CRD’s type is not registered yet. CRDs and namespaces belong in an early negative wave. This is the most common reason a fresh install fails on the first sync but works on the second, and it is easy to miss because the retry hides it.</p>\n<p><strong>Assuming waves give you cross-Application ordering. They do not.</strong> Waves order resources <em>within a single</em> <code class=\"language-text\">Application</code>. If your database and your app are separate Applications, a wave number in one says nothing about the other. That ordering problem needs a different tool, <a href=\"https://argo-cd.readthedocs.io/en/stable/user-guide/sync_windows/\">sync windows</a> or an app-of-apps structure. Reaching for waves there quietly does nothing.</p>\n<h2>What I would do differently</h2>\n<p>I over-used waves early on. After getting burned once, the instinct is to annotate everything, and you end up with a ten-wave rollout where three waves would do. Each extra wave pays the inter-wave delay, and each one is another place the sync can stall. Waves are for the orderings that genuinely cannot recover on their own:</p>\n<ul>\n<li>config before the pods that read it at boot;</li>\n<li>migrations before the code that assumes them;</li>\n<li>CRDs before their custom resources.</li>\n</ul>\n<p>For everything else, let Kubernetes retry. The controllers are good at it, and a rollout with fewer ordering constraints has fewer ways to hang.</p>\n<p>The other thing I would tell my earlier self: a stalled wave is only mysterious if you cannot see it. The health status Argo is blocking on is also worth watching over time. Pushing it into a dashboard turns “the sync is stuck” into “wave 1 has been Progressing for four minutes on a failing readiness probe.” I wrote about building that kind of view in <a href=\"/blog/2026-07-01-opensearch-workflow-monitoring/\">OpenSearch dashboards for workflow monitoring</a>.</p>\n<h2>Where this runs</h2>\n<p>This is how I roll out the <a href=\"/project/cms-workflow-operations/\">WMCore and Unified operations services</a> that schedule Monte Carlo production and data reconstruction for the CMS experiment across the Worldwide LHC Computing Grid. Moving those services onto Kubernetes and ArgoCD is what made safe updates and fast rollbacks possible in the first place. Sync waves are the reason a deploy that touches config, a schema, and a service comes up in an order I can reason about, instead of one I have to hope for. GitOps guarantees that the cluster matches Git. Sync waves and hooks let you also decide the order it gets there.</p>\n<p><em>Diagram by M. Hassan Ahmed, made for this post.</em></p>","frontmatter":{"title":"ArgoCD Sync Waves: Ordering a Kubernetes Rollout","date":"2026-07-03T00:00:00.000Z","description":"ArgoCD applies your whole app at once, so a migration races the pods that need it. How sync waves and hooks order a Kubernetes rollout, and where they stall.","thumbnail":{"childImageSharp":{"fluid":{"base64":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAABQAAAALCAIAAADwazoUAAAACXBIWXMAAAsSAAALEgHS3X78AAABvElEQVQoz4VRbW/TMBDON5bEr01iO05jJ06ctHlrt45uHUPakBAwBOL//xquRdM+IIT06HT33HN+znbAzemM8p7o92b47K+/N7tvENv9i1u+9oef3eFHNX0Zjr9y/0SLIyhfR05BRHVIFE3tSja1vx53H4b5oR/vttOpH+628wmYpr9dbj7qamKyxUn9DuuQFjEtApS2RGy5nlk+Yulj5ZBqYlHH0iHdIojyzETSceXyoilKJ1W1EjbiNsBiZMUhtScs96Qccd0i26KqxbWnmwGScwmou7x2rVkv3vS1WWsT0TLAaZOtJ2lmnm9iZrjqlF2gzMoJr2phZlFCdwE+5hVTG2EXUOLERbA2FX3p77vds6wPVHTC7v3y5MZH3dzyfGv6h3Z5roZHaXfQVfVNt/+UN0ci+ojoIGKGqQHOjrmhiUVsHXGH5YyynskOFCHJQ6xx1nG9oFUD7xSRHPjLMLc4dWmxjRMfr1oAVTNTU0iKGKSJhxchcoqYvcKapM5VjgsPBvBHQYjlFQCJEKuLj74kCuK5RXK4FywV/hFgmWQGzEMM65yH1SvkX7iQKHvLsbpCb8oAPP8H+a/Wb3tdQtPhRG8pAAAAAElFTkSuQmCC","aspectRatio":1.899441340782123,"src":"/static/072a9d5ad2239659e3f54544e74b3960/40a76/hero.png","srcSet":"/static/072a9d5ad2239659e3f54544e74b3960/c972b/hero.png 340w,\n/static/072a9d5ad2239659e3f54544e74b3960/27625/hero.png 680w,\n/static/072a9d5ad2239659e3f54544e74b3960/40a76/hero.png 1360w,\n/static/072a9d5ad2239659e3f54544e74b3960/ed396/hero.png 2000w","sizes":"(max-width: 1360px) 100vw, 1360px"}}}}}},"pageContext":{"slug":"/2026-07-03-argocd-sync-waves-ordering-rollout/","previous":"blog/2026-06-29-incremental-rag-indexing/","next":"blog/2026-07-02-hybrid-search-rag-bm25-vectors/"}},"staticQueryHashes":["32046230"]}