{"componentChunkName":"component---src-templates-blog-post-js","path":"/blog/2026-08-13-zero-downtime-reindexing-opensearch/","result":{"data":{"site":{"siteMetadata":{"title":"M.Hassan Ahmed","author":"Hassan11196"}},"markdownRemark":{"id":"4c3b6c41-28b2-5a2c-a3eb-f416cc17970b","excerpt":"You need to change the mapping of an index that is live. (The mapping is the index’s schema: each field’s type and how its text is analyzed.) Maybe a field you…","html":"<p>You need to change the mapping of an index that is live. (The mapping is the index’s schema: each field’s type and how its text is analyzed.) Maybe a field you typed as <code class=\"language-text\">text</code> should have been a <code class=\"language-text\">keyword</code> so you can group on it. Maybe an analyzer needs to change, or the index has too few shards for how much it has grown. Whatever the reason, the index is taking writes right now, something is querying it right now, and you cannot stop either one.</p>\n<p>The trap is that OpenSearch does not let you make most mapping changes in place. You can add a new field, but you cannot change an existing field’s type or analyzer, and you cannot reduce the shard count. The <a href=\"https://opensearch.org/docs/latest/field-types/\">field types docs</a> are blunt about it: once a field is mapped, that mapping is fixed for the life of the index.</p>\n<p>The reason is physical. The values are already on disk, in the <a href=\"https://opensearch.org/docs/latest/field-types/mapping-parameters/doc-values/\">inverted index and doc-values</a> structures that the original type produced. There is no cheap way to reinterpret those bytes as a different type. So the only real move is to build a new index with the mapping you want, copy the data across, and cut over to it.</p>\n<p>This post is for engineers running OpenSearch or Elasticsearch in production who need to make that change without a maintenance window. It covers the alias setup that makes it possible, the four steps of the reindex, the one gap the recipe leaves open, and the failure modes I have hit. I ran the <a href=\"/blog/2026-07-01-opensearch-workflow-monitoring/\">OpenSearch monitoring</a> for the CMS workflow stack at CERN. Those event indexes never stop taking documents, so “just delete it and reload” is not on the table. Everything below maps one to one onto Elasticsearch if that is what you run.</p>\n<h2>Aliases: separate names for reading and writing</h2>\n<p>If your clients read and write an index by its literal name, you have already lost. The name is baked into every writer and reader, and you cannot change what it points at. The fix is an <a href=\"https://opensearch.org/docs/latest/im-plugin/index-alias/\">index alias</a>: a virtual name that resolves to a real index. You can repoint an alias atomically, and no client notices.</p>\n<p>The move that makes a reindex safe is running two aliases, not one:</p>\n<ul>\n<li><strong><code class=\"language-text\">events_read</code></strong>: everything that queries goes through it.</li>\n<li><strong><code class=\"language-text\">events_write</code></strong>: everything that indexes goes through it.</li>\n</ul>\n<p>Today both point at <code class=\"language-text\">events_v1</code>. Splitting them matters because a reindex needs the write side and the read side to switch at <em>different</em> times, and a single alias cannot express that. Writes cut over early, so new data starts landing in the new index immediately. Reads cut over last, once the new index has fully caught up. If you take one thing from this post, take this: the two sides move on their own schedules, and aliases are what let them.</p>\n<p><img src=\"/3d4ba3458c6614b0b2396973357e75f8/reindex-phases.svg\" alt=\"Four phases of a zero-downtime reindex. In steady state the read and write aliases both point at events_v1. Next, events_v2 is created with the new mapping and the write alias is repointed to it so new documents land in v2 while reads still come from v1. Then the Reindex API copies v1 into v2 using op_type create. Finally the read alias is swapped to v2 and v1 is deleted. The writer and reader never stop.\"></p>\n<p>If your indexes are still addressed by their real names, the first migration is the painful one. You point an alias at the existing index, then update every client to use the alias instead. Do that once and every future reindex is invisible to clients. It is the same reason the <a href=\"/blog/2026-07-01-opensearch-workflow-monitoring/\">workflow monitoring setup</a> writes to <code class=\"language-text\">workflow-events-write</code> rather than a dated index name. The alias is what lets <a href=\"https://opensearch.org/docs/latest/im-plugin/ism/index/\">Index State Management</a> roll indices over (start a fresh index behind the same name) underneath the writer without telling it.</p>\n<p>With both aliases in place, the reindex itself is four steps.</p>\n<h2>Step 1: create the new index</h2>\n<p>Create <code class=\"language-text\">events_v2</code> with the mapping you actually want. The type change, the extra shards, or the new analyzer goes here. Nothing reads or writes the new index yet, so there is no rush and no risk. The request below declares every field explicitly and sets the new shard count:</p>\n<div class=\"gatsby-highlight\" data-language=\"json\"><pre class=\"language-json\"><code class=\"language-json\">PUT events_v<span class=\"token number\">2</span>\n<span class=\"token punctuation\">{</span>\n  <span class=\"token property\">\"settings\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">{</span> <span class=\"token property\">\"number_of_shards\"</span><span class=\"token operator\">:</span> <span class=\"token number\">6</span> <span class=\"token punctuation\">}</span><span class=\"token punctuation\">,</span>\n  <span class=\"token property\">\"mappings\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">{</span>\n    <span class=\"token property\">\"properties\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">{</span>\n      <span class=\"token property\">\"@timestamp\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">{</span> <span class=\"token property\">\"type\"</span><span class=\"token operator\">:</span> <span class=\"token string\">\"date\"</span> <span class=\"token punctuation\">}</span><span class=\"token punctuation\">,</span>\n      <span class=\"token property\">\"workflow\"</span><span class=\"token operator\">:</span>   <span class=\"token punctuation\">{</span> <span class=\"token property\">\"type\"</span><span class=\"token operator\">:</span> <span class=\"token string\">\"keyword\"</span> <span class=\"token punctuation\">}</span><span class=\"token punctuation\">,</span>\n      <span class=\"token property\">\"state\"</span><span class=\"token operator\">:</span>      <span class=\"token punctuation\">{</span> <span class=\"token property\">\"type\"</span><span class=\"token operator\">:</span> <span class=\"token string\">\"keyword\"</span> <span class=\"token punctuation\">}</span><span class=\"token punctuation\">,</span>\n      <span class=\"token property\">\"exit_code\"</span><span class=\"token operator\">:</span>  <span class=\"token punctuation\">{</span> <span class=\"token property\">\"type\"</span><span class=\"token operator\">:</span> <span class=\"token string\">\"keyword\"</span> <span class=\"token punctuation\">}</span><span class=\"token punctuation\">,</span>\n      <span class=\"token property\">\"duration_s\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">{</span> <span class=\"token property\">\"type\"</span><span class=\"token operator\">:</span> <span class=\"token string\">\"long\"</span> <span class=\"token punctuation\">}</span>\n    <span class=\"token punctuation\">}</span>\n  <span class=\"token punctuation\">}</span>\n<span class=\"token punctuation\">}</span></code></pre></div>\n<h2>Step 2: redirect writes before you copy anything</h2>\n<p>This is the step people get backwards. Repoint <code class=\"language-text\">events_write</code> to <code class=\"language-text\">events_v2</code> <em>first</em>, before the reindex. From that moment, every new document lands in the new index. Reads still go to <code class=\"language-text\">events_v1</code>, so nothing changes on the query side. One request removes the write alias from the old index and adds it to the new one:</p>\n<div class=\"gatsby-highlight\" data-language=\"json\"><pre class=\"language-json\"><code class=\"language-json\">POST _aliases\n<span class=\"token punctuation\">{</span>\n  <span class=\"token property\">\"actions\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">[</span>\n    <span class=\"token punctuation\">{</span> <span class=\"token property\">\"remove\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">{</span> <span class=\"token property\">\"index\"</span><span class=\"token operator\">:</span> <span class=\"token string\">\"events_v1\"</span><span class=\"token punctuation\">,</span> <span class=\"token property\">\"alias\"</span><span class=\"token operator\">:</span> <span class=\"token string\">\"events_write\"</span> <span class=\"token punctuation\">}</span> <span class=\"token punctuation\">}</span><span class=\"token punctuation\">,</span>\n    <span class=\"token punctuation\">{</span> <span class=\"token property\">\"add\"</span><span class=\"token operator\">:</span>    <span class=\"token punctuation\">{</span> <span class=\"token property\">\"index\"</span><span class=\"token operator\">:</span> <span class=\"token string\">\"events_v2\"</span><span class=\"token punctuation\">,</span> <span class=\"token property\">\"alias\"</span><span class=\"token operator\">:</span> <span class=\"token string\">\"events_write\"</span> <span class=\"token punctuation\">}</span> <span class=\"token punctuation\">}</span>\n  <span class=\"token punctuation\">]</span>\n<span class=\"token punctuation\">}</span></code></pre></div>\n<p>The <code class=\"language-text\">_aliases</code> endpoint applies all its actions in a <a href=\"https://opensearch.org/docs/latest/im-plugin/index-alias/#create-aliases\">single atomic operation</a>. There is no instant where <code class=\"language-text\">events_write</code> points at nothing, so no write is rejected mid-swap. Sending the remove and the add as two separate requests would open exactly that gap, so keep them in one call.</p>\n<p>Why move writes first? Because it draws a clean line. Everything written before the swap is in <code class=\"language-text\">events_v1</code> and about to be copied. Everything written after it is already in <code class=\"language-text\">events_v2</code>. The reindex only has to handle the fixed set of documents that existed at cutover, while new data flows past it into the destination on its own.</p>\n<h2>Step 3: copy the backlog with the Reindex API</h2>\n<p>Now copy <code class=\"language-text\">events_v1</code> into <code class=\"language-text\">events_v2</code> with the <a href=\"https://opensearch.org/docs/latest/api-reference/document-apis/reindex/\">Reindex API</a>, which reads documents from a source index and writes them into a destination index. The two options that matter most are <code class=\"language-text\">op_type</code> and <code class=\"language-text\">conflicts</code>:</p>\n<div class=\"gatsby-highlight\" data-language=\"json\"><pre class=\"language-json\"><code class=\"language-json\">POST _reindex?wait_for_completion=<span class=\"token boolean\">false</span>\n<span class=\"token punctuation\">{</span>\n  <span class=\"token property\">\"conflicts\"</span><span class=\"token operator\">:</span> <span class=\"token string\">\"proceed\"</span><span class=\"token punctuation\">,</span>\n  <span class=\"token property\">\"source\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">{</span> <span class=\"token property\">\"index\"</span><span class=\"token operator\">:</span> <span class=\"token string\">\"events_v1\"</span><span class=\"token punctuation\">,</span> <span class=\"token property\">\"size\"</span><span class=\"token operator\">:</span> <span class=\"token number\">2000</span> <span class=\"token punctuation\">}</span><span class=\"token punctuation\">,</span>\n  <span class=\"token property\">\"dest\"</span><span class=\"token operator\">:</span>   <span class=\"token punctuation\">{</span> <span class=\"token property\">\"index\"</span><span class=\"token operator\">:</span> <span class=\"token string\">\"events_v2\"</span><span class=\"token punctuation\">,</span> <span class=\"token property\">\"op_type\"</span><span class=\"token operator\">:</span> <span class=\"token string\">\"create\"</span> <span class=\"token punctuation\">}</span><span class=\"token punctuation\">,</span>\n  <span class=\"token property\">\"slices\"</span><span class=\"token operator\">:</span> <span class=\"token string\">\"auto\"</span>\n<span class=\"token punctuation\">}</span></code></pre></div>\n<p><code class=\"language-text\">op_type: create</code> tells the reindex to write a document only if its <code class=\"language-text\">_id</code> does not already exist in the destination. That is the safety catch for the overlap window. Suppose a document was updated after the write cutover: its fresher version already sits in <code class=\"language-text\">events_v2</code>. Without <code class=\"language-text\">create</code>, the reindex would copy the stale version from <code class=\"language-text\">events_v1</code> over it, and you would lose the update. With <code class=\"language-text\">create</code>, the reindex tries to write, sees the id exists, and steps aside.</p>\n<p>The catch is that OpenSearch counts a rejected create as a version conflict, and by default one conflict aborts the entire job. <code class=\"language-text\">conflicts: &quot;proceed&quot;</code> reverses that default. It says a conflict here is expected and fine, so keep going and just count it. Together, these two settings make the copy safe to run while writes are live.</p>\n<p><img src=\"/b78c0d80b254385c2cb015fe0aeef887/reindex-conflicts.svg\" alt=\"During the overlap window each document _id takes one of two paths. A document that only exists in events_v1 is created in events_v2 by the reindex. A document that was already rewritten to events_v2 by a live write causes a create conflict, so the reindex skips it and the fresh version survives. Setting conflicts to proceed means those expected conflicts do not abort the job.\"></p>\n<p>This only works if your document <code class=\"language-text\">_id</code> is deterministic, meaning it is derived from the data rather than auto-generated. If OpenSearch assigns a random <code class=\"language-text\">_id</code> on write, the same logical record gets one id in <code class=\"language-text\">events_v1</code> and a different one from the redirected write. The reindex then happily creates a duplicate instead of detecting a conflict. Deterministic ids are a prerequisite for a clean overlapping cutover, not a nice-to-have. If you are stuck with auto-generated ids, you have to fall back to briefly pausing writes, which is a different and less pleasant plan.</p>\n<p>The other two options control how the job runs. <code class=\"language-text\">wait_for_completion=false</code> returns a task id straight away, instead of holding the HTTP connection open for what might be an hour. You then watch the job through the <a href=\"https://opensearch.org/docs/latest/api-reference/tasks/\">Tasks API</a>:</p>\n<div class=\"gatsby-highlight\" data-language=\"json\"><pre class=\"language-json\"><code class=\"language-json\">GET _tasks/&lt;task_id></code></pre></div>\n<p><code class=\"language-text\">slices: &quot;auto&quot;</code> splits the copy into parallel sub-tasks, roughly one per source shard. That is the difference between a reindex that finishes in minutes and one that crawls. If the copy loads the cluster too hard, cap it with <code class=\"language-text\">requests_per_second</code> rather than letting it starve live traffic. A reindex that browns out your production queries is not zero-downtime in any way that matters.</p>\n<h2>Step 4: swap reads and drop the old index</h2>\n<p>When the task reports done and the destination’s document count looks right, move the read alias. It is the same atomic pattern as step 2, applied to <code class=\"language-text\">events_read</code>:</p>\n<div class=\"gatsby-highlight\" data-language=\"json\"><pre class=\"language-json\"><code class=\"language-json\">POST _aliases\n<span class=\"token punctuation\">{</span>\n  <span class=\"token property\">\"actions\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">[</span>\n    <span class=\"token punctuation\">{</span> <span class=\"token property\">\"remove\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">{</span> <span class=\"token property\">\"index\"</span><span class=\"token operator\">:</span> <span class=\"token string\">\"events_v1\"</span><span class=\"token punctuation\">,</span> <span class=\"token property\">\"alias\"</span><span class=\"token operator\">:</span> <span class=\"token string\">\"events_read\"</span> <span class=\"token punctuation\">}</span> <span class=\"token punctuation\">}</span><span class=\"token punctuation\">,</span>\n    <span class=\"token punctuation\">{</span> <span class=\"token property\">\"add\"</span><span class=\"token operator\">:</span>    <span class=\"token punctuation\">{</span> <span class=\"token property\">\"index\"</span><span class=\"token operator\">:</span> <span class=\"token string\">\"events_v2\"</span><span class=\"token punctuation\">,</span> <span class=\"token property\">\"alias\"</span><span class=\"token operator\">:</span> <span class=\"token string\">\"events_read\"</span> <span class=\"token punctuation\">}</span> <span class=\"token punctuation\">}</span>\n  <span class=\"token punctuation\">]</span>\n<span class=\"token punctuation\">}</span></code></pre></div>\n<p>Queries now resolve to <code class=\"language-text\">events_v2</code>. The reindex and the cutover are done, and nothing ever returned an error. Keep <code class=\"language-text\">events_v1</code> for a day as a rollback path: if the new mapping turns out wrong, flipping the read alias back is instant. Then delete it to reclaim the disk.</p>\n<h2>The tradeoff: reads are briefly stale</h2>\n<p>This recipe has no downtime and loses no writes, but it is not free. Between the write swap in step 2 and the read swap in step 4, readers still point at <code class=\"language-text\">events_v1</code>, so they cannot see documents written to <code class=\"language-text\">events_v2</code>. Reads are briefly stale for the newest records. They are not wrong for old records, and nothing is missing for good. Everything reconciles the instant the read alias moves.</p>\n<p>For most systems that window is fine. In a monitoring index, a state change that shows up a few minutes late is not a problem. Where it <em>is</em> a problem, such as a read-your-own-writes flow where a user must see the record they just created, you have two choices:</p>\n<ul>\n<li><strong>Keep the reindex window short</strong>, so the staleness is measured in minutes.</li>\n<li><strong>Dual-write</strong>: have the application write to both indices for the duration, so both stay current. This removes the staleness, but it costs write throughput plus some code you will delete a day later.</li>\n</ul>\n<p>Know which side of that tradeoff your data is on before you start.</p>\n<h2>Failure modes I have hit</h2>\n<p><strong>Forgetting <code class=\"language-text\">op_type: create</code>.</strong> The default is <code class=\"language-text\">index</code>, which overwrites. Run a plain reindex over an index that is taking live writes and you will clobber fresh documents with stale ones, without ever seeing an error. This is the single most expensive mistake here, and it is silent.</p>\n<p><strong>Auto-generated ids.</strong> Covered above, but it bears repeating, because it is easy to miss until you are staring at double the document count. No stable id, no safe overlap.</p>\n<p><strong>Field name collisions from a wrong dynamic mapping.</strong> With dynamic mapping on, OpenSearch guesses a type for any field you did not declare. If <code class=\"language-text\">events_v2</code> still has it on and a document arrives with an undeclared field, the guess can differ from the one <code class=\"language-text\">events_v1</code> made, and now the two indexes disagree. Declare the mapping explicitly. On objects you control, set <a href=\"https://opensearch.org/docs/latest/field-types/#dynamic-mapping\"><code class=\"language-text\">&quot;dynamic&quot;: &quot;strict&quot;</code></a> so an unexpected field is a loud rejection instead of a quiet guess.</p>\n<p><strong>Reindex starving live traffic.</strong> A big unthrottled reindex with <code class=\"language-text\">slices: auto</code> will use all the I/O it can get. If your query latency spikes during the copy, you have technically caused an outage while trying to avoid one. Throttle with <code class=\"language-text\">requests_per_second</code>, and run the copy during a quieter window if you have one.</p>\n<p><strong>Confirming the copy by task status alone.</strong> A finished task means the job ran, not that every document made it. Check the counts, and account for the conflicts you expected: <code class=\"language-text\">docs in v1</code> should roughly equal <code class=\"language-text\">created in v2</code> plus <code class=\"language-text\">version conflicts</code>. If the numbers are far off, stop and find out why before you swap reads.</p>\n<h2>What I would do differently</h2>\n<p>I wish I had put every index behind read and write aliases on day one, before there was ever a reason to reindex. Retrofitting aliases onto clients that use literal index names is the annoying part. You are editing writers and readers under load, and you pay that overhead exactly when you are already trying to fix something else. Set up the alias indirection while the index is boring and empty, and the reindex itself becomes four API calls and a wait.</p>\n<p>I would also write down the pre-flight checklist and actually follow it:</p>\n<ul>\n<li>Deterministic ids confirmed.</li>\n<li>New mapping reviewed.</li>\n<li><code class=\"language-text\">op_type: create</code> and <code class=\"language-text\">conflicts: proceed</code> set.</li>\n<li>Throttle chosen.</li>\n<li>Rollback plan (keep the old index) agreed.</li>\n</ul>\n<p>Every item on that list maps to a specific way I have watched a reindex go wrong.</p>\n<p>For the other half of the picture, the <a href=\"/blog/2026-07-01-opensearch-workflow-monitoring/\">OpenSearch dashboards post</a> covers how these event indexes get built, mapped, and alerted on in the first place. The <a href=\"/project/cms-workflow-operations/\">CMS Workflow Operations</a> writeup describes the larger system this monitoring sits inside. The same operational data feeds <a href=\"/project/archi/\">Archi</a>, the retrieval copilot I built for that team, which searches these records rather than charting them. A copilot that answers questions off a stale or half-reindexed index is worse than useless, which is a large part of why getting the reindex right matters.</p>\n<hr>\n<p><em>Diagrams by M. Hassan Ahmed, created for this post and released under CC0 (public domain). No external image was used; the figures are original work by the author.</em></p>","frontmatter":{"title":"Zero-Downtime Reindexing in OpenSearch","date":"2026-08-13T00:00:00.000Z","description":"Change an OpenSearch mapping without dropping writes or serving stale data. A step-by-step reindex with aliases, the Reindex API, and the failure modes.","thumbnail":{"childImageSharp":{"fluid":{"base64":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAABQAAAALCAIAAADwazoUAAAACXBIWXMAAAsSAAALEgHS3X78AAAB9UlEQVQoz03RWYvbMBAHcL80yUa2ZEuyDluWLB+Jc2dzbZLd5uim2y196lMLhX7/j1E5hVL4MQzD/IVgvIkp/2lSc17trtv9x8fNp+3hvN4d56vrZu8m++lilNmxKf7f97pIOx2Y+XEN+VA3+3pxspNjPT9VsxfTPA2XFzcxzT5gQ59WHZS1+/eU50f6LvOxAcQirCNiEDUB1g7mFaQW0TzABhJX75utNnUP49zHNiS5iLWVlouKisoFQJSnxZIkDZGNW2gRG7QKEGkQZh5itRocud0Wg833t28/bu+Hp9dm+2W2e5/u3pv154nrn77moyMSE6IWOJ1jNYdx1Yep5x5AcQlpgVgZpePMjN8WMyqbOBnyZJCbMU+GVNSzzZWpUQ8qEBn3BRBlfZR6D1A+BKLnC4SVtSOuhuuqzllKw2hRqENTNFowFCJeT9fXOBul5RJg0wtkHyYeQBIi5qCQo1AAxBWT10G+1PJ5oA+1fq6zpZG5ULeXy8hWA22tVD4UPSi9GMerKlkUiYmRptBQmJDgMcOvFTsV9FzG55JeXG/J21j9vu1/Xbc/z6uEyQ7gHoCcRBxjHoYxJoLGKaUJwjKIBGxJiGXbYOkj7oe8nUTSpR4C7nV91gGtD4D1wzQg7QGDKIP3OzuQaNhW44eqC9hfPb/1B023S6Sztnk7AAAAAElFTkSuQmCC","aspectRatio":1.899441340782123,"src":"/static/b78a3a6b876e8bd49503a2378a13233e/40a76/hero.png","srcSet":"/static/b78a3a6b876e8bd49503a2378a13233e/c972b/hero.png 340w,\n/static/b78a3a6b876e8bd49503a2378a13233e/27625/hero.png 680w,\n/static/b78a3a6b876e8bd49503a2378a13233e/40a76/hero.png 1360w,\n/static/b78a3a6b876e8bd49503a2378a13233e/ed396/hero.png 2000w","sizes":"(max-width: 1360px) 100vw, 1360px"}}}}}},"pageContext":{"slug":"/2026-08-13-zero-downtime-reindexing-opensearch/","previous":"blog/2026-08-14-fastapi-token-bucket-rate-limiting/","next":"blog/2026-08-17-semantic-caching-llm-apps/"}},"staticQueryHashes":["32046230"]}