Zero-Downtime Reindexing in OpenSearch
Change an OpenSearch mapping without dropping writes or serving stale data. A step-by-step reindex with aliases, the Reindex API, and the failure modes.
You need to change the mapping of an index that is live. (The mapping is the index’s schema: each field’s type and how its text is analyzed.) Maybe a field you typed as text should have been a keyword so you can group on it. Maybe an analyzer needs to change, or the index has too few shards for how much it has grown. Whatever the reason, the index is taking writes right now, something is querying it right now, and you cannot stop either one.
The trap is that OpenSearch does not let you make most mapping changes in place. You can add a new field, but you cannot change an existing field’s type or analyzer, and you cannot reduce the shard count. The field types docs are blunt about it: once a field is mapped, that mapping is fixed for the life of the index.
The reason is physical. The values are already on disk, in the inverted index and doc-values structures that the original type produced. There is no cheap way to reinterpret those bytes as a different type. So the only real move is to build a new index with the mapping you want, copy the data across, and cut over to it.
This post is for engineers running OpenSearch or Elasticsearch in production who need to make that change without a maintenance window. It covers the alias setup that makes it possible, the four steps of the reindex, the one gap the recipe leaves open, and the failure modes I have hit. I ran the OpenSearch monitoring for the CMS workflow stack at CERN. Those event indexes never stop taking documents, so “just delete it and reload” is not on the table. Everything below maps one to one onto Elasticsearch if that is what you run.
Aliases: separate names for reading and writing
If your clients read and write an index by its literal name, you have already lost. The name is baked into every writer and reader, and you cannot change what it points at. The fix is an index alias: a virtual name that resolves to a real index. You can repoint an alias atomically, and no client notices.
The move that makes a reindex safe is running two aliases, not one:
events_read: everything that queries goes through it.events_write: everything that indexes goes through it.
Today both point at events_v1. Splitting them matters because a reindex needs the write side and the read side to switch at different times, and a single alias cannot express that. Writes cut over early, so new data starts landing in the new index immediately. Reads cut over last, once the new index has fully caught up. If you take one thing from this post, take this: the two sides move on their own schedules, and aliases are what let them.
If your indexes are still addressed by their real names, the first migration is the painful one. You point an alias at the existing index, then update every client to use the alias instead. Do that once and every future reindex is invisible to clients. It is the same reason the workflow monitoring setup writes to workflow-events-write rather than a dated index name. The alias is what lets Index State Management roll indices over (start a fresh index behind the same name) underneath the writer without telling it.
With both aliases in place, the reindex itself is four steps.
Step 1: create the new index
Create events_v2 with the mapping you actually want. The type change, the extra shards, or the new analyzer goes here. Nothing reads or writes the new index yet, so there is no rush and no risk. The request below declares every field explicitly and sets the new shard count:
PUT events_v2
{
"settings": { "number_of_shards": 6 },
"mappings": {
"properties": {
"@timestamp": { "type": "date" },
"workflow": { "type": "keyword" },
"state": { "type": "keyword" },
"exit_code": { "type": "keyword" },
"duration_s": { "type": "long" }
}
}
}Step 2: redirect writes before you copy anything
This is the step people get backwards. Repoint events_write to events_v2 first, before the reindex. From that moment, every new document lands in the new index. Reads still go to events_v1, so nothing changes on the query side. One request removes the write alias from the old index and adds it to the new one:
POST _aliases
{
"actions": [
{ "remove": { "index": "events_v1", "alias": "events_write" } },
{ "add": { "index": "events_v2", "alias": "events_write" } }
]
}The _aliases endpoint applies all its actions in a single atomic operation. There is no instant where events_write points at nothing, so no write is rejected mid-swap. Sending the remove and the add as two separate requests would open exactly that gap, so keep them in one call.
Why move writes first? Because it draws a clean line. Everything written before the swap is in events_v1 and about to be copied. Everything written after it is already in events_v2. The reindex only has to handle the fixed set of documents that existed at cutover, while new data flows past it into the destination on its own.
Step 3: copy the backlog with the Reindex API
Now copy events_v1 into events_v2 with the Reindex API, which reads documents from a source index and writes them into a destination index. The two options that matter most are op_type and conflicts:
POST _reindex?wait_for_completion=false
{
"conflicts": "proceed",
"source": { "index": "events_v1", "size": 2000 },
"dest": { "index": "events_v2", "op_type": "create" },
"slices": "auto"
}op_type: create tells the reindex to write a document only if its _id does not already exist in the destination. That is the safety catch for the overlap window. Suppose a document was updated after the write cutover: its fresher version already sits in events_v2. Without create, the reindex would copy the stale version from events_v1 over it, and you would lose the update. With create, the reindex tries to write, sees the id exists, and steps aside.
The catch is that OpenSearch counts a rejected create as a version conflict, and by default one conflict aborts the entire job. conflicts: "proceed" reverses that default. It says a conflict here is expected and fine, so keep going and just count it. Together, these two settings make the copy safe to run while writes are live.
This only works if your document _id is deterministic, meaning it is derived from the data rather than auto-generated. If OpenSearch assigns a random _id on write, the same logical record gets one id in events_v1 and a different one from the redirected write. The reindex then happily creates a duplicate instead of detecting a conflict. Deterministic ids are a prerequisite for a clean overlapping cutover, not a nice-to-have. If you are stuck with auto-generated ids, you have to fall back to briefly pausing writes, which is a different and less pleasant plan.
The other two options control how the job runs. wait_for_completion=false returns a task id straight away, instead of holding the HTTP connection open for what might be an hour. You then watch the job through the Tasks API:
GET _tasks/<task_id>slices: "auto" splits the copy into parallel sub-tasks, roughly one per source shard. That is the difference between a reindex that finishes in minutes and one that crawls. If the copy loads the cluster too hard, cap it with requests_per_second rather than letting it starve live traffic. A reindex that browns out your production queries is not zero-downtime in any way that matters.
Step 4: swap reads and drop the old index
When the task reports done and the destination’s document count looks right, move the read alias. It is the same atomic pattern as step 2, applied to events_read:
POST _aliases
{
"actions": [
{ "remove": { "index": "events_v1", "alias": "events_read" } },
{ "add": { "index": "events_v2", "alias": "events_read" } }
]
}Queries now resolve to events_v2. The reindex and the cutover are done, and nothing ever returned an error. Keep events_v1 for a day as a rollback path: if the new mapping turns out wrong, flipping the read alias back is instant. Then delete it to reclaim the disk.
The tradeoff: reads are briefly stale
This recipe has no downtime and loses no writes, but it is not free. Between the write swap in step 2 and the read swap in step 4, readers still point at events_v1, so they cannot see documents written to events_v2. Reads are briefly stale for the newest records. They are not wrong for old records, and nothing is missing for good. Everything reconciles the instant the read alias moves.
For most systems that window is fine. In a monitoring index, a state change that shows up a few minutes late is not a problem. Where it is a problem, such as a read-your-own-writes flow where a user must see the record they just created, you have two choices:
- Keep the reindex window short, so the staleness is measured in minutes.
- Dual-write: have the application write to both indices for the duration, so both stay current. This removes the staleness, but it costs write throughput plus some code you will delete a day later.
Know which side of that tradeoff your data is on before you start.
Failure modes I have hit
Forgetting op_type: create. The default is index, which overwrites. Run a plain reindex over an index that is taking live writes and you will clobber fresh documents with stale ones, without ever seeing an error. This is the single most expensive mistake here, and it is silent.
Auto-generated ids. Covered above, but it bears repeating, because it is easy to miss until you are staring at double the document count. No stable id, no safe overlap.
Field name collisions from a wrong dynamic mapping. With dynamic mapping on, OpenSearch guesses a type for any field you did not declare. If events_v2 still has it on and a document arrives with an undeclared field, the guess can differ from the one events_v1 made, and now the two indexes disagree. Declare the mapping explicitly. On objects you control, set "dynamic": "strict" so an unexpected field is a loud rejection instead of a quiet guess.
Reindex starving live traffic. A big unthrottled reindex with slices: auto will use all the I/O it can get. If your query latency spikes during the copy, you have technically caused an outage while trying to avoid one. Throttle with requests_per_second, and run the copy during a quieter window if you have one.
Confirming the copy by task status alone. A finished task means the job ran, not that every document made it. Check the counts, and account for the conflicts you expected: docs in v1 should roughly equal created in v2 plus version conflicts. If the numbers are far off, stop and find out why before you swap reads.
What I would do differently
I wish I had put every index behind read and write aliases on day one, before there was ever a reason to reindex. Retrofitting aliases onto clients that use literal index names is the annoying part. You are editing writers and readers under load, and you pay that overhead exactly when you are already trying to fix something else. Set up the alias indirection while the index is boring and empty, and the reindex itself becomes four API calls and a wait.
I would also write down the pre-flight checklist and actually follow it:
- Deterministic ids confirmed.
- New mapping reviewed.
op_type: createandconflicts: proceedset.- Throttle chosen.
- Rollback plan (keep the old index) agreed.
Every item on that list maps to a specific way I have watched a reindex go wrong.
For the other half of the picture, the OpenSearch dashboards post covers how these event indexes get built, mapped, and alerted on in the first place. The CMS Workflow Operations writeup describes the larger system this monitoring sits inside. The same operational data feeds Archi, the retrieval copilot I built for that team, which searches these records rather than charting them. A copilot that answers questions off a stale or half-reindexed index is worse than useless, which is a large part of why getting the reindex right matters.
Diagrams by M. Hassan Ahmed, created for this post and released under CC0 (public domain). No external image was used; the figures are original work by the author.