OpenSearch ISM: Automate Index Rollover

Log indexes grow until a data node runs out of disk. How OpenSearch ISM rolls indexes over by size, ages them through hot and warm, then deletes on schedule.

The failure mode is boring, and it always happens the same way. You point your services at an OpenSearch cluster, ingestion works, the dashboards look great, and nobody thinks about it again for four months. Then a data node crosses its disk watermark (the disk-usage threshold at which OpenSearch starts protecting itself). OpenSearch flips the whole index to read-only, and writes start failing for every service that shared that node. The dashboards you built to catch problems are now the thing that is down.

I ran the telemetry side of CMS workflow operations at CERN. There, OpenSearch holds logs and metrics from WMCore (the CMS workload management system) and a stack of databases, and the volume only goes one direction. A single fat logs index that grows forever is the setup that eventually pages someone at 2am. Index State Management (ISM) is the OpenSearch plugin that prevents this. It turns “someone remembers to clean up old indexes” into a policy the cluster enforces on its own.

This post is for engineers who already have data flowing into OpenSearch and now need it to not fall over six months from now. (On Elasticsearch, the equivalent is ILM.) I will walk through:

  • what an ISM policy actually is;
  • how rollover works, and why it needs an alias;
  • how to attach a policy automatically so future indexes inherit it;
  • the specific ways it bites you in practice.

Why one big index is the problem

An OpenSearch index is split into shards. A shard is a Lucene index (the on-disk unit of the search library underneath OpenSearch), and it lives entirely on one node. People learn two consequences of that the hard way:

  • A shard can only grow as large as the free space on its node.
  • A huge shard is slow to recover, slow to relocate, and slow to search.

The OpenSearch sizing guidance lands on tens of gigabytes per shard for exactly this reason. A 500GB shard is an operational liability whether or not the disk can technically hold it.

So you do not want one index that grows without bound. You want a series of time-boxed indexes, each capped at a sane size, where the old ones get cheaper to keep and eventually get deleted. Doing that by hand means a cron job that creates logs-2026-09, repoints your writers, and runs DELETE on whatever is older than the retention window. ISM is that cron job, except it lives inside the cluster, survives restarts, and does not depend on a script on someone’s laptop.

A policy is a state machine

An ISM policy is a JSON document with three kinds of parts:

  • States: the stages an index moves through.
  • Actions: what runs while the index is in a given state.
  • Transitions: the conditions that move the index from one state to the next.

The mental model is a state machine walking a single index forward through its life.

An ISM policy drawn as a state machine over one index. A hot state runs rollover and holds the write index, with a self-loop labelled rollover on min_primary_shard_size 50gb or min_index_age 1d. Arrows labelled min_index_age carry the index from hot to a warm state at seven days, warm to cold at thirty days, and cold to a delete state at ninety days. Warm runs force_merge, drops replica_count to zero, and reallocates to warm nodes; cold sets read_only and low priority; delete removes the index. A note says age is measured from the rollover that created the index.

Here is a policy with the common shape. It keeps new data hot and writable, and rolls over to a fresh index once the current one gets big or old. Later it ages the index into a cheaper warm state, then deletes it after ninety days.

PUT _plugins/_ism/policies/logs_policy
{
  "policy": {
    "description": "Rollover, age, and expire log indexes",
    "default_state": "hot",
    "states": [
      {
        "name": "hot",
        "actions": [
          { "rollover": { "min_primary_shard_size": "50gb", "min_index_age": "1d" } }
        ],
        "transitions": [
          { "state_name": "warm", "conditions": { "min_index_age": "7d" } }
        ]
      },
      {
        "name": "warm",
        "actions": [
          { "replica_count": { "number_of_replicas": 1 } },
          { "force_merge": { "max_num_segments": 1 } }
        ],
        "transitions": [
          { "state_name": "delete", "conditions": { "min_index_age": "90d" } }
        ]
      },
      {
        "name": "delete",
        "actions": [ { "delete": {} } ]
      }
    ],
    "ism_template": [
      { "index_patterns": ["logs-*"], "priority": 100 }
    ]
  }
}

A few details in there deserve a close read. default_state is where a freshly attached index starts. Each rollover condition is a floor, not a target:

  • min_primary_shard_size: "50gb" means roll over once the largest primary shard reaches 50GB.
  • min_index_age: "1d", sitting next to it, means roll over after a day even if the index never fills up.

Combining the two is deliberate. Size-based rollover keeps shards bounded under heavy traffic. Age-based rollover keeps a low-traffic index from staying open for weeks. Whichever condition trips first wins.

The force_merge in the warm state is a real tradeoff, not a free win. It merges the index down to one segment (a segment is one of the immutable files a Lucene index is built from). That makes searches over old data faster and reclaims space from deleted documents. But it is CPU and IO heavy, and the index should be done taking writes before you run it. That is why it lives in warm, after rollover has already moved writes elsewhere.

Rollover needs an alias and a numbered index name

Rollover is the action people get wrong first. It has two requirements that are easy to miss, and it fails quietly when they are absent.

The write target has to be an alias, not an index. Your services write to logs, which is an alias pointing at exactly one real index. When rollover fires, OpenSearch creates the next index and repoints the alias at it. Writers never learn that anything happened; they keep sending documents to logs. That indirection is the entire trick.

Rollover shown as two panels. Before, a write alias named logs points at logs-000002, with logs-000001 sitting behind it on the read path only. After rollover, OpenSearch has created logs-000003, the write alias now points there, and both logs-000001 and logs-000002 are read-only history. A caption notes writers keep sending to the alias name, so only the index behind it changes and ingestion never pauses.

The index name has to end in a number. The index behind the alias must match the pattern ^.*-\d+$, like logs-000001. Rollover reads that trailing number, increments it, and creates logs-000002. If you bootstrap the first index without the numeric suffix, rollover has nothing to increment and the action errors.

You also have to tell each index which alias it rolls over. That goes in the plugins.index_state_management.rollover_alias setting on the index. The cleanest way to apply it is an index template, so every new logs-* index carries it. The snippet below creates that template, then bootstraps the first index with the write alias:

PUT _index_template/logs_template
{
  "index_patterns": ["logs-*"],
  "template": {
    "settings": {
      "plugins.index_state_management.rollover_alias": "logs"
    }
  }
}

PUT logs-000001
{ "aliases": { "logs": { "is_write_index": true } } }

That is_write_index: true is what makes the alias writable and gives rollover a single index to swap. Miss it and OpenSearch rejects writes to a multi-index alias, because it cannot tell which index should receive them. When rollover feels broken, the cause is almost always one of three things: the alias, the trailing number, or the rollover_alias setting.

Attaching the policy to current and future indexes

Writing the policy is half the job. The other half is making sure every index it should govern actually runs it, including indexes that do not exist yet.

The ism_template block in the policy above handles that. Its index_patterns say which index names the policy claims. priority breaks ties when several policies match the same name, and the higher number wins. With that block in place, any new index whose name matches gets the policy attached automatically at creation. This is what makes rollover sustainable: the rolled-over logs-000002 inherits the same policy that created it, so the cycle continues without anyone re-attaching anything.

Indexes that existed before the policy need it attached once, by hand:

POST _plugins/_ism/add/logs-*
{ "policy_id": "logs_policy" }

After that, ISM manages them on its own schedule. It wakes on an interval (a few minutes by default), checks each managed index against its current state’s conditions, and runs whatever is due. It is a background sweep, not an instant trigger, and that matters when you reason about timing.

Where it bites

The happy path is a dozen lines of JSON. The failures are the interesting part, and most of them come from misreading how ISM measures time and how it recovers from errors.

Age is measured from rollover, not from when you attach the policy. min_index_age counts from the index’s creation, and for a rolled-over index that is the moment rollover created it. Attach a policy with a 7d transition to a month-old index and the index moves immediately, because it is already past the threshold. This surprises people who expect the clock to start at attachment. When you backfill a policy onto historical indexes, assume the transitions fire on the next sweep.

A stuck action does not skip; it retries. Say the delete action fails because of a permissions problem, or force_merge fails under load. The index parks in a failed state, and ISM keeps retrying rather than moving on. That is the safe default, since you do not want ISM silently abandoning a deletion, but it means a broken policy quietly piles up work. When indexes are not aging the way you expect, run the Explain API first (GET _plugins/_ism/explain/logs-*). It tells you each index’s current state, its last action, and any error holding it there.

Rollover conditions are checked on a schedule, so a shard can overshoot. Between the moment a shard passes 50GB and the next ISM sweep, the shard keeps taking writes. Under a traffic spike, it can run well past the limit before rollover catches it. Set the threshold with headroom below what the node can actually hold. min_primary_shard_size is a trigger, not a hard cap.

Editing a policy does not retroactively rewrite in-flight indexes. When you change a policy, the indexes it already manages keep running the version they started under until they next transition and pick up the change. If you need every index on the new version now, the ISM API lets you retry or move indexes explicitly. By default, though, changes roll out gradually, so plan edits knowing the old behaviour lingers for a while.

Deletes are real and immediate. The delete state removes an index the moment the index reaches it, and there is no undo. Get the retention math wrong in the transition condition and you delete data you meant to keep. Before pushing a policy that ends in delete, I point the delete transition at a longer window than the real target, watch a full cycle in the Explain API, and only then tighten it. If the data matters, add a snapshot action before delete, so expiry archives the index to a repository instead of vaporizing it.

What I would set up on day one

If I were starting a cluster fresh, this is the order that saves the most pain:

  1. The index template with the rollover_alias setting.
  2. The bootstrap index, with a numeric suffix and a write alias.
  3. The policy, with an ism_template so nothing created later escapes it.

Get that scaffolding right once and every future index is born managed. You never have the conversation that starts with “why is this one node full.”

I care about this beyond tidiness because observability is only useful if it keeps running. I have written before about building OpenSearch dashboards to watch workflow health, and a dashboard is worth exactly nothing if the index behind it went read-only last Tuesday. ISM is the unglamorous layer that keeps the data you monitor from becoming the outage you are monitoring for. It is the same instinct behind sizing Kubernetes memory limits so a pod fails predictably instead of taking a node with it. Boring, automated, and enforced by the system rather than by whoever remembers: that is the version that survives a year of production.


Diagrams by M. Hassan Ahmed, released under CC0. No external image was used for this post; the figures are original work by the author.