{"componentChunkName":"component---src-templates-blog-post-js","path":"/blog/2026-07-01-opensearch-workflow-monitoring/","result":{"data":{"site":{"siteMetadata":{"title":"M.Hassan Ahmed","author":"Hassan11196"}},"markdownRemark":{"id":"78d1a54e-bbcf-587c-85db-82427cc13cbd","excerpt":"A workflow pipeline fails quietly. Jobs move through states, get scheduled across a fleet of sites, retry, and mostly finish. Then one dataset stops progressing…","html":"<p>A workflow pipeline fails quietly. Jobs move through states, get scheduled across a fleet of sites, retry, and mostly finish. Then one dataset stops progressing, and nobody notices until someone downstream asks where their output is. By then the failure is hours old. The useful context (which job, which site, which exit code) is buried in a log file on a machine you have to go find.</p>\n<p>This post is for engineers running any kind of batch or workflow system, such as a job scheduler, an ETL pipeline, or a CI fleet. If you want to see that system’s state on a screen instead of grepping logs after the fact, this covers the setup: structured events, a mapping, the queries behind each panel, alerts, and retention.</p>\n<p>I ran the workflow management stack (WMCore + Unified) for the CMS experiment at CERN. There, Monte Carlo production and reconstruction jobs run across the <a href=\"https://wlcg.web.cern.ch/\">Worldwide LHC Computing Grid</a>. Part of that work is the <a href=\"/project/cms-workflow-operations/\">OpenSearch and Grafana monitoring</a> that gives operators visibility before a dataset delay turns into a problem. The pattern below is the reusable core of it, with the CMS-specific parts stripped out.</p>\n<p>I will use <a href=\"https://opensearch.org/docs/latest/\">OpenSearch</a>, the Apache-2.0 fork of Elasticsearch and Kibana. Everything here maps almost one to one onto Elasticsearch if that is what you run.</p>\n<h2>Start with structured events, not prose logs</h2>\n<p>Most pipelines already log. The problem is that they log prose:</p>\n<div class=\"gatsby-highlight\" data-language=\"text\"><pre class=\"language-text\"><code class=\"language-text\">2026-07-01 09:14:02 job 88213 for mc-reco-2026 failed on T2_CH_CERN after 812s, exit 8021</code></pre></div>\n<p>A human reads that fine. A dashboard cannot. To count failures per site, you have to parse the sentence back into fields with a fragile regex, and that regex breaks the day someone reorders the message. The fix is to emit the fields directly and let the sentence go. Here is the same event as a JSON document:</p>\n<div class=\"gatsby-highlight\" data-language=\"json\"><pre class=\"language-json\"><code class=\"language-json\"><span class=\"token punctuation\">{</span>\n  <span class=\"token property\">\"@timestamp\"</span><span class=\"token operator\">:</span> <span class=\"token string\">\"2026-07-01T09:14:02Z\"</span><span class=\"token punctuation\">,</span>\n  <span class=\"token property\">\"workflow\"</span><span class=\"token operator\">:</span> <span class=\"token string\">\"mc-reco-2026\"</span><span class=\"token punctuation\">,</span>\n  <span class=\"token property\">\"state\"</span><span class=\"token operator\">:</span> <span class=\"token string\">\"failed\"</span><span class=\"token punctuation\">,</span>\n  <span class=\"token property\">\"site\"</span><span class=\"token operator\">:</span> <span class=\"token string\">\"T2_CH_CERN\"</span><span class=\"token punctuation\">,</span>\n  <span class=\"token property\">\"attempt\"</span><span class=\"token operator\">:</span> <span class=\"token number\">3</span><span class=\"token punctuation\">,</span>\n  <span class=\"token property\">\"duration_s\"</span><span class=\"token operator\">:</span> <span class=\"token number\">812</span><span class=\"token punctuation\">,</span>\n  <span class=\"token property\">\"exit_code\"</span><span class=\"token operator\">:</span> <span class=\"token number\">8021</span>\n<span class=\"token punctuation\">}</span></code></pre></div>\n<p>Same information, but now <code class=\"language-text\">state</code> and <code class=\"language-text\">site</code> are queryable dimensions, and <code class=\"language-text\">duration_s</code> is a number you can average. Every panel and every alert later in this post is just a question asked against these fields. Structure the event once, at the source, and everything downstream gets cheap.</p>\n<p>The diagram below shows the whole path. The pipeline emits one JSON document per state change, a shipper pushes it into OpenSearch, and OpenSearch Dashboards asks aggregate questions of the index.</p>\n<p><img src=\"/9c3d4eb33c4bb820d9a71a72b53c6c13/pipeline-to-dashboard.svg\" alt=\"A pipeline emits one flat JSON document per job state change; a log shipper such as Fluent Bit or Data Prepper bulk-loads it into an OpenSearch index with an explicit mapping and ISM rollover; OpenSearch Dashboards runs terms aggregations, percentiles, and alert thresholds against those fields\"></p>\n<p>If you get to choose field names, borrow them from the <a href=\"https://www.elastic.co/guide/en/ecs/current/ecs-reference.html\">Elastic Common Schema</a>. Naming things <code class=\"language-text\">event.duration</code> and <code class=\"language-text\">host.name</code> instead of inventing your own means integrations, and future-you, already know what the fields mean.</p>\n<h2>Define the mapping before the data arrives</h2>\n<p>The mapping is the index’s schema: the type of each field. OpenSearch will happily index a document it has never seen and guess the types. That guessing, called dynamic mapping, is the trap. Guess wrong once, and the field is stuck with the wrong type until you reindex. The classic failure is an id like <code class=\"language-text\">exit_code</code> that arrives as a number in one document and a string in another. OpenSearch rejects the second one, and you lose events without an obvious error on the dashboard.</p>\n<p>Set the types yourself with an <a href=\"https://opensearch.org/docs/latest/im-plugin/index-templates/\">index template</a>, so every new <code class=\"language-text\">workflow-events-*</code> index inherits them:</p>\n<div class=\"gatsby-highlight\" data-language=\"json\"><pre class=\"language-json\"><code class=\"language-json\">PUT _index_template/workflow-events\n<span class=\"token punctuation\">{</span>\n  <span class=\"token property\">\"index_patterns\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">[</span><span class=\"token string\">\"workflow-events-*\"</span><span class=\"token punctuation\">]</span><span class=\"token punctuation\">,</span>\n  <span class=\"token property\">\"template\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">{</span>\n    <span class=\"token property\">\"mappings\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">{</span>\n      <span class=\"token property\">\"properties\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">{</span>\n        <span class=\"token property\">\"@timestamp\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">{</span> <span class=\"token property\">\"type\"</span><span class=\"token operator\">:</span> <span class=\"token string\">\"date\"</span> <span class=\"token punctuation\">}</span><span class=\"token punctuation\">,</span>\n        <span class=\"token property\">\"workflow\"</span><span class=\"token operator\">:</span>   <span class=\"token punctuation\">{</span> <span class=\"token property\">\"type\"</span><span class=\"token operator\">:</span> <span class=\"token string\">\"keyword\"</span> <span class=\"token punctuation\">}</span><span class=\"token punctuation\">,</span>\n        <span class=\"token property\">\"state\"</span><span class=\"token operator\">:</span>      <span class=\"token punctuation\">{</span> <span class=\"token property\">\"type\"</span><span class=\"token operator\">:</span> <span class=\"token string\">\"keyword\"</span> <span class=\"token punctuation\">}</span><span class=\"token punctuation\">,</span>\n        <span class=\"token property\">\"site\"</span><span class=\"token operator\">:</span>       <span class=\"token punctuation\">{</span> <span class=\"token property\">\"type\"</span><span class=\"token operator\">:</span> <span class=\"token string\">\"keyword\"</span> <span class=\"token punctuation\">}</span><span class=\"token punctuation\">,</span>\n        <span class=\"token property\">\"attempt\"</span><span class=\"token operator\">:</span>    <span class=\"token punctuation\">{</span> <span class=\"token property\">\"type\"</span><span class=\"token operator\">:</span> <span class=\"token string\">\"integer\"</span> <span class=\"token punctuation\">}</span><span class=\"token punctuation\">,</span>\n        <span class=\"token property\">\"duration_s\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">{</span> <span class=\"token property\">\"type\"</span><span class=\"token operator\">:</span> <span class=\"token string\">\"long\"</span> <span class=\"token punctuation\">}</span><span class=\"token punctuation\">,</span>\n        <span class=\"token property\">\"exit_code\"</span><span class=\"token operator\">:</span>  <span class=\"token punctuation\">{</span> <span class=\"token property\">\"type\"</span><span class=\"token operator\">:</span> <span class=\"token string\">\"keyword\"</span> <span class=\"token punctuation\">}</span>\n      <span class=\"token punctuation\">}</span>\n    <span class=\"token punctuation\">}</span>\n  <span class=\"token punctuation\">}</span>\n<span class=\"token punctuation\">}</span></code></pre></div>\n<p>The important choice here is <code class=\"language-text\">keyword</code> versus <code class=\"language-text\">text</code>:</p>\n<ul>\n<li>A <code class=\"language-text\">text</code> field is analyzed, meaning it is broken into tokens for full-text search. It cannot be grouped on without extra work.</li>\n<li>A <code class=\"language-text\">keyword</code> field is stored whole, and it is what you aggregate on.</li>\n</ul>\n<p>Anything you want to count, filter, or group by (states, sites, workflow names) is a <code class=\"language-text\">keyword</code>. I mapped <code class=\"language-text\">exit_code</code> as a keyword on purpose, even though it looks numeric. I never do arithmetic on it; I only group by it. Treating codes as categories avoids the type-collision problem entirely.</p>\n<h2>Getting events into the index</h2>\n<p>You have two sane options for moving documents into the index.</p>\n<p><strong>A log shipper.</strong> The shipper tails your structured output and forwards it. <a href=\"https://docs.fluentbit.io/manual/pipeline/outputs/opensearch\">Fluent Bit</a> is the light one and has a native OpenSearch output. <a href=\"https://opensearch.org/docs/latest/data-prepper/\">Data Prepper</a> is the OpenSearch-native option if you want to parse and enrich events in flight. On Kubernetes the shipper runs as a DaemonSet (one copy on every node), and you barely think about it. This is the right default, because your application code stays out of the indexing path entirely: it writes JSON to stdout, and the shipper does the rest.</p>\n<p><strong>The Bulk API, directly.</strong> A small service writes to the <a href=\"https://opensearch.org/docs/latest/api-reference/document-apis/bulk/\">Bulk API</a> itself. That is worth it when you want to control batching and backpressure (holding producers back when the cluster falls behind) yourself. The snippet below uses the Python client’s bulk helper to send events in batches:</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\"><span class=\"token keyword\">from</span> opensearchpy <span class=\"token keyword\">import</span> OpenSearch<span class=\"token punctuation\">,</span> helpers\n\nclient <span class=\"token operator\">=</span> OpenSearch<span class=\"token punctuation\">(</span><span class=\"token string\">\"https://opensearch:9200\"</span><span class=\"token punctuation\">,</span> http_auth<span class=\"token operator\">=</span><span class=\"token punctuation\">(</span>user<span class=\"token punctuation\">,</span> pw<span class=\"token punctuation\">)</span><span class=\"token punctuation\">,</span> verify_certs<span class=\"token operator\">=</span><span class=\"token boolean\">True</span><span class=\"token punctuation\">)</span>\n\n<span class=\"token keyword\">def</span> <span class=\"token function\">index_events</span><span class=\"token punctuation\">(</span>events<span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n    actions <span class=\"token operator\">=</span> <span class=\"token punctuation\">(</span><span class=\"token punctuation\">{</span><span class=\"token string\">\"_index\"</span><span class=\"token punctuation\">:</span> <span class=\"token string\">\"workflow-events-write\"</span><span class=\"token punctuation\">,</span> <span class=\"token operator\">**</span>e<span class=\"token punctuation\">}</span> <span class=\"token keyword\">for</span> e <span class=\"token keyword\">in</span> events<span class=\"token punctuation\">)</span>\n    helpers<span class=\"token punctuation\">.</span>bulk<span class=\"token punctuation\">(</span>client<span class=\"token punctuation\">,</span> actions<span class=\"token punctuation\">,</span> chunk_size<span class=\"token operator\">=</span><span class=\"token number\">500</span><span class=\"token punctuation\">,</span> request_timeout<span class=\"token operator\">=</span><span class=\"token number\">30</span><span class=\"token punctuation\">)</span></code></pre></div>\n<p>Batch it. One HTTP request per event will melt the cluster under any real load. The bulk helper groups events, and 500 to a few thousand documents per request is the usual sweet spot. Write to an alias like <code class=\"language-text\">workflow-events-write</code> rather than a dated index name, so the rollover described later is invisible to the writer.</p>\n<h2>Each panel is one aggregation query</h2>\n<p>Once the fields are in, each dashboard panel is one aggregation query: a query that returns counts or statistics over many documents instead of the documents themselves. Build the panels in the Dashboards UI, but know the query underneath, so you understand what each panel costs and can reproduce it outside the UI.</p>\n<p><strong>Jobs per state.</strong> How many jobs are in each state right now is a <code class=\"language-text\">terms</code> aggregation, which counts documents per distinct value of a field:</p>\n<div class=\"gatsby-highlight\" data-language=\"json\"><pre class=\"language-json\"><code class=\"language-json\">GET workflow-events-*/_search\n<span class=\"token punctuation\">{</span>\n  <span class=\"token property\">\"size\"</span><span class=\"token operator\">:</span> <span class=\"token number\">0</span><span class=\"token punctuation\">,</span>\n  <span class=\"token property\">\"query\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">{</span> <span class=\"token property\">\"range\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">{</span> <span class=\"token property\">\"@timestamp\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">{</span> <span class=\"token property\">\"gte\"</span><span class=\"token operator\">:</span> <span class=\"token string\">\"now-1h\"</span> <span class=\"token punctuation\">}</span> <span class=\"token punctuation\">}</span> <span class=\"token punctuation\">}</span><span class=\"token punctuation\">,</span>\n  <span class=\"token property\">\"aggs\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">{</span>\n    <span class=\"token property\">\"by_state\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">{</span> <span class=\"token property\">\"terms\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">{</span> <span class=\"token property\">\"field\"</span><span class=\"token operator\">:</span> <span class=\"token string\">\"state\"</span><span class=\"token punctuation\">,</span> <span class=\"token property\">\"size\"</span><span class=\"token operator\">:</span> <span class=\"token number\">20</span> <span class=\"token punctuation\">}</span> <span class=\"token punctuation\">}</span>\n  <span class=\"token punctuation\">}</span>\n<span class=\"token punctuation\">}</span></code></pre></div>\n<p><code class=\"language-text\">size: 0</code> means return no raw documents, only the aggregated counts. That is all a panel needs, and it keeps the response small. Nest a second <code class=\"language-text\">terms</code> on <code class=\"language-text\">site</code> inside <code class=\"language-text\">by_state</code> and you get a breakdown of failures per site in the same request.</p>\n<p><strong>Throughput.</strong> Whether the pipeline is keeping up is a <code class=\"language-text\">date_histogram</code>: throughput bucketed over time. It becomes the line chart everyone stares at. This one counts completed jobs per five-minute bucket:</p>\n<div class=\"gatsby-highlight\" data-language=\"json\"><pre class=\"language-json\"><code class=\"language-json\"><span class=\"token property\">\"aggs\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">{</span>\n  <span class=\"token property\">\"over_time\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">{</span>\n    <span class=\"token property\">\"date_histogram\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">{</span> <span class=\"token property\">\"field\"</span><span class=\"token operator\">:</span> <span class=\"token string\">\"@timestamp\"</span><span class=\"token punctuation\">,</span> <span class=\"token property\">\"fixed_interval\"</span><span class=\"token operator\">:</span> <span class=\"token string\">\"5m\"</span> <span class=\"token punctuation\">}</span><span class=\"token punctuation\">,</span>\n    <span class=\"token property\">\"aggs\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">{</span> <span class=\"token property\">\"completed\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">{</span> <span class=\"token property\">\"filter\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">{</span> <span class=\"token property\">\"term\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">{</span> <span class=\"token property\">\"state\"</span><span class=\"token operator\">:</span> <span class=\"token string\">\"completed\"</span> <span class=\"token punctuation\">}</span> <span class=\"token punctuation\">}</span> <span class=\"token punctuation\">}</span> <span class=\"token punctuation\">}</span>\n  <span class=\"token punctuation\">}</span>\n<span class=\"token punctuation\">}</span></code></pre></div>\n<p><strong>Duration.</strong> How long jobs are taking is a <code class=\"language-text\">percentiles</code> aggregation on <code class=\"language-text\">duration_s</code>. Watch the p95 and p99 (the 95th and 99th percentiles), not the average. The mean hides the tail, and the tail is where the stuck jobs live. An average duration that looks healthy while p99 quietly climbs is the early signal that something is starting to back up.</p>\n<p><strong>Filtering.</strong> To filter across the whole dashboard, use the <a href=\"https://opensearch.org/docs/latest/dashboards/dql/\">Dashboards Query Language</a> bar at the top. <code class=\"language-text\">state: failed and site: T2_CH_CERN</code> narrows every panel at once. That is how you go from “something is wrong” to “it is this site” in a couple of keystrokes.</p>\n<h2>Alerting: the point of the whole exercise</h2>\n<p>A dashboard only helps when someone is looking at it. At 3 a.m., nobody is. The <a href=\"https://opensearch.org/docs/latest/observing-your-data/alerting/index/\">alerting plugin</a> closes that gap: it runs a saved query on a schedule and fires when a condition holds.</p>\n<p>A monitor is the query plus a trigger. The query finds jobs stuck in a non-terminal state for longer than they should be (here, <code class=\"language-text\">running</code> for 3600 seconds or more, checked every five minutes). The trigger says how many is too many:</p>\n<div class=\"gatsby-highlight\" data-language=\"json\"><pre class=\"language-json\"><code class=\"language-json\"><span class=\"token punctuation\">{</span>\n  <span class=\"token property\">\"name\"</span><span class=\"token operator\">:</span> <span class=\"token string\">\"stuck-jobs\"</span><span class=\"token punctuation\">,</span>\n  <span class=\"token property\">\"schedule\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">{</span> <span class=\"token property\">\"period\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">{</span> <span class=\"token property\">\"interval\"</span><span class=\"token operator\">:</span> <span class=\"token number\">5</span><span class=\"token punctuation\">,</span> <span class=\"token property\">\"unit\"</span><span class=\"token operator\">:</span> <span class=\"token string\">\"MINUTES\"</span> <span class=\"token punctuation\">}</span> <span class=\"token punctuation\">}</span><span class=\"token punctuation\">,</span>\n  <span class=\"token property\">\"inputs\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">[</span><span class=\"token punctuation\">{</span>\n    <span class=\"token property\">\"search\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">{</span>\n      <span class=\"token property\">\"indices\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">[</span><span class=\"token string\">\"workflow-events-*\"</span><span class=\"token punctuation\">]</span><span class=\"token punctuation\">,</span>\n      <span class=\"token property\">\"query\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">{</span>\n        <span class=\"token property\">\"size\"</span><span class=\"token operator\">:</span> <span class=\"token number\">0</span><span class=\"token punctuation\">,</span>\n        <span class=\"token property\">\"query\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">{</span> <span class=\"token property\">\"bool\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">{</span> <span class=\"token property\">\"must\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">[</span>\n          <span class=\"token punctuation\">{</span> <span class=\"token property\">\"term\"</span><span class=\"token operator\">:</span>  <span class=\"token punctuation\">{</span> <span class=\"token property\">\"state\"</span><span class=\"token operator\">:</span> <span class=\"token string\">\"running\"</span> <span class=\"token punctuation\">}</span> <span class=\"token punctuation\">}</span><span class=\"token punctuation\">,</span>\n          <span class=\"token punctuation\">{</span> <span class=\"token property\">\"range\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">{</span> <span class=\"token property\">\"duration_s\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">{</span> <span class=\"token property\">\"gte\"</span><span class=\"token operator\">:</span> <span class=\"token number\">3600</span> <span class=\"token punctuation\">}</span> <span class=\"token punctuation\">}</span> <span class=\"token punctuation\">}</span>\n        <span class=\"token punctuation\">]</span><span class=\"token punctuation\">}</span><span class=\"token punctuation\">}</span><span class=\"token punctuation\">,</span>\n        <span class=\"token property\">\"aggs\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">{</span> <span class=\"token property\">\"stuck\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">{</span> <span class=\"token property\">\"cardinality\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">{</span> <span class=\"token property\">\"field\"</span><span class=\"token operator\">:</span> <span class=\"token string\">\"workflow\"</span> <span class=\"token punctuation\">}</span> <span class=\"token punctuation\">}</span> <span class=\"token punctuation\">}</span>\n      <span class=\"token punctuation\">}</span>\n    <span class=\"token punctuation\">}</span>\n  <span class=\"token punctuation\">}</span><span class=\"token punctuation\">]</span>\n<span class=\"token punctuation\">}</span></code></pre></div>\n<p>The trigger condition is a small Painless script (OpenSearch’s built-in scripting language) evaluated against the result: <code class=\"language-text\">ctx.results[0].aggregations.stuck.value &gt; 5</code>. The action is a notification channel, such as Slack, email, or a webhook into your on-call tool. Set the threshold above your normal noise floor. An alert that fires every day is one people mute, and a muted alert is the same as no alert.</p>\n<h2>Retention, or the disk fills up</h2>\n<p>Workflow events never stop arriving, and an index that only grows will eventually take the cluster down with it. <a href=\"https://opensearch.org/docs/latest/im-plugin/ism/index/\">Index State Management</a> (ISM) is the answer. An ISM policy rolls the write index over to a fresh one once it hits a size or age, and deletes old indices past your retention window. The policy below rolls over at 30GB or one day, and deletes indices once they are 30 days old:</p>\n<div class=\"gatsby-highlight\" data-language=\"json\"><pre class=\"language-json\"><code class=\"language-json\"><span class=\"token punctuation\">{</span>\n  <span class=\"token property\">\"policy\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">{</span>\n    <span class=\"token property\">\"states\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">[</span>\n      <span class=\"token punctuation\">{</span> <span class=\"token property\">\"name\"</span><span class=\"token operator\">:</span> <span class=\"token string\">\"hot\"</span><span class=\"token punctuation\">,</span>\n        <span class=\"token property\">\"actions\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">[</span><span class=\"token punctuation\">{</span> <span class=\"token property\">\"rollover\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">{</span> <span class=\"token property\">\"min_size\"</span><span class=\"token operator\">:</span> <span class=\"token string\">\"30gb\"</span><span class=\"token punctuation\">,</span> <span class=\"token property\">\"min_index_age\"</span><span class=\"token operator\">:</span> <span class=\"token string\">\"1d\"</span> <span class=\"token punctuation\">}</span> <span class=\"token punctuation\">}</span><span class=\"token punctuation\">]</span><span class=\"token punctuation\">,</span>\n        <span class=\"token property\">\"transitions\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">[</span><span class=\"token punctuation\">{</span> <span class=\"token property\">\"state_name\"</span><span class=\"token operator\">:</span> <span class=\"token string\">\"delete\"</span><span class=\"token punctuation\">,</span> <span class=\"token property\">\"conditions\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">{</span> <span class=\"token property\">\"min_index_age\"</span><span class=\"token operator\">:</span> <span class=\"token string\">\"30d\"</span> <span class=\"token punctuation\">}</span> <span class=\"token punctuation\">}</span><span class=\"token punctuation\">]</span> <span class=\"token punctuation\">}</span><span class=\"token punctuation\">,</span>\n      <span class=\"token punctuation\">{</span> <span class=\"token property\">\"name\"</span><span class=\"token operator\">:</span> <span class=\"token string\">\"delete\"</span><span class=\"token punctuation\">,</span> <span class=\"token property\">\"actions\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">[</span><span class=\"token punctuation\">{</span> <span class=\"token property\">\"delete\"</span><span class=\"token operator\">:</span> <span class=\"token punctuation\">{</span><span class=\"token punctuation\">}</span> <span class=\"token punctuation\">}</span><span class=\"token punctuation\">]</span> <span class=\"token punctuation\">}</span>\n    <span class=\"token punctuation\">]</span>\n  <span class=\"token punctuation\">}</span>\n<span class=\"token punctuation\">}</span></code></pre></div>\n<p>This is why the writer targets an alias instead of a fixed index name. ISM swaps the underlying index on rollover. The write alias always points at the current index, so nothing upstream has to know the swap happened.</p>\n<h2>Failure modes I have actually hit</h2>\n<p><strong>Timezones on <code class=\"language-text\">@timestamp</code>.</strong> These bite first. Log a naive local time (one with no timezone offset) and OpenSearch assumes UTC. Your panels are then silently shifted by hours, and the 3 a.m. spike shows up at what looks like the wrong time. Always emit ISO-8601 with an explicit offset, ideally UTC, at the source.</p>\n<p><strong>Dynamic fields from free-form objects.</strong> Suppose part of your event is a free-form object, such as per-job metadata with arbitrary keys. Dynamic mapping creates a new field for every distinct key it sees. Thousands of one-off fields bloat the cluster state (the cluster-wide metadata that includes every mapping) and slow everything down. Set <code class=\"language-text\">&quot;dynamic&quot;: &quot;strict&quot;</code> on the objects you control, or store the loose part as a single stringified blob you do not aggregate on.</p>\n<p><strong>High-cardinality group-bys.</strong> People learn this one the hard way. A <code class=\"language-text\">terms</code> aggregation on individual <code class=\"language-text\">job_id</code> values asks OpenSearch to build a bucket per job, which is fine at a hundred and painful at ten million. Aggregate on the low-cardinality dimensions (those with few distinct values), like state and site. Reach for the raw documents only once you have narrowed down to a handful.</p>\n<h2>Tradeoffs, and where Grafana fits</h2>\n<p>OpenSearch Dashboards and <a href=\"https://grafana.com/docs/grafana/latest/\">Grafana</a> overlap, and I run both. Here is the split I settled on:</p>\n<ul>\n<li><strong>OpenSearch owns the event data</strong>: the discrete “job X changed to state Y” records you want to search, filter, and drill into.</li>\n<li><strong>Grafana is stronger for time-series metrics</strong> from a system like Prometheus: continuous numbers such as queue depth and CPU.</li>\n</ul>\n<p>Reaching for the wrong one is a common early mistake. Full-text-searchable events want a search engine; regularly sampled numeric series want a time-series database. Forcing high-frequency metrics into per-sample OpenSearch documents works right up until the document count buries you.</p>\n<p>The honest cost of this whole setup is that OpenSearch is a stateful cluster you now operate. It needs memory, disk, and enough attention that ISM and shard sizing do not surprise you. For a small pipeline that is real overhead, and a few structured log files plus <code class=\"language-text\">jq</code> might be all you need. The pattern earns its keep once you have enough jobs, sites, or states that no single person can hold the state of the system in their head.</p>\n<h2>What I would do differently</h2>\n<p>I would define the mapping template on day one instead of letting the first documents create it dynamically. Early on I let OpenSearch infer types. Weeks later a field arrived in a shape the guess did not cover, I hit a type collision, and I had to reindex to fix it. Writing the template first is ten minutes that saves that afternoon.</p>\n<p>I would also wire up the stuck-jobs alert before building a single pretty panel. It is tempting to spend the first day making charts, but charts only help when someone is watching them, and the alert covers the hours when nobody is. The dashboard is for investigating a problem you already know about; the alert is what tells you there is one.</p>\n<p>For fuller context, the <a href=\"/project/cms-workflow-operations/\">CMS Workflow Operations</a> writeup covers the WMCore and Unified stack this monitoring sits on top of. <a href=\"/project/archi/\">Archi</a> is the retrieval copilot I built for the same operations team. It uses the same structured operational data from the other direction, searching it instead of charting it.</p>\n<hr>\n<p><em>Image credit: workflow-monitoring diagrams by M. Hassan Ahmed, created for this post, released under CC0 (public domain).</em></p>","frontmatter":{"title":"OpenSearch Dashboards for Workflow Monitoring","date":"2026-07-01T00:00:00.000Z","description":"Build an OpenSearch dashboard that watches a job pipeline: structured logs, an index mapping, the queries behind each panel, and alerts that fire early.","thumbnail":{"childImageSharp":{"fluid":{"base64":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAABQAAAALCAIAAADwazoUAAAACXBIWXMAAAsSAAALEgHS3X78AAACD0lEQVQozyWQx47bMBRFtfJYtkV1iU1dVqWaLTu25CJkCiabAMkmi+zT1lkHmK/Jf+TLQs8ABwcX7/GCBIVV0nGkpPOP78PhIbk8R6en6PTI7Q/3yfUD7s888Al3OD7S/RV242rd8oqwBPSGRByvKqpTvZnq7pqzoWov3GV9SotD2ZxZcy6qMa/GrDxSl4krzFvCUvOWqisCOpfpggfFWSjuGzzfAcLNV5zbSnX5+VffEBYyNQiLyguJd7KdarjUcGE4jU4rFWVBPnrZ4Oejnw0k2il2ztFgaZJGNhNBXGKz6IuXn1awhf5+BdnMqBTCJFSoMDdoc4PUptsafqf5jRZ11naEw2Q0O0EUobk5pH9/B/2X67d/fVt//9TbtImCDsBcTzZ6ujXY3mwP1m6E5yu8TnCa0MO90e35zVCPm+DrR5Lfl08vu7b+8bkvwv65Hpb2Wo9aLWjkN7waOBVHdirFqQFMBRFAyYhViykwlXRvZWYzJQYol3GpWAmwEl0PEitcwziG8RqtuV3T90wf6a4wW9p+xKbLCYdbg1Y6rSVrzdJDlh1lXKm0IWZ4oPEpZh2N3vnZPsgTywk1FOpUmK9sSJMiZ8hjFi01xPhDqNsit1X4n+HKMoJf5f5PP21InOKwIGFsOrFOgteyNZfQnYS5RQktALc955bxLQM+gaaMDRlrClZlxFEA1GQEAPwPi4NYmDI7eDAAAAAASUVORK5CYII=","aspectRatio":1.899441340782123,"src":"/static/9c33da952bb8f40f57c05eb37a28efb1/40a76/hero.png","srcSet":"/static/9c33da952bb8f40f57c05eb37a28efb1/c972b/hero.png 340w,\n/static/9c33da952bb8f40f57c05eb37a28efb1/27625/hero.png 680w,\n/static/9c33da952bb8f40f57c05eb37a28efb1/40a76/hero.png 1360w,\n/static/9c33da952bb8f40f57c05eb37a28efb1/ed396/hero.png 2000w","sizes":"(max-width: 1360px) 100vw, 1360px"}}}}}},"pageContext":{"slug":"/2026-07-01-opensearch-workflow-monitoring/","previous":"blog/2026-07-02-hybrid-search-rag-bm25-vectors/","next":"blog/2026-07-05-sandboxing-llm-generated-code/"}},"staticQueryHashes":["32046230"]}