{"componentChunkName":"component---src-templates-blog-post-js","path":"/blog/2026-07-14-reliable-json-from-llms/","result":{"data":{"site":{"siteMetadata":{"title":"M.Hassan Ahmed","author":"Hassan11196"}},"markdownRemark":{"id":"a41767e3-70ee-5b7c-a2bc-ed563e723851","excerpt":"The first version always works. You write a prompt that ends with “respond with JSON like ”, call  on the reply, and it parses. You ship it. A week later, the…","html":"<p>The first version always works. You write a prompt that ends with “respond with JSON like <code class=\"language-text\">{&quot;intent&quot;: &quot;...&quot;, &quot;confidence&quot;: 0.0}</code>”, call <code class=\"language-text\">json.loads()</code> on the reply, and it parses. You ship it. A week later, the model answers some request with “Sure! Here is the JSON you asked for:” followed by a code fence. <code class=\"language-text\">json.loads()</code> throws, and the endpoint 500s.</p>\n<p>That is the whole problem in one sentence: the model produces text, and text that is <em>mostly</em> JSON is not JSON. The gap between the two is where the on-call pages come from.</p>\n<p>This post is for backend engineers who use an LLM as a component in a larger system: a classifier that returns a label, an extractor that pulls fields out of a document, a router that decides which tool to call. You are not reading the model’s output; your code is. When your code is the consumer, “close enough” is a bug.</p>\n<p>I hit every failure mode here while wiring the structured pieces of the <a href=\"/project/archi/\">Archi</a> copilot at CERN and the <a href=\"/project/llm-dev-mate/\">token-aware context bundling in LLM DevMate</a>. The fixes fall into three layers that stack on top of each other: describe the output as a schema, constrain the decoding so it always parses, and validate what comes back. This post walks through all three, and the ways each one still bites.</p>\n<h2>Why “ask nicely” is not enough</h2>\n<p>The naive version puts the whole burden on the prompt, then trusts <code class=\"language-text\">json.loads()</code>:</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\">prompt <span class=\"token operator\">=</span> <span class=\"token triple-quoted-string string\">\"\"\"Classify the ticket below. Respond with JSON:\n{\"category\": \"bug\" | \"feature\" | \"question\", \"priority\": 1-5}\n\nTicket: The export button does nothing on Safari.\"\"\"</span>\n\nresp <span class=\"token operator\">=</span> call_model<span class=\"token punctuation\">(</span>prompt<span class=\"token punctuation\">)</span>\ndata <span class=\"token operator\">=</span> json<span class=\"token punctuation\">.</span>loads<span class=\"token punctuation\">(</span>resp<span class=\"token punctuation\">)</span>   <span class=\"token comment\"># &lt;-- this line is the whole risk</span></code></pre></div>\n<p>This works often enough to feel done, and that is the trap. The failures are not random noise. They are specific, and they cluster:</p>\n<ul>\n<li><strong>Preamble.</strong> <code class=\"language-text\">Here is the classification:</code> before the object, or a trailing <code class=\"language-text\">Let me know if you need anything else.</code> after it. A <a href=\"https://commonmark.org/\">Markdown</a> code fence (<code class=\"language-text\">```json</code>) wrapped around the whole thing is the most common variant.</li>\n<li><strong>Almost-valid JSON.</strong> A trailing comma, single quotes instead of double, <code class=\"language-text\">None</code>/<code class=\"language-text\">True</code> instead of <code class=\"language-text\">null</code>/<code class=\"language-text\">true</code> because the model leaned Python, or an unescaped newline inside a string.</li>\n<li><strong>Right shape, wrong types.</strong> <code class=\"language-text\">&quot;priority&quot;: &quot;high&quot;</code> when your schema said an integer 1 to 5, or <code class=\"language-text\">&quot;confidence&quot;: &quot;0.9&quot;</code> as a string. These <em>parse</em> (<code class=\"language-text\">json.loads()</code> is happy) and then blow up three functions later when you do arithmetic on a string.</li>\n<li><strong>Extra or missing keys.</strong> A hallucinated <code class=\"language-text\">&quot;reason&quot;</code> field you never asked for, or a dropped <code class=\"language-text\">&quot;priority&quot;</code> on a ticket the model found ambiguous.</li>\n</ul>\n<p>The last two are the dangerous ones, because parsing succeeds. A <code class=\"language-text\">try/except json.JSONDecodeError</code> catches none of it. Valid JSON of the wrong shape flows into code that assumed the right shape, and the error surfaces far from its cause.</p>\n<p><img src=\"/90b60e1e5afbae4e1d7db59a78cfc2d5/json-pipeline.svg\" alt=\"Three ways to turn a model into structured JSON. Prompt-and-parse relies on the prompt and often produces text that is only nearly JSON, so the parser throws. Constrained decoding masks the token sampler to a grammar, so the string always parses but types can still drift. Validate-and-retry parses, checks against a typed schema, and feeds any error back for one corrected attempt. The reliable path uses the last two together.\"></p>\n<h2>Layer one: describe the shape as a schema, not a sentence</h2>\n<p>The first fix is to stop describing the output in prose and start describing it as <a href=\"https://json-schema.org/\">JSON Schema</a>. A schema is a machine-readable contract: field names, types, which fields are required, and which values an enum allows. In Python, the ergonomic way to write one is a <a href=\"https://docs.pydantic.dev/\">Pydantic</a> model, which gives you the schema and the validator in the same object. This model limits <code class=\"language-text\">category</code> to three values and <code class=\"language-text\">priority</code> to an integer from 1 to 5:</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\"><span class=\"token keyword\">from</span> enum <span class=\"token keyword\">import</span> Enum\n<span class=\"token keyword\">from</span> pydantic <span class=\"token keyword\">import</span> BaseModel<span class=\"token punctuation\">,</span> Field\n\n<span class=\"token keyword\">class</span> <span class=\"token class-name\">Category</span><span class=\"token punctuation\">(</span><span class=\"token builtin\">str</span><span class=\"token punctuation\">,</span> Enum<span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n    bug <span class=\"token operator\">=</span> <span class=\"token string\">\"bug\"</span>\n    feature <span class=\"token operator\">=</span> <span class=\"token string\">\"feature\"</span>\n    question <span class=\"token operator\">=</span> <span class=\"token string\">\"question\"</span>\n\n<span class=\"token keyword\">class</span> <span class=\"token class-name\">Classification</span><span class=\"token punctuation\">(</span>BaseModel<span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n    category<span class=\"token punctuation\">:</span> Category\n    priority<span class=\"token punctuation\">:</span> <span class=\"token builtin\">int</span> <span class=\"token operator\">=</span> Field<span class=\"token punctuation\">(</span>ge<span class=\"token operator\">=</span><span class=\"token number\">1</span><span class=\"token punctuation\">,</span> le<span class=\"token operator\">=</span><span class=\"token number\">5</span><span class=\"token punctuation\">)</span>\n\n<span class=\"token comment\"># Pydantic hands you the JSON Schema for free:</span>\nschema <span class=\"token operator\">=</span> Classification<span class=\"token punctuation\">.</span>model_json_schema<span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span></code></pre></div>\n<p>On its own, putting the schema in the prompt is only a marginal improvement over the prose version, because the model is still free to ignore it. The schema earns its place because the next two layers both consume it. It is the single source of truth: the decoder constrains against it, and the validator checks against it. You write the shape once, and both enforcement points read the same definition.</p>\n<h2>Layer two: constrain the decoding so the string always parses</h2>\n<p>The strongest fix does not live in the prompt at all. It lives in how the tokens are sampled.</p>\n<p>A language model generates one token at a time, drawing each from a probability distribution over its whole vocabulary. <strong>Constrained decoding</strong> puts a filter in front of that draw. At each step, it works out which tokens could still lead to a string that satisfies the schema, sets the probability of every other token to zero, and samples only from what remains. If the grammar says the next character must be a <code class=\"language-text\">}</code> or a <code class=\"language-text\">,</code>, the model physically cannot emit the word “Sure”. The output is guaranteed to parse, because a token that would break the JSON is never on the table.</p>\n<p>You rarely implement this yourself, because the providers expose it. The mechanism is the same one behind <a href=\"/blog/2026-06-30-llm-agent-tool-loop/\">tool calling</a>: a tool’s input schema is enforced exactly this way, which is why a well-defined tool call comes back as clean arguments and a free-text answer does not. Here is where each option lives:</p>\n<ul>\n<li><strong>OpenAI</strong> exposes it as <a href=\"https://platform.openai.com/docs/guides/structured-outputs\">Structured Outputs</a>. Pass your schema in <code class=\"language-text\">response_format</code> with <code class=\"language-text\">strict: true</code>, and the response is guaranteed to match it.</li>\n<li><strong>Anthropic</strong> does the same through <a href=\"https://platform.claude.com/docs/en/build-with-claude/structured-outputs\">structured outputs and strict tool use</a>. Use <code class=\"language-text\">output_config={&quot;format&quot;: {&quot;type&quot;: &quot;json_schema&quot;, &quot;schema&quot;: schema}}</code> on the Messages API, or <code class=\"language-text\">strict: true</code> on a tool definition, with <code class=\"language-text\">client.messages.parse()</code> validating the response against your model automatically.</li>\n<li><strong>Self-hosted</strong> models get it from libraries like <a href=\"https://github.com/dottxt-ai/outlines\">Outlines</a> or <a href=\"https://github.com/ggml-org/llama.cpp/blob/master/grammars/README.md\">llama.cpp’s GBNF grammars</a>, which compile a schema or grammar directly into the token mask.</li>\n</ul>\n<p>The technique is grounded research, not a vendor gimmick. The <a href=\"https://arxiv.org/abs/2307.09702\">Outlines paper on efficient guided generation</a> shows how to build the token masks from a regular expression or grammar with near-zero per-token overhead.</p>\n<p>Turning it on removes the entire first category of failures: no preamble, no code fence, no trailing comma, ever. What it does <em>not</em> remove is type and value drift, and that surprises people. The reason is that the schema a provider’s strict mode enforces is not the full Pydantic schema you wrote.</p>\n<h2>The catch: strict mode enforces only part of JSON Schema</h2>\n<p>Constrained decoding only enforces what the grammar can express, and the grammar covers a subset of JSON Schema. Anthropic’s strict mode, for example, <a href=\"https://platform.claude.com/docs/en/build-with-claude/structured-outputs\">does not enforce numeric bounds like <code class=\"language-text\">minimum</code>/<code class=\"language-text\">maximum</code> or string bounds like <code class=\"language-text\">minLength</code></a>, and it requires every object to set <code class=\"language-text\">additionalProperties: false</code>. OpenAI’s strict mode has a similar list of unsupported keywords.</p>\n<p>So what happens to the <code class=\"language-text\">Field(ge=1, le=5)</code> you wrote on <code class=\"language-text\">priority</code>? The decoder guarantees you get an integer. It does <em>not</em> guarantee the integer is between 1 and 5. The model can still return <code class=\"language-text\">priority: 9</code>, and it will be a syntactically perfect, schema-shaped <code class=\"language-text\">9</code>.</p>\n<p>Constrained decoding buys you a string that parses and a value of the right <em>type</em>. The <em>semantics</em> are still yours to check: the range, the cross-field rules, the “this enum value only makes sense when that flag is set” checks. That is the whole reason the third layer exists.</p>\n<h2>Layer three: validate what came back, and retry once with the error</h2>\n<p>Whatever the decoder gave you, run it through the full validator before anything downstream touches it. This is where Pydantic pays off a second time: the same model that generated the schema now checks the bounds the decoder ignored. The function below validates each reply and, if validation fails, sends the error back to the model for one more attempt:</p>\n<div class=\"gatsby-highlight\" data-language=\"python\"><pre class=\"language-python\"><code class=\"language-python\"><span class=\"token keyword\">from</span> pydantic <span class=\"token keyword\">import</span> ValidationError\n\n<span class=\"token keyword\">def</span> <span class=\"token function\">classify</span><span class=\"token punctuation\">(</span>ticket<span class=\"token punctuation\">:</span> <span class=\"token builtin\">str</span><span class=\"token punctuation\">,</span> retries<span class=\"token punctuation\">:</span> <span class=\"token builtin\">int</span> <span class=\"token operator\">=</span> <span class=\"token number\">1</span><span class=\"token punctuation\">)</span> <span class=\"token operator\">-</span><span class=\"token operator\">></span> Classification<span class=\"token punctuation\">:</span>\n    messages <span class=\"token operator\">=</span> <span class=\"token punctuation\">[</span><span class=\"token punctuation\">{</span><span class=\"token string\">\"role\"</span><span class=\"token punctuation\">:</span> <span class=\"token string\">\"user\"</span><span class=\"token punctuation\">,</span> <span class=\"token string\">\"content\"</span><span class=\"token punctuation\">:</span> build_prompt<span class=\"token punctuation\">(</span>ticket<span class=\"token punctuation\">)</span><span class=\"token punctuation\">}</span><span class=\"token punctuation\">]</span>\n    <span class=\"token keyword\">for</span> attempt <span class=\"token keyword\">in</span> <span class=\"token builtin\">range</span><span class=\"token punctuation\">(</span>retries <span class=\"token operator\">+</span> <span class=\"token number\">1</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">:</span>\n        raw <span class=\"token operator\">=</span> call_model<span class=\"token punctuation\">(</span>messages<span class=\"token punctuation\">,</span> schema<span class=\"token operator\">=</span>Classification<span class=\"token punctuation\">.</span>model_json_schema<span class=\"token punctuation\">(</span><span class=\"token punctuation\">)</span><span class=\"token punctuation\">)</span>\n        <span class=\"token keyword\">try</span><span class=\"token punctuation\">:</span>\n            <span class=\"token keyword\">return</span> Classification<span class=\"token punctuation\">.</span>model_validate_json<span class=\"token punctuation\">(</span>raw<span class=\"token punctuation\">)</span>\n        <span class=\"token keyword\">except</span> ValidationError <span class=\"token keyword\">as</span> e<span class=\"token punctuation\">:</span>\n            <span class=\"token keyword\">if</span> attempt <span class=\"token operator\">==</span> retries<span class=\"token punctuation\">:</span>\n                <span class=\"token keyword\">raise</span>\n            <span class=\"token comment\"># Feed the validation error back so the model can self-correct.</span>\n            messages<span class=\"token punctuation\">.</span>append<span class=\"token punctuation\">(</span><span class=\"token punctuation\">{</span><span class=\"token string\">\"role\"</span><span class=\"token punctuation\">:</span> <span class=\"token string\">\"assistant\"</span><span class=\"token punctuation\">,</span> <span class=\"token string\">\"content\"</span><span class=\"token punctuation\">:</span> raw<span class=\"token punctuation\">}</span><span class=\"token punctuation\">)</span>\n            messages<span class=\"token punctuation\">.</span>append<span class=\"token punctuation\">(</span><span class=\"token punctuation\">{</span>\n                <span class=\"token string\">\"role\"</span><span class=\"token punctuation\">:</span> <span class=\"token string\">\"user\"</span><span class=\"token punctuation\">,</span>\n                <span class=\"token string\">\"content\"</span><span class=\"token punctuation\">:</span> <span class=\"token string-interpolation\"><span class=\"token string\">f\"That failed validation:\\n</span><span class=\"token interpolation\"><span class=\"token punctuation\">{</span>e<span class=\"token punctuation\">}</span></span><span class=\"token string\">\\nReturn corrected JSON.\"</span></span><span class=\"token punctuation\">,</span>\n            <span class=\"token punctuation\">}</span><span class=\"token punctuation\">)</span></code></pre></div>\n<p>Two things make this work:</p>\n<ul>\n<li><strong>Full validation.</strong> <code class=\"language-text\">model_validate_json</code> enforces the <code class=\"language-text\">ge</code>/<code class=\"language-text\">le</code> bounds and the enum, catching the <code class=\"language-text\">priority: 9</code> that slipped past constrained decoding.</li>\n<li><strong>Error feedback.</strong> The retry hands the model its own bad output <em>plus the specific validation error</em>. Models are good at fixing a mistake when you show them exactly what broke, far better than at getting it right blind.</li>\n</ul>\n<p>One retry catches the large majority of the residual failures. A second retry has sharply diminishing returns and mostly just costs you latency and tokens.</p>\n<p>If you already return a typed object from your API layer, this validation step is nearly free to add, because you were going to construct that object anyway. You are just constructing it through the validator instead of around it.</p>\n<p><img src=\"/e2dd7e90cf09dda9c7be5b41eff86350/validate-retry-loop.svg\" alt=\"The validate-and-retry loop. The model returns a candidate, the typed schema validates it, and on success the object flows downstream. On a validation failure, the parser error is appended to the conversation and the model gets one corrected attempt before the call is allowed to raise.\"></p>\n<h2>Failure modes that survive all three layers</h2>\n<p>Even with constrained decoding and a validation loop, a few things still go wrong. These are the ones worth knowing before they page you.</p>\n<p><strong>Truncation at the token limit.</strong> If the model hits <code class=\"language-text\">max_tokens</code> mid-object, you get a valid <em>prefix</em> that is not valid JSON: an open brace with no close. Constrained decoding does not save you here, because the stop came from outside the grammar. Check the finish reason. If it is <code class=\"language-text\">max_tokens</code> (or <code class=\"language-text\">length</code>), treat it as a failure and either raise or retry with a higher limit; do not try to parse the fragment. Give large or deeply nested outputs generous headroom.</p>\n<p><strong>Refusals.</strong> If the model declines a request for safety reasons, the response is a refusal, not your schema. Providers signal this with a <code class=\"language-text\">refusal</code> stop reason on the response. You have to branch on it <em>before</em> you try to parse, or you will feed a polite apology into <code class=\"language-text\">json.loads()</code>. This matters more when the text being classified is itself untrusted, as with anything downstream of a RAG (retrieval-augmented generation) crawl.</p>\n<p><strong>First-request compilation lag.</strong> Provider strict modes compile your schema into a grammar the first time they see it, and then cache it (Anthropic caches for 24 hours). A brand-new or frequently changing schema pays that compile cost on the first call, which shows up as a latency spike. Keep schemas stable and let the cache work. Do not regenerate a structurally identical schema with reordered keys on every request.</p>\n<p><strong>Over-constraining the model into worse answers.</strong> This is the subtle one. Forcing a rigid schema can <em>lower</em> the quality of the content inside it, because you have taken away the model’s room to reason. A classifier that must emit <code class=\"language-text\">{&quot;category&quot;: ...}</code> as its very first tokens has no space to think first. If accuracy matters, give it a scratchpad: add a <code class=\"language-text\">reasoning</code> string field <em>before</em> the answer fields in the schema, or let it produce free text and make a cheap second call to structure that text. Order matters, because a field the model fills in first cannot depend on a field that comes later.</p>\n<p><strong>Enums that drift with the world.</strong> An enum locks the model to a fixed set of values, which is exactly what you want for a router, until the set changes. If categories come from a database, generate the enum from that source at startup rather than hardcoding it. Otherwise, the day someone adds a category is the day the model can never route to it.</p>\n<h2>What I would reach for, and when</h2>\n<p>For anything where my code consumes the output, I now turn on the provider’s constrained decoding by default. It is a request parameter, the cost is near zero, and it deletes an entire class of parse errors that used to be the top line in the error dashboard. Then I validate with the real schema, because constrained decoding is a syntax guarantee, not a semantic one, and the bounds it skips are usually the ones that matter. The one-retry-with-the-error loop is the cheap insurance on top.</p>\n<p>The judgment call is layer two on self-hosted models. Wiring up Outlines or GBNF grammars is real work. If you control the model and throughput matters, it is worth it. If you are calling a hosted API, the same guarantee is one parameter away, so there is no reason not to use it.</p>\n<p>None of this is exotic. It is the same discipline you would apply to any untrusted input crossing a boundary into your system: constrain what can come in, validate it against a typed contract, and have a defined path for when it is wrong. An LLM is just an unusually fluent source of untrusted input. The same approach sits behind the <a href=\"/blog/2026-06-30-llm-agent-tool-loop/\">agent tool loop</a>, the structured fields in the <a href=\"/project/archi/\">Archi copilot</a>, and the token accounting in <a href=\"/project/llm-dev-mate/\">LLM DevMate</a>: anywhere the model’s answer has to be data before it can be useful.</p>\n<hr>\n<p><em>Diagrams by M. Hassan Ahmed, released under CC0. Image credit: original work by the author.</em></p>","frontmatter":{"title":"Getting Reliable JSON Out of an LLM","date":"2026-07-14T00:00:00.000Z","description":"Getting an LLM to return JSON is easy; getting valid JSON every time is not. How to use JSON Schema, constrained decoding, and validation to make it reliable.","thumbnail":{"childImageSharp":{"fluid":{"base64":"data:image/png;base64,iVBORw0KGgoAAAANSUhEUgAAABQAAAALCAIAAADwazoUAAAACXBIWXMAAAsSAAALEgHS3X78AAAByElEQVQoz12RW4/aMBCFs2hhA7GdCyHO1XbiBAgJNIE0yyIQsCtVVaW+tE997p/pQ5/6izskq6220qcja3zOzFhWCi57jlVzbffXj/vn9um8a18eD8ClaY/V7tI8nqrm+GHX1091s2IJRJSRzh50BurFVb69LOvLojonxWFZnfP6AoekPEAxWZ9Evk/Xx3x7jebNkESQUpDJeiZ6CIz1EE1jgy4MJ9OdFNRy5ybNXDfF01glQYffRxTNiN7oS2CKsy1PaxqVNCpcVlBWsDDDVmw60pjFyOK9X+kH3maSgNixPkuQGXl+6riS2BJZorsN1G4pjxW2N0cm7yMQDl6BZSwYy4fwGMzvCLvD4QCzoc41k2ngxv5oQlXkTTCYbxEFMm9oRgBrp4HYJbJO0g3nteBbwbHho2mkRQkRqRYIYGJH0EsZE6+ny3uaxX+2iz9f1r8+lb8/g65/1HPLDBzK5KZcNVuRL/xMmh5TsauMsfvKLe9NjPBJiu8b+bWQ39byZREPxhSubDNYOaL05JLGiRXMDL8P0/+AP783+L3OBwQezMyZMGxBbIGmfKQ5I0QfsKt2TkVFzjuwoxEX6S66KYXzu9YI9J/5LxqzRjEZv3uyAAAAAElFTkSuQmCC","aspectRatio":1.899441340782123,"src":"/static/c4ceb8e65853708427d78040856e38b7/40a76/hero.png","srcSet":"/static/c4ceb8e65853708427d78040856e38b7/c972b/hero.png 340w,\n/static/c4ceb8e65853708427d78040856e38b7/27625/hero.png 680w,\n/static/c4ceb8e65853708427d78040856e38b7/40a76/hero.png 1360w,\n/static/c4ceb8e65853708427d78040856e38b7/ed396/hero.png 2000w","sizes":"(max-width: 1360px) 100vw, 1360px"}}}}}},"pageContext":{"slug":"/2026-07-14-reliable-json-from-llms/","previous":"blog/2026-07-11-llm-api-rate-limits-retries-backoff/","next":"blog/2026-07-13-streaming-markdown-react-llm/"}},"staticQueryHashes":["32046230"]}