Getting Reliable JSON Out of an LLM

Getting an LLM to return JSON is easy; getting valid JSON every time is not. How to use JSON Schema, constrained decoding, and validation to make it reliable.

The first version always works. You write a prompt that ends with “respond with JSON like {"intent": "...", "confidence": 0.0}”, call json.loads() on the reply, and it parses. You ship it. A week later, the model answers some request with “Sure! Here is the JSON you asked for:” followed by a code fence. json.loads() throws, and the endpoint 500s.

That is the whole problem in one sentence: the model produces text, and text that is mostly JSON is not JSON. The gap between the two is where the on-call pages come from.

This post is for backend engineers who use an LLM as a component in a larger system: a classifier that returns a label, an extractor that pulls fields out of a document, a router that decides which tool to call. You are not reading the model’s output; your code is. When your code is the consumer, “close enough” is a bug.

I hit every failure mode here while wiring the structured pieces of the Archi copilot at CERN and the token-aware context bundling in LLM DevMate. The fixes fall into three layers that stack on top of each other: describe the output as a schema, constrain the decoding so it always parses, and validate what comes back. This post walks through all three, and the ways each one still bites.

Why “ask nicely” is not enough

The naive version puts the whole burden on the prompt, then trusts json.loads():

prompt = """Classify the ticket below. Respond with JSON:
{"category": "bug" | "feature" | "question", "priority": 1-5}

Ticket: The export button does nothing on Safari."""

resp = call_model(prompt)
data = json.loads(resp)   # <-- this line is the whole risk

This works often enough to feel done, and that is the trap. The failures are not random noise. They are specific, and they cluster:

  • Preamble. Here is the classification: before the object, or a trailing Let me know if you need anything else. after it. A Markdown code fence (```json) wrapped around the whole thing is the most common variant.
  • Almost-valid JSON. A trailing comma, single quotes instead of double, None/True instead of null/true because the model leaned Python, or an unescaped newline inside a string.
  • Right shape, wrong types. "priority": "high" when your schema said an integer 1 to 5, or "confidence": "0.9" as a string. These parse (json.loads() is happy) and then blow up three functions later when you do arithmetic on a string.
  • Extra or missing keys. A hallucinated "reason" field you never asked for, or a dropped "priority" on a ticket the model found ambiguous.

The last two are the dangerous ones, because parsing succeeds. A try/except json.JSONDecodeError catches none of it. Valid JSON of the wrong shape flows into code that assumed the right shape, and the error surfaces far from its cause.

Three ways to turn a model into structured JSON. Prompt-and-parse relies on the prompt and often produces text that is only nearly JSON, so the parser throws. Constrained decoding masks the token sampler to a grammar, so the string always parses but types can still drift. Validate-and-retry parses, checks against a typed schema, and feeds any error back for one corrected attempt. The reliable path uses the last two together.

Layer one: describe the shape as a schema, not a sentence

The first fix is to stop describing the output in prose and start describing it as JSON Schema. A schema is a machine-readable contract: field names, types, which fields are required, and which values an enum allows. In Python, the ergonomic way to write one is a Pydantic model, which gives you the schema and the validator in the same object. This model limits category to three values and priority to an integer from 1 to 5:

from enum import Enum
from pydantic import BaseModel, Field

class Category(str, Enum):
    bug = "bug"
    feature = "feature"
    question = "question"

class Classification(BaseModel):
    category: Category
    priority: int = Field(ge=1, le=5)

# Pydantic hands you the JSON Schema for free:
schema = Classification.model_json_schema()

On its own, putting the schema in the prompt is only a marginal improvement over the prose version, because the model is still free to ignore it. The schema earns its place because the next two layers both consume it. It is the single source of truth: the decoder constrains against it, and the validator checks against it. You write the shape once, and both enforcement points read the same definition.

Layer two: constrain the decoding so the string always parses

The strongest fix does not live in the prompt at all. It lives in how the tokens are sampled.

A language model generates one token at a time, drawing each from a probability distribution over its whole vocabulary. Constrained decoding puts a filter in front of that draw. At each step, it works out which tokens could still lead to a string that satisfies the schema, sets the probability of every other token to zero, and samples only from what remains. If the grammar says the next character must be a } or a ,, the model physically cannot emit the word “Sure”. The output is guaranteed to parse, because a token that would break the JSON is never on the table.

You rarely implement this yourself, because the providers expose it. The mechanism is the same one behind tool calling: a tool’s input schema is enforced exactly this way, which is why a well-defined tool call comes back as clean arguments and a free-text answer does not. Here is where each option lives:

  • OpenAI exposes it as Structured Outputs. Pass your schema in response_format with strict: true, and the response is guaranteed to match it.
  • Anthropic does the same through structured outputs and strict tool use. Use output_config={"format": {"type": "json_schema", "schema": schema}} on the Messages API, or strict: true on a tool definition, with client.messages.parse() validating the response against your model automatically.
  • Self-hosted models get it from libraries like Outlines or llama.cpp’s GBNF grammars, which compile a schema or grammar directly into the token mask.

The technique is grounded research, not a vendor gimmick. The Outlines paper on efficient guided generation shows how to build the token masks from a regular expression or grammar with near-zero per-token overhead.

Turning it on removes the entire first category of failures: no preamble, no code fence, no trailing comma, ever. What it does not remove is type and value drift, and that surprises people. The reason is that the schema a provider’s strict mode enforces is not the full Pydantic schema you wrote.

The catch: strict mode enforces only part of JSON Schema

Constrained decoding only enforces what the grammar can express, and the grammar covers a subset of JSON Schema. Anthropic’s strict mode, for example, does not enforce numeric bounds like minimum/maximum or string bounds like minLength, and it requires every object to set additionalProperties: false. OpenAI’s strict mode has a similar list of unsupported keywords.

So what happens to the Field(ge=1, le=5) you wrote on priority? The decoder guarantees you get an integer. It does not guarantee the integer is between 1 and 5. The model can still return priority: 9, and it will be a syntactically perfect, schema-shaped 9.

Constrained decoding buys you a string that parses and a value of the right type. The semantics are still yours to check: the range, the cross-field rules, the “this enum value only makes sense when that flag is set” checks. That is the whole reason the third layer exists.

Layer three: validate what came back, and retry once with the error

Whatever the decoder gave you, run it through the full validator before anything downstream touches it. This is where Pydantic pays off a second time: the same model that generated the schema now checks the bounds the decoder ignored. The function below validates each reply and, if validation fails, sends the error back to the model for one more attempt:

from pydantic import ValidationError

def classify(ticket: str, retries: int = 1) -> Classification:
    messages = [{"role": "user", "content": build_prompt(ticket)}]
    for attempt in range(retries + 1):
        raw = call_model(messages, schema=Classification.model_json_schema())
        try:
            return Classification.model_validate_json(raw)
        except ValidationError as e:
            if attempt == retries:
                raise
            # Feed the validation error back so the model can self-correct.
            messages.append({"role": "assistant", "content": raw})
            messages.append({
                "role": "user",
                "content": f"That failed validation:\n{e}\nReturn corrected JSON.",
            })

Two things make this work:

  • Full validation. model_validate_json enforces the ge/le bounds and the enum, catching the priority: 9 that slipped past constrained decoding.
  • Error feedback. The retry hands the model its own bad output plus the specific validation error. Models are good at fixing a mistake when you show them exactly what broke, far better than at getting it right blind.

One retry catches the large majority of the residual failures. A second retry has sharply diminishing returns and mostly just costs you latency and tokens.

If you already return a typed object from your API layer, this validation step is nearly free to add, because you were going to construct that object anyway. You are just constructing it through the validator instead of around it.

The validate-and-retry loop. The model returns a candidate, the typed schema validates it, and on success the object flows downstream. On a validation failure, the parser error is appended to the conversation and the model gets one corrected attempt before the call is allowed to raise.

Failure modes that survive all three layers

Even with constrained decoding and a validation loop, a few things still go wrong. These are the ones worth knowing before they page you.

Truncation at the token limit. If the model hits max_tokens mid-object, you get a valid prefix that is not valid JSON: an open brace with no close. Constrained decoding does not save you here, because the stop came from outside the grammar. Check the finish reason. If it is max_tokens (or length), treat it as a failure and either raise or retry with a higher limit; do not try to parse the fragment. Give large or deeply nested outputs generous headroom.

Refusals. If the model declines a request for safety reasons, the response is a refusal, not your schema. Providers signal this with a refusal stop reason on the response. You have to branch on it before you try to parse, or you will feed a polite apology into json.loads(). This matters more when the text being classified is itself untrusted, as with anything downstream of a RAG (retrieval-augmented generation) crawl.

First-request compilation lag. Provider strict modes compile your schema into a grammar the first time they see it, and then cache it (Anthropic caches for 24 hours). A brand-new or frequently changing schema pays that compile cost on the first call, which shows up as a latency spike. Keep schemas stable and let the cache work. Do not regenerate a structurally identical schema with reordered keys on every request.

Over-constraining the model into worse answers. This is the subtle one. Forcing a rigid schema can lower the quality of the content inside it, because you have taken away the model’s room to reason. A classifier that must emit {"category": ...} as its very first tokens has no space to think first. If accuracy matters, give it a scratchpad: add a reasoning string field before the answer fields in the schema, or let it produce free text and make a cheap second call to structure that text. Order matters, because a field the model fills in first cannot depend on a field that comes later.

Enums that drift with the world. An enum locks the model to a fixed set of values, which is exactly what you want for a router, until the set changes. If categories come from a database, generate the enum from that source at startup rather than hardcoding it. Otherwise, the day someone adds a category is the day the model can never route to it.

What I would reach for, and when

For anything where my code consumes the output, I now turn on the provider’s constrained decoding by default. It is a request parameter, the cost is near zero, and it deletes an entire class of parse errors that used to be the top line in the error dashboard. Then I validate with the real schema, because constrained decoding is a syntax guarantee, not a semantic one, and the bounds it skips are usually the ones that matter. The one-retry-with-the-error loop is the cheap insurance on top.

The judgment call is layer two on self-hosted models. Wiring up Outlines or GBNF grammars is real work. If you control the model and throughput matters, it is worth it. If you are calling a hosted API, the same guarantee is one parameter away, so there is no reason not to use it.

None of this is exotic. It is the same discipline you would apply to any untrusted input crossing a boundary into your system: constrain what can come in, validate it against a typed contract, and have a defined path for when it is wrong. An LLM is just an unusually fluent source of untrusted input. The same approach sits behind the agent tool loop, the structured fields in the Archi copilot, and the token accounting in LLM DevMate: anywhere the model’s answer has to be data before it can be useful.


Diagrams by M. Hassan Ahmed, released under CC0. Image credit: original work by the author.