# Making a model emit valid JSON, every time > Asking nicely gets you valid JSON almost always, and almost is a broken integration. The fix is in the sampler: delete every invalid token before you sample. Source: https://rfriedmann.de/blog/constrained-decoding-valid-json/ Published: 2026-09-16 · Track: log · Level: Expert Ask a good model for JSON and you get JSON, maybe 98 times in a hundred. The other two come wrapped in a markdown fence, or with a cheerful "Here you go!" in front, or with a trailing comma, or with the whole object fine except one string that never closed because generation hit the token limit. If that output feeds a parser, you do not have a 98% success rate. You have a pager. No amount of prompt engineering closes that gap, because the gap is structural: the model is a probability distribution over the next token, and every token with non-zero probability will eventually come up. The fix is not to ask better. It is to make the invalid tokens unrepresentable. ## The sampler is a place you are allowed to edit Recall what a decode step actually produces: a vector of logits, one real number per token in the vocabulary. Those get shaped by temperature, top-k and top-p, turned into probabilities, and one token is drawn. That is where [the dice roll lives](/blog/why-different-answers/). Constrained decoding slots one operation in ahead of all of that. Before the softmax, set the logit of every token the grammar forbids to negative infinity. After the softmax those tokens hold exactly zero probability, and no sampler, at any temperature, can pick them.
The mask is arithmetic, not persuasion
[diagram omitted — see the page for the chart]
Six of roughly 150,000 candidates. The highest-scoring one is forbidden: a quote, the model about to open the next key and forget the comma in front of it. That is the near-miss a parser would have thrown out, and no prompt would have known to prevent it. After the mask, all the probability mass redistributes across the two legal tokens.
That is the whole idea, and the guarantee it buys is absolute rather than statistical. Output that violates the schema is not unlikely; it is unreachable. ## The awkward part: grammars are over characters, models emit tokens If a model emitted one character at a time, this would be a tutorial exercise. Build a finite state machine for the grammar, ask it which characters are legal in the current state, mask the rest. But a model does not see characters, [it sees tokens](/blog/why-models-cant-spell/), and a token is an arbitrary byte string chosen by the tokenizer for its frequency, with no respect whatsoever for your grammar's boundaries. Real vocabularies contain tokens like `",`, `"}`, `": "` and `\n "`, single tokens that carry a string terminator, a structural character and the start of the next key all at once. So the question a mask has to answer is not "which characters may come next" but: > from this automaton state, which of the 150,000 token strings can I feed through > the automaton, byte by byte, without ever falling out of the language? Answer that naively, per step, per sequence in the batch, and you are walking the whole vocabulary through a state machine between every forward pass — on the CPU, while a very expensive GPU sits idle waiting for a token. The production answer is precomputation. Compile the schema into an automaton, then build an index from automaton state to the set of allowed token IDs, so that at decode time the mask is a lookup rather than a search. Regular structures (a regex, an enum, a date format) compile to a plain FSM. JSON needs a little more, because nesting is not regular: you carry a stack, or at least a depth counter, which makes it a pushdown automaton with a bounded stack in practice. Most of the transitions are shared across states, and most tokens are decidable without knowing the stack at all, which is what the current generation of libraries exploits to keep the per-step cost down to microseconds. ## What it costs an inference engine A mask over a 151,936-token vocabulary, held as a bitmask, is 18.5 KiB. Per sequence. Per step. That number is small enough to be fine and large enough to matter, and it lands in the middle of every fast path the engine has. **It is per sequence, not per batch.** With [continuous batching](/blog/continuous-batching/) every stream in flight sits in a different grammar state, so the mask is a matrix with one row per sequence, rebuilt every step as each one advances. **It wants to stall the pipeline.** The mask depends on the token you just sampled, which you only know after the forward pass finished. Done in the obvious order, every step gains a host round trip: copy the token down, advance the automaton, build the mask, copy it up. Engines hide this by overlapping the mask construction for step *n* with the forward pass of step *n*, and by keeping the mask in a fixed device buffer that is overwritten in place rather than reallocated, so a captured CUDA graph stays replayable. The same discipline that keeps [MoE routing from breaking the graph](/blog/mixture-of-experts-explained/) applies here. **It complicates speculation.** [Speculative decoding](/blog/speculative-decoding/) proposes several tokens before any of them are verified, so the grammar has to be advanced speculatively too, and rolled back for every rejected token. The bookkeeping is real, and an engine that gets it subtly wrong produces output that is valid right up until the first rejection. None of this is exotic. `llama.cpp` has shipped GBNF grammars for years, vLLM and friends wire in XGrammar or similar, Outlines does the regex-to-FSM compilation, and the strict-schema modes of the big provider APIs are this mechanism running on their side of the wire. It is standard equipment. It is just not free. ## Valid is not correct Now the caveats, which are the reason this post exists rather than a one-liner. **A schema constrains shape, not truth.** The mask guarantees that `"port": 8080` parses and that `port` is an integer. It has nothing to say about whether 8080 is the right port. You have replaced a parse error, which is loud and obvious, with a type-correct wrong value, which is silent. That is a better failure mode to build on but a worse one to notice, and it needs the same [verification](/blog/when-not-to-use-ai/) everything else does. **Masking moves the distribution, and sometimes somewhere worse.** Forcing a token the model rated unlikely puts it in a context its training rarely produced, and quality can degrade from there. The classic own-goal is a schema whose first field is `answer` and whose second is `reasoning`. The model is now required to commit to the answer before it has written a word of the working-out, which is precisely the [computation it was going to buy with those tokens](/blog/why-models-think-before-answering/). Put the free-text field first and the effect reverses. Field names and ordering are part of the prompt, whether or not you were treating them that way. **The grammar cannot stop you running out of tokens.** Nothing in the automaton prevents a 900-token string value in an object with a 1,000-token budget. You still need a limit policy: force the closing tokens when the budget runs low, or treat a truncated generation as a failed call. Silently handing a half-object to a retry loop is how an [agent](/blog/how-an-agent-uses-tools-and-memory/) ends up burning its context on the same broken call. The reason I keep coming back to this one is that it is such a clean example of where the leverage actually is. The model is a distribution; the engine owns the sampler; therefore an entire class of failure can be deleted with an addition of negative infinity, before it is ever a bug to handle. That is a much better bargain than any amount of asking the model to please, please return only JSON.