Saturday, October 3, 2026

 How Jev Works

Parallel constrained decoding is a superior way to make an LLM produce a structured JSON schema. The core idea is that instead of treating the schema as a sequence of tokens to be generated autoregressively — which forces the model to emit brackets, quotes, commas, and enum values one token at a time — you treat the schema as a set of independent fields whose values can be inferred in parallel from a single shared representation of the input document.

In the traditional autoregressive path, the model consumes your OCR’d document X as the prompt, then begins generating the JSON schema token by token. Even a tiny schema like {"risk_level": ..., "requires_review": ..., "action_tier": ...} requires dozens or hundreds of forward passes because each token depends on the previous one. This sequential dependency is slow, brittle, and prone to malformed JSON. Any hallucinated comma or missing quote breaks the entire output.

An alternative approach reframes the problem. You first “prefill” the model: run the Transformer decoder once over the concatenation of the context (your document X) and the JSON schema template. During this pass, every layer writes its keys and values into the KV cache. That cache now contains the full contextualized representation of both the document and the schema structure.

Once the prefill is complete, each field in the schema becomes a small, isolated classification problem. For a field like risk_level, you append a short suffix token sequence (e.g., the field name) and run a single forward pass using the cached KV values. The decoder produces a final hidden state — a dense embedding vector, shown as 1,536 dimensions in the diagram — which is fed through the shared language modeling head. This head produces logits over the entire vocabulary, but you immediately mask out everything except the valid tokens for that field: HIGH, MEDIUM, LOW, NONE. After applying softmax over just those candidates, the highest probability token becomes the field’s value. The same process applies to booleans like requires_review and enums like action_tier.

Because the KV cache is reused, each field evaluation is extremely cheap: one forward pass, no sequential dependency, no need to generate syntactic scaffolding. The model never has to “write” JSON; it only selects from allowed values. The result is fast, deterministic, and always syntactically valid. Prefill happens once, and all fields are resolved independently and in parallel. This yields near instant scoring of each field — for example, risk_level = HIGH (p = 0.99), requires_review = true (p = 1.00), action_tier = TIER_2 (p = 0.98) — without ever generating a malformed structure.

The engineering insight is that JSON schema generation can be reframed as constrained classification over a shared contextual embedding rather than free form text generation. By leveraging the Transformer’s KV cache and restricting the output space per field, you eliminate the fragility of autoregressive decoding and achieve predictable, high throughput structured inference.


No comments:

Post a Comment