Video · 8 min read
Getting Reliable JSON from LLMs: The 2026 Guide to Structured Outputs
Stop fighting with LLM outputs. Learn 7 proven techniques for guaranteed JSON schema compliance from OpenAI, Anthropic, and Google APIs.
The Bottom Line
Getting reliable, structured data from LLMs has evolved from a frustrating prompt-engineering exercise into a solved problem—if you know the right techniques.
Without constraints, asking an LLM for JSON fails 30-70% of the time on complex schemas. With native structured outputs, you get 100% schema compliance. The difference? Understanding when to use guaranteed constraints versus best-effort approaches.
Why This Matters
Modern applications need LLMs to power APIs, populate databases, and integrate with existing systems. A JSON parsing failure at 2 AM cascades into customer-facing outages.
The core problem: LLMs generate text token-by-token based on probability distributions. They're trained to produce helpful, conversational responses—not machine-parseable data structures.
Two fundamental approaches exist:
| Approach | Guarantee | Best For |
|---|---|---|
| Guaranteed constraints | 100% schema compliance | Production systems |
| Best-effort constraints | ~70-85% compliance | Prototyping, models without native support |
Guaranteed constraints modify token generation itself—invalid tokens are mathematically impossible. Best-effort approaches guide through prompting and hope the model complies.
Provider-Native Structured Outputs
All three major AI providers now offer built-in structured output features. Here's how they compare:
Comparison Table
| Feature | OpenAI | Anthropic | Gemini |
|---|---|---|---|
| Release date | Aug 2024 | Nov 2025 | Nov 2025 (enhanced) |
Union types (anyOf) | No | No | Yes |
| Recursive schemas | No | No | Yes |
| Numeric constraints | No | No | Yes |
| Property ordering | No | No | Yes (Gemini 2.5+) |
| Streaming support | Yes | Yes | Yes |
OpenAI: Most Mature
OpenAI reports 100% schema compliance vs ~35% with prompting alone. Works through constrained decoding—the API masks out tokens that would violate your schema.
from openai import OpenAI
from pydantic import BaseModel
class CalendarEvent(BaseModel):
name: str
date: str
participants: list[str]
client = OpenAI()
completion = client.beta.chat.completions.parse(
model="gpt-4o-2024-08-06",
messages=[
{"role": "system", "content": "Extract event information."},
{"role": "user", "content": "Alice and Bob are going to a science fair on Friday."}
],
response_format=CalendarEvent
)
event = completion.choices[0].message.parsed # Already a CalendarEvent object
Limitations: No anyOf/oneOf, no recursive schemas, no numeric constraints. All fields must be required.
Anthropic Claude: Newer but Capable
Released November 2025 as public beta. Uses compiled grammar artifacts for enforcement.
from anthropic import Anthropic
client = Anthropic()
response = client.beta.messages.create(
model="claude-sonnet-4-5",
betas=["structured-outputs-2025-11-13"],
max_tokens=1024,
messages=[
{"role": "user", "content": "Extract: John Smith (john@email.com) wants a demo."}
],
output_format={
"type": "json_schema",
"schema": {
"type": "object",
"properties": {
"name": {"type": "string"},
"email": {"type": "string"},
"demo_requested": {"type": "boolean"}
},
"required": ["name", "email", "demo_requested"]
}
}
)
Note: Requires beta header (anthropic-beta: structured-outputs-2025-11-13).
Google Gemini: Most Flexible
Most advanced JSON Schema support. Unique features: anyOf for union types, $ref for recursive schemas, minimum/maximum for numeric constraints.
from google import genai
from pydantic import BaseModel
class Recipe(BaseModel):
recipe_name: str
ingredients: list[str]
instructions: list[str]
client = genai.Client()
response = client.models.generate_content(
model="gemini-2.5-flash",
contents="Extract the recipe from: Pancakes - mix flour, eggs, milk. Cook on griddle.",
config={
"response_mime_type": "application/json",
"response_json_schema": Recipe.model_json_schema()
}
)
7 Techniques for Constraining Output
1. JSON Mode vs Structured Outputs
Don't confuse them:
| Mode | Guarantees |
|---|---|
| JSON mode | Valid JSON syntax only |
| Structured outputs | Valid JSON AND exact schema compliance |
# JSON mode - might return {"status": "ok"} when you expected {"name": "...", "age": ...}
response_format={"type": "json_object"}
# Structured outputs - guarantees exact schema
response_format={"type": "json_schema", "json_schema": {...}}
2. Schema Design That Works
Well-designed schemas dramatically improve reliability:
from pydantic import BaseModel, Field
from typing import Literal, Optional
from enum import Enum
class SentimentLevel(str, Enum):
positive = "positive"
negative = "negative"
neutral = "neutral"
class ProductReview(BaseModel):
"""Analysis of a customer product review."""
product_name: str = Field(
description="The product being reviewed, exactly as mentioned"
)
sentiment: SentimentLevel = Field(
description="Overall emotional tone of the review"
)
rating_inferred: Optional[int] = Field(
default=None,
description="Estimated star rating 1-5 if determinable, null otherwise"
)
key_points: list[str] = Field(
description="Main points mentioned, maximum 5 items",
max_length=5
)
Key principles:
- Use descriptive field names
- Add
descriptionattributes to clarify ambiguous fields - Use enums or
Literaltypes to constrain categories - Keep nesting shallow—deeply nested schemas have higher failure rates
3. Regex Constraints
For simple patterns like emails, dates, classifications:
import outlines
model = outlines.models.transformers("mistralai/Mistral-7B-Instruct-v0.2")
# Only these exact words are possible outputs
classifier = outlines.generate.regex(model, r"(positive|negative|neutral)")
sentiment = classifier("Classify this review: 'Amazing product!' Sentiment:")
# Output: "positive" (guaranteed)
4. Grammar-Based Constraints
For nested structures, recursion, and code generation:
# GBNF grammar for JSON arrays (llama.cpp format)
grammar = """
root ::= "[" ws (object ("," ws object)*)? ws "]"
object ::= "{" ws "\"name\":" ws string "," ws "\"value\":" ws number ws "}"
string ::= "\"" ([^"\\] | "\\" .)* "\""
number ::= "-"? [0-9]+ ("." [0-9]+)?
ws ::= [ \t\n]*
"""
5. Few-Shot Examples
When you can't use constrained decoding:
Extract product information as JSON.
Example 1:
Input: "iPhone 15 Pro - $999, 256GB storage, Space Black"
Output: {"name": "iPhone 15 Pro", "price": 999, "storage": "256GB", "color": "Space Black"}
Example 2:
Input: "Samsung Galaxy S24 Ultra priced at $1199 with 512GB"
Output: {"name": "Samsung Galaxy S24 Ultra", "price": 1199, "storage": "512GB", "color": null}
Now extract:
Input: "OnePlus 12 - 256GB Flowy Emerald edition for $799"
Output:
Best practices: Use 2-5 examples. Cover edge cases. Place the most important example last.
6. Explicit Format Instructions
Clear, specific instructions significantly improve compliance:
system_prompt = """You are a data extraction assistant.
CRITICAL RULES:
1. Return ONLY valid JSON - no markdown, no explanations
2. Use exactly these field names: name, age, email, skills
3. For missing information, use null (not "unknown")
4. age must be an integer, not a string
5. skills must be an array, even if there's only one skill
Expected structure:
{
"name": "<full name as string>",
"age": <integer or null>,
"email": "<email address or null>",
"skills": ["<skill1>", "<skill2>"]
}"""
7. Why Negative Constraints Often Fail
"Don't do X" instructions can make unwanted behavior more likely (the "pink elephant problem"):
# Less effective
prompt = "Summarize this. Don't include bullet points. Don't exceed 200 words."
# More effective - positive framing
prompt = "Summarize this in 2-3 flowing paragraphs. Target 150 words. Use prose format."
Tools That Make This Practical
| Tool | Best For | Monthly Downloads |
|---|---|---|
| Instructor | API models, multi-provider | 3M+ |
| Outlines | Local models, guaranteed compliance | - |
| LangChain | When already in LangChain ecosystem | - |
| Guidance | Token-level control, research | - |
Instructor Example
import instructor
from pydantic import BaseModel
class User(BaseModel):
name: str
age: int
# Automatically retries on validation failure
client = instructor.from_provider("openai/gpt-4o")
user = client.chat.completions.create(
response_model=User,
max_retries=3,
messages=[{"role": "user", "content": "Extract: Jason is 30 years old"}]
)
Outlines Example
import outlines
from pydantic import BaseModel
class Character(BaseModel):
name: str
age: int
armor: str
model = outlines.models.transformers("microsoft/Phi-3-mini-128k-instruct")
generator = outlines.generate.json(model, Character)
character = generator("Generate a fantasy RPG character:")
# Guaranteed to be a valid Character instance
When Things Go Wrong
The Truncation Trap
The model hits max_tokens before completing JSON:
def safe_llm_call(prompt: str, schema: dict):
response = client.chat.completions.create(...)
# ALWAYS check finish_reason
if response.choices[0].finish_reason == "length":
raise ValueError("Response truncated - increase max_tokens")
return response.choices[0].message.content
Schema Compliance ≠ Content Accuracy
Structured outputs guarantee format, not truth. Always validate content:
from pydantic import field_validator
class ExtractedData(BaseModel):
company_name: str
founded_year: int
@field_validator('founded_year')
def reasonable_year(cls, v):
if v < 1800 or v > 2026:
raise ValueError(f'Year {v} seems unlikely')
return v
Graceful Degradation Pattern
def extract_with_fallback(text: str) -> dict:
# Tier 1: Strict schema
try:
return call_with_strict_schema(text).model_dump()
except ValidationError:
pass
# Tier 2: Relaxed schema
try:
return call_with_relaxed_schema(text).model_dump()
except ValidationError:
pass
# Tier 3: JSON mode without schema
try:
return json.loads(call_with_json_mode(text))
except json.JSONDecodeError:
pass
# Tier 4: Return partial/default data
return {"raw_text": text, "extraction_failed": True}
Key Takeaways
For API-based applications:
- Use native structured outputs (OpenAI, Anthropic, or Gemini)
- Combine with Instructor for automatic retries
- Always check
finish_reasonfor truncation
For local model deployment:
- Use Outlines for guaranteed compliance
- Consider grammar constraints for code generation
For all cases:
- Design schemas simply with clear field descriptions
- Implement retry logic with exponential backoff
- Build graceful degradation paths
- Test edge cases before production
The tools exist. "The LLM returned invalid JSON" is no longer an acceptable production failure mode.
Building real-estate growth tools without a Frankenstack? Talk to us.
Keep going