Skip to content

Video · 8 min read

Getting Reliable JSON from LLMs: The 2026 Guide to Structured Outputs

Stop fighting with LLM outputs. Learn 7 proven techniques for guaranteed JSON schema compliance from OpenAI, Anthropic, and Google APIs.

The Bottom Line

Getting reliable, structured data from LLMs has evolved from a frustrating prompt-engineering exercise into a solved problem—if you know the right techniques.

Without constraints, asking an LLM for JSON fails 30-70% of the time on complex schemas. With native structured outputs, you get 100% schema compliance. The difference? Understanding when to use guaranteed constraints versus best-effort approaches.


Why This Matters

Modern applications need LLMs to power APIs, populate databases, and integrate with existing systems. A JSON parsing failure at 2 AM cascades into customer-facing outages.

The core problem: LLMs generate text token-by-token based on probability distributions. They're trained to produce helpful, conversational responses—not machine-parseable data structures.

Two fundamental approaches exist:

ApproachGuaranteeBest For
Guaranteed constraints100% schema complianceProduction systems
Best-effort constraints~70-85% compliancePrototyping, models without native support

Guaranteed constraints modify token generation itself—invalid tokens are mathematically impossible. Best-effort approaches guide through prompting and hope the model complies.


Provider-Native Structured Outputs

All three major AI providers now offer built-in structured output features. Here's how they compare:

Comparison Table

FeatureOpenAIAnthropicGemini
Release dateAug 2024Nov 2025Nov 2025 (enhanced)
Union types (anyOf)NoNoYes
Recursive schemasNoNoYes
Numeric constraintsNoNoYes
Property orderingNoNoYes (Gemini 2.5+)
Streaming supportYesYesYes

OpenAI: Most Mature

OpenAI reports 100% schema compliance vs ~35% with prompting alone. Works through constrained decoding—the API masks out tokens that would violate your schema.

Copyable text
from openai import OpenAI
from pydantic import BaseModel

class CalendarEvent(BaseModel):
    name: str
    date: str
    participants: list[str]

client = OpenAI()

completion = client.beta.chat.completions.parse(
    model="gpt-4o-2024-08-06",
    messages=[
        {"role": "system", "content": "Extract event information."},
        {"role": "user", "content": "Alice and Bob are going to a science fair on Friday."}
    ],
    response_format=CalendarEvent
)

event = completion.choices[0].message.parsed  # Already a CalendarEvent object

Limitations: No anyOf/oneOf, no recursive schemas, no numeric constraints. All fields must be required.

Anthropic Claude: Newer but Capable

Released November 2025 as public beta. Uses compiled grammar artifacts for enforcement.

Copyable text
from anthropic import Anthropic

client = Anthropic()

response = client.beta.messages.create(
    model="claude-sonnet-4-5",
    betas=["structured-outputs-2025-11-13"],
    max_tokens=1024,
    messages=[
        {"role": "user", "content": "Extract: John Smith (john@email.com) wants a demo."}
    ],
    output_format={
        "type": "json_schema",
        "schema": {
            "type": "object",
            "properties": {
                "name": {"type": "string"},
                "email": {"type": "string"},
                "demo_requested": {"type": "boolean"}
            },
            "required": ["name", "email", "demo_requested"]
        }
    }
)

Note: Requires beta header (anthropic-beta: structured-outputs-2025-11-13).

Google Gemini: Most Flexible

Most advanced JSON Schema support. Unique features: anyOf for union types, $ref for recursive schemas, minimum/maximum for numeric constraints.

Copyable text
from google import genai
from pydantic import BaseModel

class Recipe(BaseModel):
    recipe_name: str
    ingredients: list[str]
    instructions: list[str]

client = genai.Client()

response = client.models.generate_content(
    model="gemini-2.5-flash",
    contents="Extract the recipe from: Pancakes - mix flour, eggs, milk. Cook on griddle.",
    config={
        "response_mime_type": "application/json",
        "response_json_schema": Recipe.model_json_schema()
    }
)

7 Techniques for Constraining Output

1. JSON Mode vs Structured Outputs

Don't confuse them:

ModeGuarantees
JSON modeValid JSON syntax only
Structured outputsValid JSON AND exact schema compliance
Copyable text
# JSON mode - might return {"status": "ok"} when you expected {"name": "...", "age": ...}
response_format={"type": "json_object"}

# Structured outputs - guarantees exact schema
response_format={"type": "json_schema", "json_schema": {...}}

2. Schema Design That Works

Well-designed schemas dramatically improve reliability:

Copyable text
from pydantic import BaseModel, Field
from typing import Literal, Optional
from enum import Enum

class SentimentLevel(str, Enum):
    positive = "positive"
    negative = "negative"
    neutral = "neutral"

class ProductReview(BaseModel):
    """Analysis of a customer product review."""

    product_name: str = Field(
        description="The product being reviewed, exactly as mentioned"
    )
    sentiment: SentimentLevel = Field(
        description="Overall emotional tone of the review"
    )
    rating_inferred: Optional[int] = Field(
        default=None,
        description="Estimated star rating 1-5 if determinable, null otherwise"
    )
    key_points: list[str] = Field(
        description="Main points mentioned, maximum 5 items",
        max_length=5
    )

Key principles:

  • Use descriptive field names
  • Add description attributes to clarify ambiguous fields
  • Use enums or Literal types to constrain categories
  • Keep nesting shallow—deeply nested schemas have higher failure rates

3. Regex Constraints

For simple patterns like emails, dates, classifications:

Copyable text
import outlines

model = outlines.models.transformers("mistralai/Mistral-7B-Instruct-v0.2")

# Only these exact words are possible outputs
classifier = outlines.generate.regex(model, r"(positive|negative|neutral)")
sentiment = classifier("Classify this review: 'Amazing product!' Sentiment:")
# Output: "positive" (guaranteed)

4. Grammar-Based Constraints

For nested structures, recursion, and code generation:

Copyable text
# GBNF grammar for JSON arrays (llama.cpp format)
grammar = """
root ::= "[" ws (object ("," ws object)*)? ws "]"
object ::= "{" ws "\"name\":" ws string "," ws "\"value\":" ws number ws "}"
string ::= "\"" ([^"\\] | "\\" .)* "\""
number ::= "-"? [0-9]+ ("." [0-9]+)?
ws ::= [ \t\n]*
"""

5. Few-Shot Examples

When you can't use constrained decoding:

Copyable text
Extract product information as JSON.

Example 1:
Input: "iPhone 15 Pro - $999, 256GB storage, Space Black"
Output: {"name": "iPhone 15 Pro", "price": 999, "storage": "256GB", "color": "Space Black"}

Example 2:
Input: "Samsung Galaxy S24 Ultra priced at $1199 with 512GB"
Output: {"name": "Samsung Galaxy S24 Ultra", "price": 1199, "storage": "512GB", "color": null}

Now extract:
Input: "OnePlus 12 - 256GB Flowy Emerald edition for $799"
Output:

Best practices: Use 2-5 examples. Cover edge cases. Place the most important example last.

6. Explicit Format Instructions

Clear, specific instructions significantly improve compliance:

Copyable text
system_prompt = """You are a data extraction assistant.

CRITICAL RULES:
1. Return ONLY valid JSON - no markdown, no explanations
2. Use exactly these field names: name, age, email, skills
3. For missing information, use null (not "unknown")
4. age must be an integer, not a string
5. skills must be an array, even if there's only one skill

Expected structure:
{
    "name": "<full name as string>",
    "age": <integer or null>,
    "email": "<email address or null>",
    "skills": ["<skill1>", "<skill2>"]
}"""

7. Why Negative Constraints Often Fail

"Don't do X" instructions can make unwanted behavior more likely (the "pink elephant problem"):

Copyable text
# Less effective
prompt = "Summarize this. Don't include bullet points. Don't exceed 200 words."

# More effective - positive framing
prompt = "Summarize this in 2-3 flowing paragraphs. Target 150 words. Use prose format."

Tools That Make This Practical

ToolBest ForMonthly Downloads
InstructorAPI models, multi-provider3M+
OutlinesLocal models, guaranteed compliance-
LangChainWhen already in LangChain ecosystem-
GuidanceToken-level control, research-

Instructor Example

Copyable text
import instructor
from pydantic import BaseModel

class User(BaseModel):
    name: str
    age: int

# Automatically retries on validation failure
client = instructor.from_provider("openai/gpt-4o")
user = client.chat.completions.create(
    response_model=User,
    max_retries=3,
    messages=[{"role": "user", "content": "Extract: Jason is 30 years old"}]
)

Outlines Example

Copyable text
import outlines
from pydantic import BaseModel

class Character(BaseModel):
    name: str
    age: int
    armor: str

model = outlines.models.transformers("microsoft/Phi-3-mini-128k-instruct")
generator = outlines.generate.json(model, Character)

character = generator("Generate a fantasy RPG character:")
# Guaranteed to be a valid Character instance

When Things Go Wrong

The Truncation Trap

The model hits max_tokens before completing JSON:

Copyable text
def safe_llm_call(prompt: str, schema: dict):
    response = client.chat.completions.create(...)

    # ALWAYS check finish_reason
    if response.choices[0].finish_reason == "length":
        raise ValueError("Response truncated - increase max_tokens")

    return response.choices[0].message.content

Schema Compliance ≠ Content Accuracy

Structured outputs guarantee format, not truth. Always validate content:

Copyable text
from pydantic import field_validator

class ExtractedData(BaseModel):
    company_name: str
    founded_year: int

    @field_validator('founded_year')
    def reasonable_year(cls, v):
        if v < 1800 or v > 2026:
            raise ValueError(f'Year {v} seems unlikely')
        return v

Graceful Degradation Pattern

Copyable text
def extract_with_fallback(text: str) -> dict:
    # Tier 1: Strict schema
    try:
        return call_with_strict_schema(text).model_dump()
    except ValidationError:
        pass

    # Tier 2: Relaxed schema
    try:
        return call_with_relaxed_schema(text).model_dump()
    except ValidationError:
        pass

    # Tier 3: JSON mode without schema
    try:
        return json.loads(call_with_json_mode(text))
    except json.JSONDecodeError:
        pass

    # Tier 4: Return partial/default data
    return {"raw_text": text, "extraction_failed": True}

Key Takeaways

For API-based applications:

  • Use native structured outputs (OpenAI, Anthropic, or Gemini)
  • Combine with Instructor for automatic retries
  • Always check finish_reason for truncation

For local model deployment:

  • Use Outlines for guaranteed compliance
  • Consider grammar constraints for code generation

For all cases:

  • Design schemas simply with clear field descriptions
  • Implement retry logic with exponential backoff
  • Build graceful degradation paths
  • Test edge cases before production

The tools exist. "The LLM returned invalid JSON" is no longer an acceptable production failure mode.


Building real-estate growth tools without a Frankenstack? Talk to us.

Keep going

One useful next step.

Find your next practical guide