How to Implement Structured Output Agents with Pydantic: A Complete Guide

Structured output agents with Pydantic combine LLM reasoning with automatic schema validation, converting raw model responses into type-safe Python objects using the output_type parameter in the OpenAI Agents SDK.

Structured output agents with Pydantic enable developers to enforce strict data contracts on LLM responses, eliminating manual JSON parsing and validation errors. This pattern, implemented throughout the awesome-llm-apps repository by Shubhamsaboo, leverages the OpenAI Agents SDK to automatically convert natural language into validated Pydantic models. By defining schemas as Python classes, you create type-safe pipelines that guarantee consistent data structures for downstream applications.

Understanding Structured Output Agents with Pydantic

A structured output agent is an LLM agent configured to return responses that conform to a predefined Pydantic schema rather than free-form text. When you set the output_type parameter to a Pydantic model class, the SDK automatically injects schema instructions into the prompt, parses the LLM's JSON response, and validates it against your model definition.

This approach provides compile-time type safety for runtime LLM outputs. If the model returns malformed JSON or violates schema constraints (such as enum values or numeric ranges), Pydantic raises validation errors immediately rather than propagating invalid data through your application.

Building a Structured Output Agent: Step-by-Step

Step 1: Define the Pydantic Schema

Create a Pydantic model that describes the exact JSON structure you expect from the LLM. Use Field descriptions to guide the model, and define enums for categorical data to constrain outputs.

In ai_agent_framework_crash_course/openai_sdk_crash_course/2_structured_output_agent/support_ticket_agent.py, the SupportTicket model (lines 33-53) demonstrates this pattern:

from enum import Enum
from typing import List, Optional
from pydantic import BaseModel, Field

class Priority(str, Enum):
    LOW = "low"
    MEDIUM = "medium"
    HIGH = "high"
    CRITICAL = "critical"

class SupportTicket(BaseModel):
    title: str = Field(description="Short issue summary")
    description: str = Field(description="Full problem description")
    priority: Priority = Field(description="Ticket priority")
    category: str = Field(description="Support category")
    customer_name: Optional[str] = Field(default=None)
    steps_to_reproduce: Optional[List[str]] = Field(default=None)
    estimated_resolution_time: str = Field(description="How long it will take")
    urgency_keywords: List[str] = Field(default=[])

Step 2: Configure the Agent with output_type

Instantiate the Agent class from the OpenAI Agents SDK, passing your Pydantic model to the output_type parameter. The SDK automatically appends schema instructions to your system prompt.

From support_ticket_agent.py (lines 55-94):

from agents import Agent

support_ticket_agent = Agent(
    name="Support Ticket Creator",
    instructions="""
    You are a support-ticket writer. Extract the requested fields and
    ALWAYS return a JSON object that matches the SupportTicket schema.
    """,
    output_type=SupportTicket,  # Pydantic model passed here

)

Step 3: Execute with Runner and Access Typed Results

Use Runner.run_sync() (or the async variant) to execute the agent. The method returns a RunResult object where final_output contains your validated Pydantic instance.

From support_ticket_agent.py (lines 96-115):

from agents import Runner

def create_ticket(complaint: str) -> SupportTicket:
    """Runs the agent and returns a typed SupportTicket."""
    result = Runner.run_sync(support_ticket_agent, complaint)
    return result.final_output

# Example usage

if __name__ == "__main__":
    sample = (
        "Hi, I'm Jane. My credit card was charged twice for the "
        "subscription this month. Please fix this ASAP."
    )
    ticket = create_ticket(sample)
    print(ticket.json(indent=2))

Complete Implementation Examples from awesome-llm-apps

Support Ticket Agent with Flat Schema

The support_ticket_agent.py file demonstrates a flat schema approach suitable for straightforward data extraction. It uses enums for the Priority field and optional fields for data that may not appear in every input. The demonstrate_support_tickets function shows how to batch process multiple complaints into typed SupportTicket objects.

Product Review Agent with Nested Schema and Validators

For complex scenarios requiring hierarchical data, product_review_agent.py (lines 60-74) implements nested Pydantic models:

from pydantic import BaseModel, Field, validator

class ProductInfo(BaseModel):
    name: Optional[str] = Field(default=None)
    category: ProductCategory = Field(description="Inferred category")
    brand: Optional[str] = None
    price_mentioned: Optional[str] = None

class ReviewMetrics(BaseModel):
    rating: int = Field(ge=1, le=5, description="Star rating")
    sentiment: Sentiment = Field(description="Overall sentiment")
    confidence_score: float = Field(ge=0, le=1)
    word_count: int = Field(ge=0)

class ProductReview(BaseModel):
    product_info: ProductInfo
    metrics: ReviewMetrics
    aspects: ReviewAspects
    main_positives: List[str] = []
    main_negatives: List[str] = []
    would_recommend: Optional[bool] = None
    summary: str
    key_phrases: List[str] = []

    @validator("key_phrases")
    def limit_key_phrases(cls, v):
        return v[:5]  # Enforce maximum 5 phrases

This example demonstrates Field constraints (ge, le for numeric ranges), custom validators (limit_key_phrases), and nested composition (ProductInfo, ReviewMetrics inside ProductReview). The demonstrate_review_analysis function (lines 27-61) shows how to process raw review text into this complex structure using Runner.run_sync(product_review_agent, text).

Advanced Pydantic Patterns for LLM Outputs

Enforcing Constraints with Field Validators

Pydantic validators ensure LLM outputs meet business logic requirements beyond basic type checking. In product_review_agent.py, the limit_key_phrases validator truncates the list to five items maximum, preventing context window bloat. You can also use Pydantic's root_validator to enforce cross-field dependencies (e.g., ensuring estimated_resolution_time matches the priority level).

Handling Complex Nested Models

Nested schemas allow you to extract hierarchical relationships from unstructured text. When defining nested models for structured output agents with Pydantic, ensure every sub-model inherits from BaseModel and that you import Optional from typing for fields that may be absent. The SDK recursively generates JSON schemas for nested classes, allowing the LLM to understand complex object graphs like product_info.category or metrics.confidence_score.

Summary

  • Structured output agents with Pydantic enforce type safety on LLM responses by validating JSON against Python schemas before it reaches application logic.
  • The OpenAI Agents SDK simplifies implementation through the output_type parameter, which automatically injects schema instructions and parses responses.
  • Pydantic models act as the contract between your application and the LLM, supporting enums, optional fields, nested objects, and custom validators.
  • Reference implementations in awesome-llm-apps demonstrate both flat schemas (support_ticket_agent.py) and complex nested schemas (product_review_agent.py) with field constraints.
  • Using Runner.run_sync() returns a typed final_output object that provides IDE auto-completion and runtime validation, eliminating fragile string parsing.

Frequently Asked Questions

What is the difference between structured output and function calling?

Structured output constrains the LLM's final response to match a specific JSON schema, returning data directly to your application as a typed object. Function calling allows the model to invoke external tools or APIs by generating JSON arguments, but the final response to the user may still be free-form text. With structured output agents using Pydantic, you eliminate the need for manual JSON parsing while ensuring type safety, whereas function calling focuses on extending the agent's capabilities through external integrations.

How does Pydantic validation improve LLM reliability?

Pydantic validation acts as a strict contract enforcement layer that catches malformed outputs, type mismatches, and constraint violations before they propagate through your application. When an LLM returns JSON that violates the schema—such as a string where an integer is expected or an invalid enum value—Pydantic raises a validation error immediately. This fail-fast approach prevents subtle bugs caused by inconsistent data shapes and ensures that downstream logic receives exactly the typed structure it expects, making structured output agents with Pydantic significantly more robust than manual parsing approaches.

Can I use nested Pydantic models with the OpenAI Agents SDK?

Yes, the OpenAI Agents SDK fully supports nested Pydantic models for complex hierarchical data extraction, as demonstrated in the product_review_agent.py implementation. You can define sub-models like ProductInfo and ReviewMetrics, then compose them into a parent ProductReview model; the SDK automatically generates the complete JSON schema including all nested properties. This allows you to extract structured relationships from unstructured text, such as linking specific metrics to product categories, while maintaining full type safety and IDE auto-completion throughout the nested object hierarchy.

What happens if the LLM returns invalid JSON?

If the LLM returns malformed JSON or data that violates your Pydantic schema constraints, the OpenAI Agents SDK will raise a validation error during the parsing phase before returning control to your application. The SDK attempts to coerce common type mismatches automatically, but strict constraints like enum values, numeric ranges (defined via ge/le in Field), or custom validators (like limit_key_phrases) will trigger ValidationError exceptions if violated. This behavior ensures that your application never processes corrupted or unexpected data structures, allowing you to implement error handling logic to retry the request or prompt the user for clarification when validation fails.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →