How to Use Guardrails and Validation for AI Agents: A Complete Implementation Guide
AI agents use guardrails and validation to intercept unsafe inputs before processing and filter harmful outputs before delivery, using decorators like @input_guardrail and @output_guardrail combined with Pydantic schemas for type-safe enforcement.
The Shubhamsaboo/awesome-llm-apps repository demonstrates production-ready patterns for implementing guardrails and validation for AI agents across multiple frameworks. This guide extracts the architectural patterns from the OpenAI Agents SDK crash course and the RAG-based math tutor implementation to show you how to build robust safety layers.
Understanding the Guardrail Architecture for AI Agents
Guardrails function as a pipeline pattern that wraps your main agent execution. The architecture intercepts requests at two critical points:
- Input Guardrails – Validate user queries before the agent processes them
- Output Guardrails – Validate generated responses before returning them to users
If any guardrail detects a violation, the framework raises InputGuardrailTripwireTriggered or OutputGuardrailTripwireTriggered, immediately halting execution and preventing unsafe content from reaching the user.
Implementing Input Guardrails to Filter User Requests
Input guardrails use the @input_guardrail decorator to register validation functions. These functions receive the raw user input, run lightweight validation models, and return a GuardrailFunctionOutput.
In ai_agent_framework_crash_course/openai_sdk_crash_course/6_guardrails_validation/agent.py, the implementation uses Pydantic schemas to enforce structured outputs:
from pydantic import BaseModel
from agents import Agent, input_guardrail, GuardrailFunctionOutput
class MathHomeworkCheck(BaseModel):
is_math_homework: bool
reasoning: str
confidence: float
@input_guardrail
async def math_homework_guardrail(ctx, agent, input_data):
# Run a lightweight guardrail agent to classify intent
guardrail_agent = Agent(
name="Input Guardrail",
instructions="Classify if this is a math homework question",
output_type=MathHomeworkCheck,
)
result = await Runner.run(guardrail_agent, input_data, context=ctx.context)
output = result.final_output_as(MathHomeworkCheck)
return GuardrailFunctionOutput(
output_info=output,
tripwire_triggered=not output.is_math_homework
)
The tripwire_triggered boolean determines whether to abort the main workflow. When set to True, the framework raises InputGuardrailTripwireTriggered.
Implementing Output Guardrails to Validate Agent Responses
Output guardrails follow the same pattern but inspect the agent's generated response. They use the @output_guardrail decorator and typically enforce content safety policies or format validation.
From the same source file, here is an output guardrail that checks for inappropriate content:
from agents import output_guardrail
class ContentSafetyCheck(BaseModel):
is_inappropriate: bool
reasoning: str
severity: str
@output_guardrail
async def content_safety_guardrail(ctx, agent, output):
guardrail_agent = Agent(
name="Output Guardrail",
instructions="Detect harmful or sensitive content in the response",
output_type=ContentSafetyCheck,
)
result = await Runner.run(guardrail_agent, output, context=ctx.context)
safety_check = result.final_output_as(ContentSafetyCheck)
return GuardrailFunctionOutput(
output_info=safety_check,
tripwire_triggered=safety_check.is_inappropriate
)
When attached to a root agent, this guardrail automatically intercepts every response. If is_inappropriate is True, the framework raises OutputGuardrailTripwireTriggered before the user receives the content.
Building Validator Classes for RAG Pipelines
For Retrieval-Augmented Generation (RAG) systems, the repository provides a hybrid approach using InputValidator and OutputValidator classes. These encapsulate rule-based or ML-driven checks outside the decorator pattern.
In rag_tutorials/agentic_rag_math_agent/rag/query_router.py, validators filter queries and answers in a math tutor pipeline:
from rag.guardrails import InputValidator, OutputValidator
def answer_math_question(question: str):
input_validator = InputValidator()
output_validator = OutputValidator()
# Input validation – reject non-math queries
if not input_validator.forward(question):
return "⚠️ This assistant only answers math-related academic questions."
# ... run KB retrieval or web search, generate answer ...
answer = generate_answer(question)
# Output validation – ensure answer passes safety policy
if not output_validator.forward(question, answer):
# Fallback to safer source if validation fails
web_content = query_web(question)
answer = explain_with_openai(question, web_content)
return answer
This demonstrates a hybrid guardrail approach where lightweight validators handle routine checks, while full guardrail agents handle complex policy enforcement.
Complete Implementation Example
To deploy guardrails and validation for AI agents in production, combine the decorator pattern with structured output types. Here is the complete setup from ai_agent_framework_crash_course/openai_sdk_crash_course/6_guardrails_validation/agent.py:
from agents import Agent, Runner
import asyncio
# Attach guardrails to the main agent
root_agent = Agent(
name="Protected Customer Support Agent",
instructions="You are a helpful support agent...",
input_guardrails=[math_homework_guardrail],
output_guardrails=[content_safety_guardrail],
output_type=AgentResponse,
)
async def test_input_guardrail():
try:
result = await Runner.run(
root_agent,
"Can you solve this calculus problem?"
)
print("❌ Guardrail should have triggered")
except InputGuardrailTripwireTriggered:
print("✅ Input guardrail correctly triggered for math homework")
if __name__ == "__main__":
asyncio.run(test_input_guardrail())
The Runner.run helper executes guardrail agents automatically when listed in input_guardrails or output_guardrails. If any guardrail returns tripwire_triggered=True, the framework immediately raises the appropriate exception, preventing unsafe execution.
Summary
- Guardrails intercept requests at two critical points: before processing (input) and after generation (output), using the
@input_guardrailand@output_guardraildecorators. - Pydantic schemas provide type-safe validation outputs, enabling structured checks for content safety, topic relevance, and format compliance.
- Tripwire exceptions (
InputGuardrailTripwireTriggered,OutputGuardrailTripwireTriggered) halt execution immediately when validation fails, preventing unsafe content delivery. - Validator classes in RAG pipelines offer a hybrid approach, combining lightweight rule-based checks with full guardrail agents for complex policy enforcement.
- Source implementations in
ai_agent_framework_crash_course/openai_sdk_crash_course/6_guardrails_validation/agent.pyandrag_tutorials/agentic_rag_math_agent/rag/query_router.pydemonstrate production-ready patterns for the OpenAI Agents SDK.
Frequently Asked Questions
What is the difference between input and output guardrails for AI agents?
Input guardrails validate user queries before the main agent processes them, typically using the @input_guardrail decorator to check intent, topic relevance, or safety. Output guardrails validate generated responses before delivery to users, using @output_guardrail to detect harmful content, hallucinations, or policy violations. Both return a GuardrailFunctionOutput with a tripwire_triggered boolean that raises an exception if validation fails.
How do guardrails handle validation failures in production?
When a guardrail function returns tripwire_triggered=True, the framework immediately raises InputGuardrailTripwireTriggered or OutputGuardrailTripwireTriggered. This exception halts the main agent execution before any unsafe content is processed or delivered. In RAG pipelines, as shown in rag_tutorials/agentic_rag_math_agent/rag/query_router.py, you can catch these failures to trigger fallback logic, such as switching data sources or returning a safe default message.
Can guardrails use custom validation logic beyond LLM-based checks?
Yes, guardrails support hybrid validation approaches. While the decorator pattern often uses lightweight LLM agents for semantic checks, you can implement Validator Classes with rule-based logic, regex patterns, or traditional ML models. The repository demonstrates this in the RAG math tutor, where InputValidator and OutputValidator classes perform fast rule-based filtering before invoking heavier guardrail agents, optimizing both latency and safety coverage.
What Pydantic schemas are required for structured guardrail outputs?
Guardrail agents must return structured data via Pydantic models to enable type-safe validation. Common schemas include MathHomeworkCheck (with fields like is_math_homework: bool, reasoning: str, confidence: float) for input validation, and ContentSafetyCheck (with is_inappropriate: bool, severity: str) for output validation. These models are passed to the output_type parameter of the guardrail agent and extracted using result.final_output_as(ModelClass) to inspect validation results programmatically.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →