How to Implement Guardrails and Evals for Agent Output Validation in Agno
Implement guardrails in Agno by subclassing BaseGuardrail to create validators that check inputs and outputs via check() or async_check(), attach them to agents using add_guardrail(), and verify behavior with pytest integration tests.
Agno provides a robust guardrail system that enables automatic validation and sanitization of agent inputs and outputs without modifying core agent logic. The agno-agi/agno repository implements this through a lightweight inheritance pattern where guardrails hook into the agent execution pipeline at pre- and post-processing stages. This architecture allows developers to enforce safety, compliance, and quality constraints across individual agents, entire teams, or complex workflows.
Understanding the Guardrail Architecture
Agno's guardrail system centers on the abstract BaseGuardrail class defined in libs/agno/agno/guardrails/base.py. Every guardrail must implement two core methods to participate in the validation pipeline:
check(run_input: RunInput) → None– Synchronous validation that raisesInputCheckErrororOutputCheckErrorwhen data violates rules.async_check(run_input: RunInput) → Awaitable[None]– Asynchronous counterpart for non-blocking validation contexts.
When an agent executes, the framework constructs a RunInput object and runs all attached guardrails before the LLM call (input validation) and after response generation (output validation). This sequence ensures that only vetted data reaches the model and that generated content meets safety standards before returning to users.
Built-in Guardrails for Common Validation Needs
The Agno framework ships with three production-ready guardrails located in libs/agno/agno/guardrails/:
PromptInjectionGuardrail (prompt_injection.py) detects attempts to override system instructions using built-in pattern matching against known injection signatures.
PIIDetectionGuardrail and PIIMaskingGuardrail (pii.py) identify sensitive data including SSNs, credit cards, emails, and phone numbers, with optional automatic masking of detected entities.
OpenAIModerationGuardrail (openai.py) integrates with OpenAI's moderation API to filter content across categories like hate, harassment, and self-harm.
Each guardrail accepts configuration parameters at initialization, such as custom regex patterns or API credentials, making them adaptable to specific organizational policies.
Adding Guardrails to Agents and Teams
Attach guardrails to any Agent instance using the add_guardrail method defined in libs/agno/agno/agent/agent.py. The following example demonstrates attaching both prompt injection detection and PII masking:
from agno.agent import Agent
from agno.guardrails import PromptInjectionGuardrail, PIIDetectionGuardrail
agent = Agent(
name="support-bot",
model="gpt-4o-mini",
)
# Instantiate and attach guardrails
prompt_guard = PromptInjectionGuardrail()
pii_guard = PIIDetectionGuardrail(mask_pii=True)
agent.add_guardrail(prompt_guard)
agent.add_guardrail(pii_guard)
# Guardrails execute automatically during run()
response = agent.run("My phone number is 555-123-4567.")
For organization-wide enforcement, attach guardrails to Team instances in libs/agno/agno/team/team.py. All agents belonging to that team inherit the guardrails automatically:
from agno.team import Team
from agno.guardrails import OpenAIModerationGuardrail
team = Team(name="sales")
team.add_guardrail(OpenAIModerationGuardrail(api_key="YOUR_OPENAI_KEY"))
Creating Custom Guardrails
Domain-specific validation requires subclassing BaseGuardrail. The following implementation in cookbook/02_agents/08_guardrails/custom_guardrail.py demonstrates a guardrail that blocks outputs containing specific keywords:
from agno.guardrails.base import BaseGuardrail
from agno.exceptions import OutputCheckError
import re
class SpamDetectionGuardrail(BaseGuardrail):
"""Reject outputs containing the word 'spam' (case-insensitive)."""
_spam_pattern = re.compile(r"\bspam\b", re.I)
def check(self, run_input):
if self._spam_pattern.search(run_input.output):
raise OutputCheckError("Spam detected in agent output.")
Attach custom guardrails identically to built-ins:
from cookbook.02_agents.08_guardrails.custom_guardrail import SpamDetectionGuardrail
agent.add_guardrail(SpamDetectionGuardrail())
For input validation, inspect run_input.prompt instead of run_input.output. The following MinTokenGuardrail enforces minimum prompt length:
from agno.guardrails.base import BaseGuardrail
from agno.exceptions import InputCheckError
class MinTokenGuardrail(BaseGuardrail):
def __init__(self, min_tokens: int = 5):
self.min_tokens = min_tokens
def check(self, run_input):
token_count = len(run_input.prompt.split())
if token_count < self.min_tokens:
raise InputCheckError(
f"Prompt too short: need at least {self.min_tokens} tokens."
)
Testing and Evaluating Guardrails
Validate guardrail behavior using the integration test suite located at libs/agno/tests/integration/agent/test_guardrails.py. These tests demonstrate the standard evaluation pattern: creating fixtures with guardrail-equipped agents, feeding safe and unsafe prompts, and asserting that check raises appropriate exceptions.
The following pytest example evaluates PII protection:
import pytest
from agno.agent import Agent
from agno.guardrails import PIIDetectionGuardrail
from agno.exceptions import OutputCheckError
@pytest.fixture
def pii_agent():
ag = Agent(name="support", model="gpt-4o-mini")
ag.add_guardrail(PIIDetectionGuardrail(mask_pii=False))
return ag
def test_no_pii_leak(pii_agent):
# Safe request should not raise
safe_output = pii_agent.run("What is the weather today?").content
# Unsafe request containing SSN should raise OutputCheckError
with pytest.raises(OutputCheckError):
pii_agent.run("My SSN is 123-45-6789.")
Run the complete validation suite using:
source .venv/bin/activate
pytest libs/agno/tests/integration/agent/test_guardrails.py -q
The test test_hook_with_guardrail_exceptions in libs/agno/tests/integration/agent/test_hooks.py verifies that guardrail exceptions surface correctly in agent responses without breaking execution contracts.
Guardrails in Teams and Workflows
Team-level guardrails provide centralized policy enforcement. When attached to a Team object, every agent within that team automatically inherits the validation logic, ensuring consistent PII masking or content moderation across multi-agent systems.
Workflows in libs/agno/agno/workflow/ also support guardrails via the add_guardrail method, enabling validation at every step of complex pipelines including parallel execution branches and iterative loops.
Summary
- Inherit from
BaseGuardrail(libs/agno/agno/guardrails/base.py) to create custom validators implementingcheck()andasync_check(). - Use built-in options for prompt injection detection, PII handling, and OpenAI moderation without writing custom code.
- Attach via
add_guardrail()onAgent,Team, or workflow instances; team-level attachments propagate to all member agents. - Validate with pytest using the patterns in
libs/agno/tests/integration/agent/test_guardrails.pyto ensure guardrails catch violations and allow legitimate traffic. - Handle exceptions by catching
InputCheckErrororOutputCheckErrorto implement fallback logic when validation fails.
Frequently Asked Questions
How do guardrails handle asynchronous agent execution?
Guardrails provide both synchronous check() and asynchronous async_check() methods. When an agent runs in async mode, the framework automatically invokes async_check() for all attached guardrails, preventing blocking I/O during validation. The integration tests in test_guardrails.py include async variants to verify this behavior.
Can I apply different guardrails to different agents within the same team?
Yes. While team-level guardrails inherited via Team.add_guardrail() apply to all agents, you can override or extend protection by calling add_guardrail() on individual agent instances. Agent-specific guardrails execute alongside inherited ones, allowing base organizational policies plus specialized validation for specific roles.
What happens when a guardrail detects a violation?
The guardrail raises either InputCheckError (for pre-processing violations) or OutputCheckError (for post-processing violations). According to the test suite in test_hooks.py, these exceptions are captured and returned as part of the agent's response structure, allowing calling code to detect validation failures and implement retry logic or user notifications.
How do I test custom guardrails before deploying to production?
Create pytest fixtures that instantiate agents with your custom guardrail, then write test cases that assert check raises the appropriate exception for invalid inputs/outputs and passes silently for valid data. Reference libs/agno/tests/integration/agent/test_guardrails.py for the exact patterns used to validate sync and async guardrail behavior, including edge cases like empty inputs or malformed regex patterns.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →