How to Implement Guardrails in Embabel: A Complete Guide to Input and Output Validation
Implement guardrails in Embabel by extending UserInputGuardRail or AssistantMessageGuardRail, registering instances globally via application.properties or per-call through PromptRunner.withGuardRails(), and returning ValidationResult.critical() to block execution when safety checks fail.
The embabel/embabel-agent framework provides a type-safe, extensible architecture for enforcing safety and quality constraints on LLM interactions. Implementing guardrails in Embabel allows you to validate user messages before they reach the model and scrutinize assistant outputs before they are returned to users, ensuring compliance with content policies and structural requirements.
Understanding the Guardrail Architecture
The guardrail system is built around a small set of core interfaces that define validation contracts. At runtime, the framework orchestrates these validations through the GlobalGuardRailsRegistry and the PromptRunner execution engine.
The GuardRail Base Interface
All guardrails implement the GuardRail interface, which defines the core validation contract:
interface GuardRail {
fun validate(input: String, blackboard: Blackboard): ValidationResult
val description: String
}
The validate method receives the content to inspect and a Blackboard instance for accessing shared context. It returns a ValidationResult indicating whether the content passes, warns, or fails critically.
UserInputGuardRail for Pre-Processing
To validate messages before they are sent to the LLM, implement UserInputGuardRail. This interface extends GuardRail and provides default helper methods for handling multimodal content:
combineMessages(List<UserMessage>): Concatenates user messages into a single validation stringvalidate(MultimodalContent, Blackboard): Handles image and text inputs
Input guardrails execute sequentially before the LLM API call. If any guardrail returns a CRITICAL result, the framework immediately throws a GuardRailViolationException and cancels the request.
AssistantMessageGuardRail for Post-Processing
To validate LLM outputs, implement AssistantMessageGuardRail. This interface processes AssistantMessage objects after generation but before tool execution or user delivery.
Assistant guardrails are essential for enforcing output schemas, detecting hallucinations, or filtering sensitive content in generated responses.
Registering Guardrails Globally
The GlobalGuardRailsRegistry manages guardrail lifecycle and instantiation. At startup, it scans application.properties (or application.yml) for fully-qualified class names and instantiates guardrails using the default constructor.
Add global guardrails to your configuration:
# Global user input validation
embabel.agent.guardrails.user-input=com.example.ProhibitedTermsGuardRail,com.example.InputLengthValidator
# Global assistant output validation
embabel.agent.guardrails.assistant-message=com.example.JsonSchemaValidator,com.example.ToxicityDetector
# Strict mode: fail on any critical validation
embabel.agent.guardrails.fail-on-error=true
The registry exposes two methods used by the runtime:
getUserInputGuardRails(): Returns all registered input validatorsgetAssistantMessageGuardRails(): Returns all output validators
Implementing Custom Guardrails
Guardrails can be implemented in Kotlin or Java. Each implementation must provide the description property and override the appropriate validate method.
Blocking Profanity in User Input
Create a Kotlin guardrail that blocks messages containing prohibited terms:
package com.example.guardrails
import com.embabel.agent.api.validation.guardrails.UserInputGuardRail
import com.embabel.agent.api.validation.guardrails.ValidationResult
import com.embabel.agent.core.Blackboard
class ProfanityFilterGuardRail : UserInputGuardRail {
private val prohibitedPattern = Regex("\\b(badword|offensive)\\b", RegexOption.IGNORE_CASE)
override val description = "Blocks messages containing profanity or prohibited terms"
override fun validate(input: String, blackboard: Blackboard): ValidationResult {
return if (prohibitedPattern.containsMatchIn(input)) {
ValidationResult.critical("Content violates safety policy: prohibited terms detected")
} else {
ValidationResult.success()
}
}
}
Register this in application.properties:
embabel.agent.guardrails.user-input=com.example.guardrails.ProfanityFilterGuardRail
Validating JSON Schema in LLM Outputs
Implement an assistant guardrail in Java that validates structural output against a JSON schema:
package com.example.guardrails;
import com.embabel.agent.api.validation.guardrails.AssistantMessageGuardRail;
import com.embabel.agent.api.validation.guardrails.ValidationResult;
import com.embabel.agent.core.Blackboard;
import com.embabel.chat.AssistantMessage;
import com.networknt.schema.JsonSchema;
import com.networknt.schema.JsonSchemaFactory;
import com.networknt.schema.SpecVersion;
import com.networknt.schema.ValidationMessage;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;
import java.util.Set;
public class StructuredOutputGuardRail implements AssistantMessageGuardRail {
private static final ObjectMapper MAPPER = new ObjectMapper();
private final JsonSchema schema;
public StructuredOutputGuardRail() {
this.schema = JsonSchemaFactory.getInstance(SpecVersion.VersionFlag.V201909)
.getSchema(getClass().getResourceAsStream("/schemas/output-schema.json"));
}
@Override
public ValidationResult validate(AssistantMessage assistantMessage, Blackboard blackboard) {
try {
JsonNode jsonNode = MAPPER.readTree(assistantMessage.getContent());
Set<ValidationMessage> errors = schema.validate(jsonNode);
if (!errors.isEmpty()) {
return ValidationResult.critical("Output schema violation: " + errors);
}
return ValidationResult.success();
} catch (Exception e) {
return ValidationResult.critical("Invalid JSON in assistant message: " + e.getMessage());
}
}
@Override
public String getDescription() {
return "Validates assistant messages against JSON schema";
}
}
Tracking Costs via Blackboard
Access the Blackboard to implement context-aware validation, such as cost limits:
package com.example.guardrails
import com.embabel.agent.api.validation.guardrails.UserInputGuardRail
import com.embabel.agent.api.validation.guardrails.ValidationResult
import com.embabel.agent.core.Blackboard
import com.embabel.metrics.Counter
class CostLimitGuardRail(private val maxTokens: Int) : UserInputGuardRail {
override val description = "Enforces maximum token budget per session"
override fun validate(input: String, blackboard: Blackboard): ValidationResult {
val tokenCounter = blackboard.get(Counter::class.java)
?: return ValidationResult.critical("Cost counter not initialized")
return if (tokenCounter.value() > maxTokens) {
ValidationResult.critical("Token budget exceeded: ${tokenCounter.value()} > $maxTokens")
} else {
ValidationResult.success()
}
}
}
Integrating Guardrails at Runtime
While global guardrails apply to all LLM operations, you can attach additional guardrails to specific calls using the PromptRunner API. The runner merges global guardrails with per-call additions, executing them in sequence.
import com.embabel.agent.api.PromptRunner
import com.embabel.agent.api.ChatMessage
val runner = PromptRunner.builder()
.withGuardRails(StructuredOutputGuardRail())
.withGuardRails(CostLimitGuardRail(maxTokens = 1000))
.build()
val response = runner.generate(
messages = listOf(ChatMessage.user("Generate a user profile as JSON"))
)
Execution flow:
- All
UserInputGuardRailinstances run; if any returnsCRITICAL, throwGuardRailViolationException - Send request to LLM
- All
AssistantMessageGuardRailinstances run; if any returnsCRITICAL, throwGuardRailViolationExceptionbefore returning content
The ValidationResult supports three severity levels:
- INFO: Logged but does not affect execution
- WARN: Logged as warning but allows continuation
- CRITICAL: Immediately aborts the operation with a
GuardRailViolationException
Summary
- Extend the correct interface: Use
UserInputGuardRailfor pre-LLM validation andAssistantMessageGuardRailfor post-generation checks. - Leverage the GlobalGuardRailsRegistry: Configure global guardrails in
application.propertiesfor application-wide policy enforcement. - Return appropriate severity: Use
ValidationResult.critical()to block unsafe content,warn()for soft alerts, andsuccess()for clean content. - Access context via Blackboard: Read and write shared state such as cost counters, conversation history, or custom metadata during validation.
- Combine global and per-call guards: Use
PromptRunner.withGuardRails()for scenario-specific validation while maintaining baseline protections globally.
Frequently Asked Questions
What is the difference between UserInputGuardRail and AssistantMessageGuardRail?
UserInputGuardRail validates content before sending it to the LLM, inspecting UserMessage objects and MultimodalContent. AssistantMessageGuardRail validates the LLM's response after generation, inspecting AssistantMessage objects before they are returned to the user or processed by tools. Both extend the base GuardRail interface but operate at different stages of the request lifecycle.
How do I prevent an LLM call from executing when a guardrail fails?
Return ValidationResult.critical() from your guardrail's validate method. When the PromptRunner or GlobalGuardRailsRegistry encounters a CRITICAL result, it immediately throws a GuardRailViolationException, preventing the LLM API call (for input guards) or returning the response to the user (for output guards). Set embabel.agent.guardrails.fail-on-error=true to ensure strict enforcement.
Can I implement guardrails in Java, or are they Kotlin-only?
Guardrails are fully interoperable with Java. The interfaces are written in Kotlin but compile to standard JVM bytecode. Java implementations must override validate() and getDescription(), and can use all framework features including the Blackboard API and ValidationResult static factory methods.
How do I access conversation history or session state inside a guardrail?
Use the Blackboard parameter passed to the validate method. The Blackboard acts as a shared context store; retrieve objects by type using blackboard.get(Counter::class.java) or blackboard.get(MyContext::class.java). This allows guardrails to implement stateful logic such as rate limiting across multiple turns or accumulating cost metrics throughout a session.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →