How to Implement Guardrails in Embabel: A Complete Guide to Input and Output Validation

Implement guardrails in Embabel by extending UserInputGuardRail or AssistantMessageGuardRail, registering instances globally via application.properties or per-call through PromptRunner.withGuardRails(), and returning ValidationResult.critical() to block execution when safety checks fail.

The embabel/embabel-agent framework provides a type-safe, extensible architecture for enforcing safety and quality constraints on LLM interactions. Implementing guardrails in Embabel allows you to validate user messages before they reach the model and scrutinize assistant outputs before they are returned to users, ensuring compliance with content policies and structural requirements.

Understanding the Guardrail Architecture

The guardrail system is built around a small set of core interfaces that define validation contracts. At runtime, the framework orchestrates these validations through the GlobalGuardRailsRegistry and the PromptRunner execution engine.

The GuardRail Base Interface

All guardrails implement the GuardRail interface, which defines the core validation contract:

interface GuardRail {
    fun validate(input: String, blackboard: Blackboard): ValidationResult
    val description: String
}

The validate method receives the content to inspect and a Blackboard instance for accessing shared context. It returns a ValidationResult indicating whether the content passes, warns, or fails critically.

UserInputGuardRail for Pre-Processing

To validate messages before they are sent to the LLM, implement UserInputGuardRail. This interface extends GuardRail and provides default helper methods for handling multimodal content:

  • combineMessages(List<UserMessage>): Concatenates user messages into a single validation string
  • validate(MultimodalContent, Blackboard): Handles image and text inputs

Input guardrails execute sequentially before the LLM API call. If any guardrail returns a CRITICAL result, the framework immediately throws a GuardRailViolationException and cancels the request.

AssistantMessageGuardRail for Post-Processing

To validate LLM outputs, implement AssistantMessageGuardRail. This interface processes AssistantMessage objects after generation but before tool execution or user delivery.

Assistant guardrails are essential for enforcing output schemas, detecting hallucinations, or filtering sensitive content in generated responses.

Registering Guardrails Globally

The GlobalGuardRailsRegistry manages guardrail lifecycle and instantiation. At startup, it scans application.properties (or application.yml) for fully-qualified class names and instantiates guardrails using the default constructor.

Add global guardrails to your configuration:


# Global user input validation

embabel.agent.guardrails.user-input=com.example.ProhibitedTermsGuardRail,com.example.InputLengthValidator

# Global assistant output validation  

embabel.agent.guardrails.assistant-message=com.example.JsonSchemaValidator,com.example.ToxicityDetector

# Strict mode: fail on any critical validation

embabel.agent.guardrails.fail-on-error=true

The registry exposes two methods used by the runtime:

  • getUserInputGuardRails(): Returns all registered input validators
  • getAssistantMessageGuardRails(): Returns all output validators

Implementing Custom Guardrails

Guardrails can be implemented in Kotlin or Java. Each implementation must provide the description property and override the appropriate validate method.

Blocking Profanity in User Input

Create a Kotlin guardrail that blocks messages containing prohibited terms:

package com.example.guardrails

import com.embabel.agent.api.validation.guardrails.UserInputGuardRail
import com.embabel.agent.api.validation.guardrails.ValidationResult
import com.embabel.agent.core.Blackboard

class ProfanityFilterGuardRail : UserInputGuardRail {
    private val prohibitedPattern = Regex("\\b(badword|offensive)\\b", RegexOption.IGNORE_CASE)
    
    override val description = "Blocks messages containing profanity or prohibited terms"

    override fun validate(input: String, blackboard: Blackboard): ValidationResult {
        return if (prohibitedPattern.containsMatchIn(input)) {
            ValidationResult.critical("Content violates safety policy: prohibited terms detected")
        } else {
            ValidationResult.success()
        }
    }
}

Register this in application.properties:

embabel.agent.guardrails.user-input=com.example.guardrails.ProfanityFilterGuardRail

Validating JSON Schema in LLM Outputs

Implement an assistant guardrail in Java that validates structural output against a JSON schema:

package com.example.guardrails;

import com.embabel.agent.api.validation.guardrails.AssistantMessageGuardRail;
import com.embabel.agent.api.validation.guardrails.ValidationResult;
import com.embabel.agent.core.Blackboard;
import com.embabel.chat.AssistantMessage;
import com.networknt.schema.JsonSchema;
import com.networknt.schema.JsonSchemaFactory;
import com.networknt.schema.SpecVersion;
import com.networknt.schema.ValidationMessage;
import com.fasterxml.jackson.databind.JsonNode;
import com.fasterxml.jackson.databind.ObjectMapper;

import java.util.Set;

public class StructuredOutputGuardRail implements AssistantMessageGuardRail {
    private static final ObjectMapper MAPPER = new ObjectMapper();
    private final JsonSchema schema;

    public StructuredOutputGuardRail() {
        this.schema = JsonSchemaFactory.getInstance(SpecVersion.VersionFlag.V201909)
                .getSchema(getClass().getResourceAsStream("/schemas/output-schema.json"));
    }

    @Override
    public ValidationResult validate(AssistantMessage assistantMessage, Blackboard blackboard) {
        try {
            JsonNode jsonNode = MAPPER.readTree(assistantMessage.getContent());
            Set<ValidationMessage> errors = schema.validate(jsonNode);
            
            if (!errors.isEmpty()) {
                return ValidationResult.critical("Output schema violation: " + errors);
            }
            return ValidationResult.success();
        } catch (Exception e) {
            return ValidationResult.critical("Invalid JSON in assistant message: " + e.getMessage());
        }
    }

    @Override
    public String getDescription() {
        return "Validates assistant messages against JSON schema";
    }
}

Tracking Costs via Blackboard

Access the Blackboard to implement context-aware validation, such as cost limits:

package com.example.guardrails

import com.embabel.agent.api.validation.guardrails.UserInputGuardRail
import com.embabel.agent.api.validation.guardrails.ValidationResult
import com.embabel.agent.core.Blackboard
import com.embabel.metrics.Counter

class CostLimitGuardRail(private val maxTokens: Int) : UserInputGuardRail {
    override val description = "Enforces maximum token budget per session"

    override fun validate(input: String, blackboard: Blackboard): ValidationResult {
        val tokenCounter = blackboard.get(Counter::class.java)
            ?: return ValidationResult.critical("Cost counter not initialized")
            
        return if (tokenCounter.value() > maxTokens) {
            ValidationResult.critical("Token budget exceeded: ${tokenCounter.value()} > $maxTokens")
        } else {
            ValidationResult.success()
        }
    }
}

Integrating Guardrails at Runtime

While global guardrails apply to all LLM operations, you can attach additional guardrails to specific calls using the PromptRunner API. The runner merges global guardrails with per-call additions, executing them in sequence.

import com.embabel.agent.api.PromptRunner
import com.embabel.agent.api.ChatMessage

val runner = PromptRunner.builder()
    .withGuardRails(StructuredOutputGuardRail())
    .withGuardRails(CostLimitGuardRail(maxTokens = 1000))
    .build()

val response = runner.generate(
    messages = listOf(ChatMessage.user("Generate a user profile as JSON"))
)

Execution flow:

  1. All UserInputGuardRail instances run; if any returns CRITICAL, throw GuardRailViolationException
  2. Send request to LLM
  3. All AssistantMessageGuardRail instances run; if any returns CRITICAL, throw GuardRailViolationException before returning content

The ValidationResult supports three severity levels:

  • INFO: Logged but does not affect execution
  • WARN: Logged as warning but allows continuation
  • CRITICAL: Immediately aborts the operation with a GuardRailViolationException

Summary

  • Extend the correct interface: Use UserInputGuardRail for pre-LLM validation and AssistantMessageGuardRail for post-generation checks.
  • Leverage the GlobalGuardRailsRegistry: Configure global guardrails in application.properties for application-wide policy enforcement.
  • Return appropriate severity: Use ValidationResult.critical() to block unsafe content, warn() for soft alerts, and success() for clean content.
  • Access context via Blackboard: Read and write shared state such as cost counters, conversation history, or custom metadata during validation.
  • Combine global and per-call guards: Use PromptRunner.withGuardRails() for scenario-specific validation while maintaining baseline protections globally.

Frequently Asked Questions

What is the difference between UserInputGuardRail and AssistantMessageGuardRail?

UserInputGuardRail validates content before sending it to the LLM, inspecting UserMessage objects and MultimodalContent. AssistantMessageGuardRail validates the LLM's response after generation, inspecting AssistantMessage objects before they are returned to the user or processed by tools. Both extend the base GuardRail interface but operate at different stages of the request lifecycle.

How do I prevent an LLM call from executing when a guardrail fails?

Return ValidationResult.critical() from your guardrail's validate method. When the PromptRunner or GlobalGuardRailsRegistry encounters a CRITICAL result, it immediately throws a GuardRailViolationException, preventing the LLM API call (for input guards) or returning the response to the user (for output guards). Set embabel.agent.guardrails.fail-on-error=true to ensure strict enforcement.

Can I implement guardrails in Java, or are they Kotlin-only?

Guardrails are fully interoperable with Java. The interfaces are written in Kotlin but compile to standard JVM bytecode. Java implementations must override validate() and getDescription(), and can use all framework features including the Blackboard API and ValidationResult static factory methods.

How do I access conversation history or session state inside a guardrail?

Use the Blackboard parameter passed to the validate method. The Blackboard acts as a shared context store; retrieve objects by type using blackboard.get(Counter::class.java) or blackboard.get(MyContext::class.java). This allows guardrails to implement stateful logic such as rate limiting across multiple turns or accumulating cost metrics throughout a session.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →