How embabel-agent Handles Conversational AI: A Deep Dive into the Agent-Engine Architecture

embabel-agent handles conversational AI through a Spring AI-powered agent engine that orchestrates prompts, pluggable LLM providers, dynamic tool execution, and OpenTelemetry observability—all wired together via Spring Boot auto-configuration.

The embabel-agent framework, developed by Embabel, provides a structured approach to building conversational AI applications in the JVM ecosystem. Rather than wrapping a single LLM provider, it decouples the conversation abstraction (Chatbot) from the underlying inference layer, enabling developers to swap models, extend capabilities with tools, and maintain full observability without boilerplate.


Core Conversational Architecture

The conversational flow in embabel-agent follows six distinct phases, each handled by dedicated components across the framework's modules.

1. Prompt Construction and Management

Every conversation starts with a Prompt object that encapsulates user messages, system instructions, and conversation history. The framework supports both simple text prompts and structured ChatPrompt instances for multi-turn dialogs.

The prompt layer abstracts away provider-specific formatting, allowing the same conversation logic to work across OpenAI, Anthropic, Ollama, and other supported backends.

2. LLM Provider Selection and Auto-Configuration

embabel-agent uses Spring Boot auto-configuration to wire LLM providers without manual bean definition. Each provider module exposes a ChatModel bean that conforms to Spring AI's org.springframework.ai.chat.model.ChatModel interface.

Key auto-configuration classes:

Provider Auto-Configuration Class Source Location
OpenAI AgentOpenAiAutoConfiguration embabel-agent-autoconfigure/models/embabel-agent-openai-autoconfigure/src/main/java/com/embabel/agent/autoconfigure/models/openai/AgentOpenAiAutoConfiguration.java
Anthropic AgentAnthropicAutoConfiguration embabel-agent-autoconfigure/models/embabel-agent-anthropic-autoconfigure/src/main/java/com/embabel/agent/autoconfigure/models/anthropic/AgentAnthropicAutoConfiguration.java
Ollama AgentOllamaAutoConfiguration embabel-agent-autoconfigure/models/embabel-agent-ollama-autoconfigure/src/main/java/com/embabel/agent/autoconfigure/models/ollama/AgentOllamaAutoConfiguration.java
Amazon Bedrock AgentBedrockConverseAutoConfiguration embabel-agent-autoconfigure/models/embabel-agent-bedrock-converse-autoconfigure/src/main/java/com/embabel/agent/autoconfigure/models/bedrock/converse/AgentBedrockConverseAutoConfiguration.java

The AgentLlmService abstraction selects the appropriate ChatModel based on runtime configuration, enabling multi-model deployments where different conversations route to different providers.

3. LLM Invocation and Response Handling

Once a model is selected, the core engine invokes chatModel.call(prompt) and processes the resulting ChatResponse. This response may contain:

  • Direct assistant message — returned to the user immediately
  • Tool call request — triggering the tool execution loop
  • Streaming chunks — for real-time response delivery via Stream<ChatResponse>

The invocation is wrapped in resilience patterns (timeouts, retries) configured at the AgentLlmService level.

4. Dynamic Tool Execution and Re-Planning

Tools in embabel-agent are first-class citizens implementing the Tool interface:

// Tool interface definition
package com.embabel.agent.api.tool

interface Tool {
    val name: String
    val description: String
    val parameters: ToolParameters
    
    suspend fun invoke(args: Map<String, Any?>): Any
}

Tool execution flow:

  1. LLM response includes tool call (e.g., {"tool": "weather", "args": {"location": "Berlin"}})
  2. ToolGroup resolves the tool by name from registered tools
  3. Tool executes asynchronously (supporting coroutines via suspend)
  4. Result injected into conversation history as a new message
  5. Re-planning: Engine re-invokes LLM with updated context, allowing multi-step reasoning

The ToolGroup builder pattern enables composable tool sets:

val tools = ToolGroup.builder()
    .add(WeatherTool())
    .add(CalculatorTool())
    .add(WebSearchTool())
    .build()

When a tool signals that re-planning is needed, the engine throws ReplanRequestedException, caught internally to continue the conversation loop transparently.

5. Observability with OpenTelemetry and Micrometer

Every LLM call is automatically instrumented via ChatModelObservationFilter, producing detailed spans without developer intervention.

Key observability components:

Captured span attributes:

Attribute Description Conditional
embabel.llm.model Model identifier (e.g., gpt-4o, claude-3-opus-20240229) Always
embabel.llm.input.tokens Input token count from response metadata When available
embabel.llm.output.tokens Output token count from response metadata When available
embabel.llm.input.content Full prompt content captureMessageContent = true
embabel.llm.output.content Full response content captureMessageContent = true

Configuration example:


# application.properties

embabel.agent.observability.captureMessageContent=true
embabel.agent.observability.sensitive-keywords=password,ssn,creditCard

6. The Chatbot Abstraction and Session Management

The Chatbot interface provides the highest-level API for conversational applications:

// Chatbot interface definition
package com.embabel.chat

interface Chatbot {
    suspend fun chat(message: String): String
    
    suspend fun chat(sessionId: String, message: String): String
    
    fun streamChat(message: String): Flow<String>
}

Session management is handled internally via ChatSession objects that persist conversation state across turns. Sessions can be:

  • In-memory — for single-instance deployments
  • Redis-backed — for distributed, multi-node agent clusters
  • Custom store — via ChatSessionRepository interface implementation

The AgentBuilder DSL fluent-API constructs configured agent instances:

val agent = AgentBuilder.builder()
    .withModel("gpt-4o")                    // Specific model selection
    .withTools(tools)                       // ToolGroup integration
    .withSystemPrompt("You are a helpful assistant.")
    .withSessionStore(redisSessionStore)    // Distributed sessions
    .withMaxToolIterations(5)               // Prevent infinite tool loops
    .build()

Built-in Skills and RAG Capabilities

Beyond core conversation, embabel-agent provides pre-built skills and retrieval-augmented generation infrastructure.

Skills Module

The embabel-agent-skills module contains ready-to-use tool implementations:

RAG Pipeline

The retrieval-augmented generation subsystem (embabel-agent-rag) provides:


Putting It Together: Complete Example

import com.embabel.chat.Chatbot
import com.embabel.agent.api.builder.AgentBuilder
import com.embabel.agent.tools.ToolGroup
import com.embabel.agent.api.tool.Tool
import kotlinx.coroutines.flow.Flow

// Custom tool: Query internal HR system
class HrLookupTool : Tool {
    override val name = "hr_lookup"
    override val description = "Search employee directory by name or department"
    override val parameters = ToolParameters.builder()
        .addString("query", "Search term (name or department)")
        .addOptionalString("location", "Office location filter")
        .build()
    
    override suspend fun invoke(args: Map<String, Any?>): Any {
        val query = args["query"] as String
        val location = args["location"] as String?
        return hrService.search(query, location) // Returns JSON-serializable result
    }
}

// Build the conversational agent
val hrAgent = AgentBuilder.builder()
    .withModel("claude-3-opus-20240229")
    .withTools(ToolGroup.builder().add(HrLookupTool()).build())
    .withSystemPrompt("""
        You are an HR assistant. Use the hr_lookup tool when users ask about 
        employees. Always verify information before providing contact details.
    """.trimIndent())
    .withMaxToolIterations(3)
    .build()

// Implement Chatbot interface for application integration
class HrAssistantChatbot(private val agent: Agent) : Chatbot {
    override suspend fun chat(message: String): String = 
        agent.invoke(message)
    
    override suspend fun chat(sessionId: String, message: String): String =
        agent.invoke(sessionId, message)
    
    override fun streamChat(message: String): Flow<String> =
        agent.stream(message)
}

// Deploy as Spring Bean
@Configuration
class HrAssistantConfig {
    @Bean
    fun hrAssistant(): Chatbot = HrAssistantChatbot(hrAgent)
}

This example demonstrates the complete conversational AI pipeline: tool definition, agent construction, session handling, and Spring integration—all provided by embabel-agent's architecture.


Summary

  • embabel-agent decouples conversation logic from LLM providers through the Chatbot interface and Spring AI ChatModel abstraction
  • Auto-configuration modules (OpenAI, Anthropic, Ollama, Bedrock) enable zero-code model swapping via application.properties
  • Dynamic tool execution with automatic re-planning supports multi-step reasoning workflows
  • Built-in observability (OpenTelemetry/Micrometer) captures model calls, token usage, and optional message content without instrumentation code
  • Skills and RAG modules extend core capabilities with sandboxed execution and document retrieval
  • Session persistence scales from in-memory to distributed Redis-backed stores

Frequently Asked Questions

How does embabel-agent compare to using Spring AI directly?

Spring AI provides the foundational ChatModel abstraction for LLM calls, while embabel-agent adds higher-level orchestration: the Chatbot interface for conversation management, built-in tool execution loops, automatic re-planning, session persistence, and pre-configured observability. You can use Spring AI standalone for simple use cases; embabel-agent becomes valuable when you need structured multi-turn conversations with tools and production observability.

Can I use embabel-agent with my own fine-tuned models?

Yes. Any model exposed through a Spring AI ChatModel bean integrates automatically. For self-hosted or fine-tuned models, use the Ollama auto-configuration (AgentOllamaAutoConfiguration) or implement a custom ChatModel bean—the AgentLlmService will discover and route to it based on model name or default configuration.

How does session management work in distributed deployments?

Sessions implement the ChatSessionRepository interface. The framework includes a Redis-backed implementation for distributed scenarios. Configure via AgentBuilder.withSessionStore(redisSessionStore) or set embabel.agent.session.store.type=redis in properties. In-memory storage is the default for development and single-node deployments.

What happens when a tool takes too long or fails?

Tool execution is wrapped in configurable timeouts and circuit breakers at the ToolGroup level. Failed tool calls are captured as error results in the conversation history, allowing the LLM to respond gracefully. For critical failures, you can configure maxToolIterations to prevent infinite re-planning loops and ToolFailureHandler for custom error recovery logic.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →