How to Implement Retrieval Augmented Generation (RAG) with Embabel Agent

Implement Retrieval Augmented Generation (RAG) in Embabel Agent by instantiating ToolishRag to wrap a search backend like Lucene, configuring retrieval behavior through RagOptions, and registering the component as a standard LLM tool that the model invokes automatically during generation.

Embabel Agent provides a modular, tool-first architecture for building RAG pipelines without complex orchestration layers. By treating retrieval as a standard tool call, the framework allows large language models to request external knowledge dynamically. This guide demonstrates how to configure ingestion, indexing, and quality evaluation using the actual source implementation from the embabel/embabel-agent repository.

Understanding the Tool-First RAG Architecture

The modern RAG implementation in Embabel centers on ToolishRag, defined in embabel-agent-rag/embabel-agent-rag-core/src/main/kotlin/com/embabel/agent/rag/tools/ToolishRag.kt. This class wraps any implementation of search operations and exposes a consistent interface that the LLM consumes as a tool.

The legacy RagService interface (embabel-agent-rag/embabel-agent-rag-pipeline/src/main/kotlin/com/embabel/agent/rag/service/RagService.kt) remains available for backward compatibility but is deprecated for new development. When using ToolishRag, the LLM receives a tool schema—typically a search(query: String, topK: Int) function—and decides independently when to retrieve context based on the conversation state.

Setting Up the Lucene Search Backend

The reference implementation uses LuceneSearchOperations (embabel-agent-rag/embabel-agent-rag-lucene/src/main/kotlin/com/embabel/agent/rag/lucene/LuceneSearchOperations.kt) to manage indices. This backend supports both pure keyword search and hybrid vector search when paired with an embedding service.

To implement RAG using only Lucene's keyword index:

import com.embabel.agent.rag.lucene.LuceneSearchOperations
import com.embabel.agent.rag.tools.ToolishRag
import com.embabel.agent.rag.tools.RagOptions
import java.nio.file.Paths
import java.time.Duration

// Initialize index at local path
val lucene = LuceneSearchOperations(
    indexPath = Paths.get("rag-index"),
    embeddingService = null
)

// Index raw text documents
val docs = listOf(
    "Embabel Agent provides tool-first LLM integration.",
    "Retrieval Augmented Generation reduces hallucinations."
)
lucene.indexDocuments(docs.map { it.toByteArray() })

// Wrap in ToolishRag with 5-second timeout
val ragTool = ToolishRag(
    name = "lucene-rag",
    description = "Keyword search over product documentation",
    searchOperations = lucene,
    options = RagOptions(maxResults = 5, timeout = Duration.ofSeconds(5))
)

Hybrid Search with Custom Embeddings

Enable semantic search by injecting an EmbeddingService into the Lucene operations:

import com.embabel.agent.rag.embeddings.OpenAiEmbeddingService

// Configure embedding provider
val embeddingService = OpenAiEmbeddingService(
    apiKey = System.getenv("EMBABEL_OPENAI_API_KEY")
)

// Create hybrid index
val lucene = LuceneSearchOperations(
    indexPath = Paths.get("rag-hybrid"),
    embeddingService = embeddingService
)

// Enable dual-shot retrieval (search + synthesis)
val hybridTool = ToolishRag(
    name = "hybrid-rag",
    description = "Hybrid keyword and vector search",
    searchOperations = lucene,
    options = RagOptions(maxResults = 3, dualShot = true)
)

Configuring Retrieval Parameters

Control retrieval behavior using RagOptions, located in embabel-agent-rag/embabel-agent-rag-pipeline/src/main/kotlin/com/embabel/agent/rag/tools/RagOptions.kt. This immutable configuration class accepts:

  • maxResults: Integer specifying the number of chunks to return
  • dualShot: Boolean enabling an additional LLM synthesis step on retrieved results
  • timeout: Duration object limiting search execution time

These parameters are set during ToolishRag construction but can be overridden per-request when calling the tool programmatically.

Ingesting Documents with Tika

For complex document types (PDF, DOCX, HTML), use TikaHierarchicalContentReader from embabel-agent-rag/embabel-agent-rag-tika/src/main/kotlin/com/embabel/agent/rag/ingestion/TikaHierarchicalContentReader.kt. This component parses files into hierarchical LeafSection objects that preserve document structure before chunking.

The reader extracts text while maintaining semantic boundaries, ensuring that code blocks, sections, and paragraphs remain intact during the ingestion phase prior to indexing.

Registering the RAG Tool

After configuration, register the tool with your agent's registry. The LLM automatically handles invocation:

import com.embabel.agent.model.ChatModel;
import com.embabel.agent.model.ChatResponse;

ChatModel chat = EmbabelChatModel.builder()
    .modelName("gpt-4o-mini")
    .build();

chat.registerTool(ragTool);

ChatResponse response = chat.generate(
    "What are the benefits of tool-first RAG architectures?"
);
System.out.println(response.getContent());

During execution, the model internally calls the search method with an appropriate query, retrieves the top-k chunks, and incorporates that context into the final generated response.

Evaluating Quality with RAGAS Scores

Embabel automatically computes quality metrics for every retrieval operation. The RagResponse class in embabel-agent-rag/embabel-agent-rag-core/src/main/kotlin/com/embabel/agent/rag/service/RagResponse.kt aggregates precision, recall, context relevancy, and latency into a harmonic-mean RAGAS score.

Access the overall quality metric programmatically:

val result = ragTool.search("Explain hybrid search", RagOptions())
val qualityScore = result.overallScore // Range 0.0 to 1.0
println("Retrieval quality: $qualityScore")

The system also emits RagPipelineEvent instances (defined in embabel-agent-rag/embabel-agent-rag-pipeline/src/main/kotlin/com/embabel/agent/rag/pipeline/event/PipelineRagEvents.kt) at each pipeline stage. These events integrate with Embabel's tracing framework for comprehensive observability.

Summary

  • ToolishRag is the primary modern implementation for RAG in Embabel Agent, wrapping search backends as standard LLM tools.
  • Configure retrieval behavior using RagOptions to set maxResults, dualShot synthesis, and timeout durations.
  • Use LuceneSearchOperations for keyword or hybrid search by optionally providing an EmbeddingService.
  • Parse complex documents with TikaHierarchicalContentReader to generate hierarchical LeafSection objects for semantic chunking.
  • Register the configured tool with your agent; the LLM automatically invokes retrieval when the conversation requires external knowledge.
  • Monitor retrieval quality through built-in RAGAS scoring and pipeline events.

Frequently Asked Questions

What is the difference between ToolishRag and RagService?

ToolishRag implements the current tool-first architecture where retrieval is exposed as a standard function that the LLM calls, while RagService represents the legacy pipeline-based API. New projects should use ToolishRag for better integration with the agent's tool registry and observability systems.

How do I implement hybrid search with custom embeddings?

Pass an implementation of EmbeddingService—such as OpenAiEmbeddingService—to the LuceneSearchOperations constructor. When the embedding service is non-null, the system automatically combines Lucene's keyword index with vector similarity search for hybrid retrieval.

How can I monitor RAG pipeline performance?

Access the overallScore field on any RagResponse object to retrieve the RAGAS quality metric, which aggregates precision, recall, and relevance. Additionally, subscribe to RagPipelineEvent emissions to trace individual pipeline stages through Embabel's integrated observability framework.

When does the LLM invoke the RAG tool?

The LLM receives the tool's schema—including the function signature and description—as part of its context. Based on the user query and the tool's description, the model autonomously decides whether to call search() to retrieve external knowledge before generating a response, requiring no explicit routing logic in your application code.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →