# How to Implement Retrieval Augmented Generation (RAG) with Embabel Agent

> Implement Retrieval Augmented Generation RAG with Embabel Agent. Wrap search backends like Lucene, configure retrieval, and register as an LLM tool for automatic invocation. Get started now.

- Repository: [Embabel/embabel-agent](https://github.com/embabel/embabel-agent)
- Tags: how-to-guide
- Published: 2026-08-09

---

**Implement Retrieval Augmented Generation (RAG) in Embabel Agent by instantiating `ToolishRag` to wrap a search backend like Lucene, configuring retrieval behavior through `RagOptions`, and registering the component as a standard LLM tool that the model invokes automatically during generation.**

Embabel Agent provides a modular, tool-first architecture for building RAG pipelines without complex orchestration layers. By treating retrieval as a standard tool call, the framework allows large language models to request external knowledge dynamically. This guide demonstrates how to configure ingestion, indexing, and quality evaluation using the actual source implementation from the `embabel/embabel-agent` repository.

## Understanding the Tool-First RAG Architecture

The modern RAG implementation in Embabel centers on **`ToolishRag`**, defined in [`embabel-agent-rag/embabel-agent-rag-core/src/main/kotlin/com/embabel/agent/rag/tools/ToolishRag.kt`](https://github.com/embabel/embabel-agent/blob/main/embabel-agent-rag/embabel-agent-rag-core/src/main/kotlin/com/embabel/agent/rag/tools/ToolishRag.kt). This class wraps any implementation of search operations and exposes a consistent interface that the LLM consumes as a tool.

The legacy **`RagService`** interface ([`embabel-agent-rag/embabel-agent-rag-pipeline/src/main/kotlin/com/embabel/agent/rag/service/RagService.kt`](https://github.com/embabel/embabel-agent/blob/main/embabel-agent-rag/embabel-agent-rag-pipeline/src/main/kotlin/com/embabel/agent/rag/service/RagService.kt)) remains available for backward compatibility but is deprecated for new development. When using `ToolishRag`, the LLM receives a tool schema—typically a `search(query: String, topK: Int)` function—and decides independently when to retrieve context based on the conversation state.

## Setting Up the Lucene Search Backend

The reference implementation uses **`LuceneSearchOperations`** ([`embabel-agent-rag/embabel-agent-rag-lucene/src/main/kotlin/com/embabel/agent/rag/lucene/LuceneSearchOperations.kt`](https://github.com/embabel/embabel-agent/blob/main/embabel-agent-rag/embabel-agent-rag-lucene/src/main/kotlin/com/embabel/agent/rag/lucene/LuceneSearchOperations.kt)) to manage indices. This backend supports both pure keyword search and hybrid vector search when paired with an embedding service.

### Pure Keyword Search

To implement RAG using only Lucene's keyword index:

```kotlin
import com.embabel.agent.rag.lucene.LuceneSearchOperations
import com.embabel.agent.rag.tools.ToolishRag
import com.embabel.agent.rag.tools.RagOptions
import java.nio.file.Paths
import java.time.Duration

// Initialize index at local path
val lucene = LuceneSearchOperations(
    indexPath = Paths.get("rag-index"),
    embeddingService = null
)

// Index raw text documents
val docs = listOf(
    "Embabel Agent provides tool-first LLM integration.",
    "Retrieval Augmented Generation reduces hallucinations."
)
lucene.indexDocuments(docs.map { it.toByteArray() })

// Wrap in ToolishRag with 5-second timeout
val ragTool = ToolishRag(
    name = "lucene-rag",
    description = "Keyword search over product documentation",
    searchOperations = lucene,
    options = RagOptions(maxResults = 5, timeout = Duration.ofSeconds(5))
)

```

### Hybrid Search with Custom Embeddings

Enable semantic search by injecting an `EmbeddingService` into the Lucene operations:

```kotlin
import com.embabel.agent.rag.embeddings.OpenAiEmbeddingService

// Configure embedding provider
val embeddingService = OpenAiEmbeddingService(
    apiKey = System.getenv("EMBABEL_OPENAI_API_KEY")
)

// Create hybrid index
val lucene = LuceneSearchOperations(
    indexPath = Paths.get("rag-hybrid"),
    embeddingService = embeddingService
)

// Enable dual-shot retrieval (search + synthesis)
val hybridTool = ToolishRag(
    name = "hybrid-rag",
    description = "Hybrid keyword and vector search",
    searchOperations = lucene,
    options = RagOptions(maxResults = 3, dualShot = true)
)

```

## Configuring Retrieval Parameters

Control retrieval behavior using **`RagOptions`**, located in [`embabel-agent-rag/embabel-agent-rag-pipeline/src/main/kotlin/com/embabel/agent/rag/tools/RagOptions.kt`](https://github.com/embabel/embabel-agent/blob/main/embabel-agent-rag/embabel-agent-rag-pipeline/src/main/kotlin/com/embabel/agent/rag/tools/RagOptions.kt). This immutable configuration class accepts:

- **`maxResults`**: Integer specifying the number of chunks to return
- **`dualShot`**: Boolean enabling an additional LLM synthesis step on retrieved results
- **`timeout`**: `Duration` object limiting search execution time

These parameters are set during `ToolishRag` construction but can be overridden per-request when calling the tool programmatically.

## Ingesting Documents with Tika

For complex document types (PDF, DOCX, HTML), use **`TikaHierarchicalContentReader`** from [`embabel-agent-rag/embabel-agent-rag-tika/src/main/kotlin/com/embabel/agent/rag/ingestion/TikaHierarchicalContentReader.kt`](https://github.com/embabel/embabel-agent/blob/main/embabel-agent-rag/embabel-agent-rag-tika/src/main/kotlin/com/embabel/agent/rag/ingestion/TikaHierarchicalContentReader.kt). This component parses files into hierarchical `LeafSection` objects that preserve document structure before chunking.

The reader extracts text while maintaining semantic boundaries, ensuring that code blocks, sections, and paragraphs remain intact during the ingestion phase prior to indexing.

## Registering the RAG Tool

After configuration, register the tool with your agent's registry. The LLM automatically handles invocation:

```java
import com.embabel.agent.model.ChatModel;
import com.embabel.agent.model.ChatResponse;

ChatModel chat = EmbabelChatModel.builder()
    .modelName("gpt-4o-mini")
    .build();

chat.registerTool(ragTool);

ChatResponse response = chat.generate(
    "What are the benefits of tool-first RAG architectures?"
);
System.out.println(response.getContent());

```

During execution, the model internally calls the `search` method with an appropriate query, retrieves the top-k chunks, and incorporates that context into the final generated response.

## Evaluating Quality with RAGAS Scores

Embabel automatically computes quality metrics for every retrieval operation. The **`RagResponse`** class in [`embabel-agent-rag/embabel-agent-rag-core/src/main/kotlin/com/embabel/agent/rag/service/RagResponse.kt`](https://github.com/embabel/embabel-agent/blob/main/embabel-agent-rag/embabel-agent-rag-core/src/main/kotlin/com/embabel/agent/rag/service/RagResponse.kt) aggregates precision, recall, context relevancy, and latency into a harmonic-mean **RAGAS** score.

Access the overall quality metric programmatically:

```kotlin
val result = ragTool.search("Explain hybrid search", RagOptions())
val qualityScore = result.overallScore // Range 0.0 to 1.0
println("Retrieval quality: $qualityScore")

```

The system also emits **`RagPipelineEvent`** instances (defined in [`embabel-agent-rag/embabel-agent-rag-pipeline/src/main/kotlin/com/embabel/agent/rag/pipeline/event/PipelineRagEvents.kt`](https://github.com/embabel/embabel-agent/blob/main/embabel-agent-rag/embabel-agent-rag-pipeline/src/main/kotlin/com/embabel/agent/rag/pipeline/event/PipelineRagEvents.kt)) at each pipeline stage. These events integrate with Embabel's tracing framework for comprehensive observability.

## Summary

- **`ToolishRag`** is the primary modern implementation for RAG in Embabel Agent, wrapping search backends as standard LLM tools.
- Configure retrieval behavior using **`RagOptions`** to set `maxResults`, `dualShot` synthesis, and `timeout` durations.
- Use **`LuceneSearchOperations`** for keyword or hybrid search by optionally providing an `EmbeddingService`.
- Parse complex documents with **`TikaHierarchicalContentReader`** to generate hierarchical `LeafSection` objects for semantic chunking.
- Register the configured tool with your agent; the LLM automatically invokes retrieval when the conversation requires external knowledge.
- Monitor retrieval quality through built-in **RAGAS** scoring and pipeline events.

## Frequently Asked Questions

### What is the difference between ToolishRag and RagService?

**`ToolishRag`** implements the current tool-first architecture where retrieval is exposed as a standard function that the LLM calls, while **`RagService`** represents the legacy pipeline-based API. New projects should use `ToolishRag` for better integration with the agent's tool registry and observability systems.

### How do I implement hybrid search with custom embeddings?

Pass an implementation of `EmbeddingService`—such as `OpenAiEmbeddingService`—to the `LuceneSearchOperations` constructor. When the embedding service is non-null, the system automatically combines Lucene's keyword index with vector similarity search for hybrid retrieval.

### How can I monitor RAG pipeline performance?

Access the `overallScore` field on any `RagResponse` object to retrieve the RAGAS quality metric, which aggregates precision, recall, and relevance. Additionally, subscribe to **`RagPipelineEvent`** emissions to trace individual pipeline stages through Embabel's integrated observability framework.

### When does the LLM invoke the RAG tool?

The LLM receives the tool's schema—including the function signature and description—as part of its context. Based on the user query and the tool's description, the model autonomously decides whether to call `search()` to retrieve external knowledge before generating a response, requiring no explicit routing logic in your application code.