# What Is RAG (Retrieval-Augmented Generation) in AI Agent Knowledge Integration?

> Discover RAG Retrieval-Augmented Generation an AI pattern enabling agents to access external knowledge for dynamic responses. Enhance your AI agent capabilities.

- Repository: [Bojie Li/ai-agent-book](https://github.com/bojieli/ai-agent-book)
- Tags: deep-dive
- Published: 2026-08-25

---

**RAG (Retrieval-Augmented Generation) is a design pattern that lets LLM-driven agents pull external knowledge at inference time and inject it into the generation step, bridging the gap between static pre-trained models and up-to-date external data.**

In the `bojieli/ai-agent-book` repository, RAG is implemented as a production-ready pipeline that gives AI agents reliable access to facts without retraining the underlying language model. This article breaks down the architecture, code paths, and practical usage based on the source implementation.

## Core RAG Pipeline Architecture

The repository structures RAG as a **three-stage pipeline** with clear separation between retrieval, reranking, and generation. Each stage has dedicated modules with measurable I/O contracts.

### Stage 1: Dense and Sparse Retrieval

Documents and conversation chunks are indexed in two parallel backends:

- **Dense retrieval** — vector embeddings stored in a vector database
- **Sparse retrieval** — BM25 index for keyword matching

The `chapter3/retrieval-pipeline/` directory runs both backends simultaneously and produces a ranked candidate list. In [`fusion.py`](https://github.com/bojieli/ai-agent-book/blob/main/fusion.py), scores from both indexes are normalized and combined before the next stage.

```bash

# Run pure RAG from the repository root

uv run python chapter3/retrieval-pipeline/evaluate.py --query "What is the purpose of the BM25 algorithm?"

```

This CLI command starts both dense and sparse services, queries both indexes, applies the reranker, and outputs the top-3 passages. The implementation demonstrates **always-on retrieval** — every user query triggers the full pipeline.

### Stage 2: Neural Reranking

Raw retrieval scores are refined using a cross-encoder reranker (default: `bge-reranker-v2`). The [`fusion.py`](https://github.com/bojieli/ai-agent-book/blob/main/fusion.py) module implements this as a second-pass scoring layer:

- First pass: fast candidate generation from dense + sparse indexes
- Second pass: accurate relevance scoring from the neural reranker

This hybrid approach balances **recall** (finding all relevant documents) with **precision** (ranking the best ones highest).

### Stage 3: Prompt Construction and Generation

The top-k passages are formatted into a retrieval-augmented context block and prepended to the user query. The LLM receives a composite prompt:

```

[System: Retrieved context chunks]

User: [original query]

```

This keeps the prompt size bounded while maximizing factual grounding.

## Agentic RAG: When the Agent Decides What to Retrieve

Beyond the fixed pipeline, the repository implements **agentic RAG** — a variant where the agent itself decides whether retrieval is necessary. This lives in [`chapter3/agentic-rag/main.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter3/agentic-rag/main.py) with the `AgenticRAG` class.

```python
from chapter3.agentic_rag.main import AgenticRAG

agent = AgenticRAG(
    llm="gpt-4o-mini",
    retriever_endpoint="http://localhost:4242",   # retrieval-pipeline HTTP service

    max_rounds=3,
)

response = agent.run("Explain the difference between dense and sparse retrieval.")
print(response)

```

The `AgenticRAG` class calls `self.should_retrieve()` after each LLM turn. If the generated answer lacks sufficient evidence, it performs **iterative retrieval rounds** (up to `max_rounds`). This avoids wasteful retrievals for questions the LLM can answer from its parametric knowledge.

## RAG Variants in the Repository

The `bojieli/ai-agent-book` codebase implements three RAG deployment patterns:

| Variant | Trigger | Use Case | Key File |
|---------|---------|----------|----------|
| **Pure RAG** | Every turn | High-stakes factual domains where hallucination is unacceptable | [`chapter3/retrieval-pipeline/evaluate.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter3/retrieval-pipeline/evaluate.py) |
| **Hybrid RAG** | Open-ended questions only | Balanced latency and accuracy; core facts kept in context | Lesson 13 materials in [`slides/lesson-13.md`](https://github.com/bojieli/ai-agent-book/blob/main/slides/lesson-13.md) |
| **Structured RAG (GraphRAG/RAPTOR)** | Multi-hop reasoning queries | Complex relationships requiring entity traversal | [`chapter3/structured-index/main.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter3/structured-index/main.py) |

### Structured RAG and GraphRAG

For knowledge-intensive tasks requiring **multi-hop reasoning**, the repository provides `GraphRAG` in [`chapter3/structured-index/graph_rag.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter3/structured-index/graph_rag.py):

```python
from chapter3.structured_index.graph_rag import GraphRAG

gr = GraphRAG(
    index_path="data/graph_rag_index",
    reranker="bge-reranker-v2"
)

answer = gr.query(
    "How does the RAPTOR algorithm improve document search in a knowledge graph?",
    top_k=5
)
print(answer)

```

Instead of flat text chunks, `GraphRAG` indexes **entity-relation triples**. The retrieval process:
1. Finds relevant seed entities
2. Traverses graph relationships (RAPTOR nodes → GraphRAG communities)
3. Gathers connected facts into the context window

This handles questions like "What techniques did the authors of the BM25 paper later apply to neural retrieval?" — queries that span multiple document boundaries.

## RAG Integration with User Memory

The repository extends RAG with **two-tier memory** in `chapter3/contextual-retrieval-for-user-memory/`. This combines:

- **Episodic memory** — user-specific facts stored as memory cards
- **Semantic memory** — general domain knowledge via standard RAG

The retrieval pipeline queries both tiers in parallel, allowing personalized responses grounded in both the user's history and external knowledge sources.

## Key Implementation Files

| Component | Path | Purpose |
|-----------|------|---------|
| Retrieval pipeline entry | [`chapter3/retrieval-pipeline/evaluate.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter3/retrieval-pipeline/evaluate.py) | CLI for pure RAG evaluation |
| Score fusion & reranking | [`chapter3/retrieval-pipeline/fusion.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter3/retrieval-pipeline/fusion.py) | Combines dense, sparse, and neural scores |
| Agentic decision layer | [`chapter3/agentic-rag/main.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter3/agentic-rag/main.py) | `AgenticRAG` class with `should_retrieve()` |
| Graph-based retrieval | [`chapter3/structured-index/main.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter3/structured-index/main.py) | `GraphRAG` with multi-hop traversal |
| Memory integration | [`chapter3/contextual-retrieval-for-user-memory/README.md`](https://github.com/bojieli/ai-agent-book/blob/main/chapter3/contextual-retrieval-for-user-memory/README.md) | Two-tier memory architecture |
| Teaching materials | [`slides/lesson-12.md`](https://github.com/bojieli/ai-agent-book/blob/main/slides/lesson-12.md), [`slides/lesson-13.md`](https://github.com/bojieli/ai-agent-book/blob/main/slides/lesson-13.md) | Design trade-offs and when to use each variant |

## Summary

- **RAG** injects external knowledge into LLM prompts at inference time without model retraining.
- The `bojieli/ai-agent-book` pipeline uses **dense + sparse retrieval → neural reranker → prompt assembly**.
- **Agentic RAG** adds runtime decision-making to avoid unnecessary retrievals.
- **GraphRAG** replaces flat chunks with structured entity graphs for multi-hop reasoning.
- All components are runnable via the provided CLI and Python APIs in `chapter3/`.

## Frequently Asked Questions

### What is the difference between pure RAG and agentic RAG?

**Pure RAG** runs the retrieval pipeline on every user turn, ensuring maximum factuality at the cost of latency. **Agentic RAG** uses the `AgenticRAG` class in [`chapter3/agentic-rag/main.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter3/agentic-rag/main.py) to let the LLM judge whether external evidence is needed — reducing latency for questions answerable from parametric knowledge while preserving retrieval for uncertain claims.

### When should I use GraphRAG instead of standard RAG?

Use **GraphRAG** when your queries require **multi-hop reasoning** across connected entities — for example, finding relationships between authors, papers, and techniques. Standard RAG with flat chunks fails on these boundary-spanning questions. The `GraphRAG` implementation in [`chapter3/structured-index/main.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter3/structured-index/main.py) indexes entity-relation triples and performs graph traversal before LLM generation.

### How does the hybrid retrieval scoring work?

The [`fusion.py`](https://github.com/bojieli/ai-agent-book/blob/main/fusion.py) module combines three signals: (1) dense vector similarity from the neural encoder, (2) sparse BM25 keyword scores, and (3) cross-encoder reranker logits. Scores are normalized and fused into a single ranking. This outperforms any single method on the repository's evaluation benchmarks in [`chapter3/retrieval-pipeline/evaluate.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter3/retrieval-pipeline/evaluate.py).

### Can RAG work with user-specific memories?

Yes — the **contextual retrieval for user memory** implementation in `chapter3/contextual-retrieval-for-user-memory/` demonstrates a two-tier architecture. The retrieval pipeline queries both a user memory store (personal facts) and a general knowledge index (domain documents), merging results before prompt construction.