What Is RAG (Retrieval-Augmented Generation) in AI Agent Knowledge Integration?

RAG (Retrieval-Augmented Generation) is a design pattern that lets LLM-driven agents pull external knowledge at inference time and inject it into the generation step, bridging the gap between static pre-trained models and up-to-date external data.

In the bojieli/ai-agent-book repository, RAG is implemented as a production-ready pipeline that gives AI agents reliable access to facts without retraining the underlying language model. This article breaks down the architecture, code paths, and practical usage based on the source implementation.

Core RAG Pipeline Architecture

The repository structures RAG as a three-stage pipeline with clear separation between retrieval, reranking, and generation. Each stage has dedicated modules with measurable I/O contracts.

Stage 1: Dense and Sparse Retrieval

Documents and conversation chunks are indexed in two parallel backends:

  • Dense retrieval — vector embeddings stored in a vector database
  • Sparse retrieval — BM25 index for keyword matching

The chapter3/retrieval-pipeline/ directory runs both backends simultaneously and produces a ranked candidate list. In fusion.py, scores from both indexes are normalized and combined before the next stage.


# Run pure RAG from the repository root

uv run python chapter3/retrieval-pipeline/evaluate.py --query "What is the purpose of the BM25 algorithm?"

This CLI command starts both dense and sparse services, queries both indexes, applies the reranker, and outputs the top-3 passages. The implementation demonstrates always-on retrieval — every user query triggers the full pipeline.

Stage 2: Neural Reranking

Raw retrieval scores are refined using a cross-encoder reranker (default: bge-reranker-v2). The fusion.py module implements this as a second-pass scoring layer:

  • First pass: fast candidate generation from dense + sparse indexes
  • Second pass: accurate relevance scoring from the neural reranker

This hybrid approach balances recall (finding all relevant documents) with precision (ranking the best ones highest).

Stage 3: Prompt Construction and Generation

The top-k passages are formatted into a retrieval-augmented context block and prepended to the user query. The LLM receives a composite prompt:


[System: Retrieved context chunks]

User: [original query]

This keeps the prompt size bounded while maximizing factual grounding.

Agentic RAG: When the Agent Decides What to Retrieve

Beyond the fixed pipeline, the repository implements agentic RAG — a variant where the agent itself decides whether retrieval is necessary. This lives in chapter3/agentic-rag/main.py with the AgenticRAG class.

from chapter3.agentic_rag.main import AgenticRAG

agent = AgenticRAG(
    llm="gpt-4o-mini",
    retriever_endpoint="http://localhost:4242",   # retrieval-pipeline HTTP service

    max_rounds=3,
)

response = agent.run("Explain the difference between dense and sparse retrieval.")
print(response)

The AgenticRAG class calls self.should_retrieve() after each LLM turn. If the generated answer lacks sufficient evidence, it performs iterative retrieval rounds (up to max_rounds). This avoids wasteful retrievals for questions the LLM can answer from its parametric knowledge.

RAG Variants in the Repository

The bojieli/ai-agent-book codebase implements three RAG deployment patterns:

Variant Trigger Use Case Key File
Pure RAG Every turn High-stakes factual domains where hallucination is unacceptable chapter3/retrieval-pipeline/evaluate.py
Hybrid RAG Open-ended questions only Balanced latency and accuracy; core facts kept in context Lesson 13 materials in slides/lesson-13.md
Structured RAG (GraphRAG/RAPTOR) Multi-hop reasoning queries Complex relationships requiring entity traversal chapter3/structured-index/main.py

Structured RAG and GraphRAG

For knowledge-intensive tasks requiring multi-hop reasoning, the repository provides GraphRAG in chapter3/structured-index/graph_rag.py:

from chapter3.structured_index.graph_rag import GraphRAG

gr = GraphRAG(
    index_path="data/graph_rag_index",
    reranker="bge-reranker-v2"
)

answer = gr.query(
    "How does the RAPTOR algorithm improve document search in a knowledge graph?",
    top_k=5
)
print(answer)

Instead of flat text chunks, GraphRAG indexes entity-relation triples. The retrieval process:

  1. Finds relevant seed entities
  2. Traverses graph relationships (RAPTOR nodes → GraphRAG communities)
  3. Gathers connected facts into the context window

This handles questions like "What techniques did the authors of the BM25 paper later apply to neural retrieval?" — queries that span multiple document boundaries.

RAG Integration with User Memory

The repository extends RAG with two-tier memory in chapter3/contextual-retrieval-for-user-memory/. This combines:

  • Episodic memory — user-specific facts stored as memory cards
  • Semantic memory — general domain knowledge via standard RAG

The retrieval pipeline queries both tiers in parallel, allowing personalized responses grounded in both the user's history and external knowledge sources.

Key Implementation Files

Component Path Purpose
Retrieval pipeline entry chapter3/retrieval-pipeline/evaluate.py CLI for pure RAG evaluation
Score fusion & reranking chapter3/retrieval-pipeline/fusion.py Combines dense, sparse, and neural scores
Agentic decision layer chapter3/agentic-rag/main.py AgenticRAG class with should_retrieve()
Graph-based retrieval chapter3/structured-index/main.py GraphRAG with multi-hop traversal
Memory integration chapter3/contextual-retrieval-for-user-memory/README.md Two-tier memory architecture
Teaching materials slides/lesson-12.md, slides/lesson-13.md Design trade-offs and when to use each variant

Summary

  • RAG injects external knowledge into LLM prompts at inference time without model retraining.
  • The bojieli/ai-agent-book pipeline uses dense + sparse retrieval → neural reranker → prompt assembly.
  • Agentic RAG adds runtime decision-making to avoid unnecessary retrievals.
  • GraphRAG replaces flat chunks with structured entity graphs for multi-hop reasoning.
  • All components are runnable via the provided CLI and Python APIs in chapter3/.

Frequently Asked Questions

What is the difference between pure RAG and agentic RAG?

Pure RAG runs the retrieval pipeline on every user turn, ensuring maximum factuality at the cost of latency. Agentic RAG uses the AgenticRAG class in chapter3/agentic-rag/main.py to let the LLM judge whether external evidence is needed — reducing latency for questions answerable from parametric knowledge while preserving retrieval for uncertain claims.

When should I use GraphRAG instead of standard RAG?

Use GraphRAG when your queries require multi-hop reasoning across connected entities — for example, finding relationships between authors, papers, and techniques. Standard RAG with flat chunks fails on these boundary-spanning questions. The GraphRAG implementation in chapter3/structured-index/main.py indexes entity-relation triples and performs graph traversal before LLM generation.

How does the hybrid retrieval scoring work?

The fusion.py module combines three signals: (1) dense vector similarity from the neural encoder, (2) sparse BM25 keyword scores, and (3) cross-encoder reranker logits. Scores are normalized and fused into a single ranking. This outperforms any single method on the repository's evaluation benchmarks in chapter3/retrieval-pipeline/evaluate.py.

Can RAG work with user-specific memories?

Yes — the contextual retrieval for user memory implementation in chapter3/contextual-retrieval-for-user-memory/ demonstrates a two-tier architecture. The retrieval pipeline queries both a user memory store (personal facts) and a general knowledge index (domain documents), merging results before prompt construction.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →