# What RAG and Vector Database Tools are Integrated in the agentbook `rag` Extra?

> Explore RAG tools in the agentbook `rag` extra. Discover integrations for vector stores like ChromaDB FAISS, embedding models, sparse retrieval BM25, ANN indexes, and production utilities. Boost your AI agent capabilities.

- Repository: [Bojie Li/ai-agent-book](https://github.com/bojieli/ai-agent-book)
- Tags: deep-dive
- Published: 2026-08-23

---

**The `rag` optional dependency in `bojieli/ai-agent-book` bundles 12+ specialized packages across vector stores (ChromaDB, FAISS), embedding models (Sentence-Transformers, FlagEmbedding), sparse retrieval (BM25), ANN indexes (Annoy, HNSWLIB), and production utilities for token counting, logging, and async I/O.**

The `bojieli/ai-agent-book` repository consolidates its entire Retrieval-Augmented Generation (RAG) stack into a single pip-installable extra. Defined in [`pyproject.toml`](https://github.com/bojieli/ai-agent-book/blob/main/pyproject.toml) at lines 146-168, this curated dependency group eliminates version conflicts by providing compatible vector search backends, embedding generators, and preprocessing libraries in one command.

## Core Components of the rag Extra

The `rag` extra is organized into functional categories that cover the complete retrieval pipeline—from text tokenization to vector persistence.

### Vector Stores and ANN Indexes

For **vector storage** and **approximate nearest neighbor** search, the extra includes:

- **ChromaDB** (`chromadb` ≥ 0.5.0) — A persistent vector database with DuckDB+Parquet backend for production document storage
- **FAISS** (`faiss-cpu` ≥ 1.7.4) — Facebook AI's high-performance library for in-memory similarity search on dense vectors
- **Annoy** (`annoy` ≥ 1.17.3) — Spotify's tree-based ANN index for read-heavy workloads
- **HNSWLIB** (`hnswlib` ≥ 0.8) — Header-only implementation of Hierarchical Navigable Small World graphs for fast approximate search

### Embedding Models

To generate dense vector representations, the stack provides:

- **Sentence-Transformers** (`sentence-transformers` ≥ 2.2.2) — Framework for state-of-the-art sentence and text embeddings
- **FlagEmbedding** (`FlagEmbedding` ≥ 1.2.11) — BGE (BAAI General Embedding) model support for multilingual retrieval
- **Hugging Face Hub** (`huggingface-hub` ≥ 0.34.0) — Client library for downloading and caching transformer models

### Sparse Retrieval and Text Processing

For **lexical retrieval** and preprocessing:

- **Rank-BM25** (`rank-bm25` ≥ 0.2.2) — Classical probabilistic retrieval for keyword-based search
- **Jieba** (`jieba` ≥ 0.42.1) — Chinese text tokenization for multilingual RAG pipelines
- **NetworkX** (`networkx` ≥ 3.2) — Graph-based text processing and analysis utilities
- **UMAP** (`umap-learn` ≥ 0.5.4) — Dimensionality reduction for visualization and index optimization

### Infrastructure and Utilities

Supporting infrastructure includes **agentbook[tokens]** (providing `tiktoken` for fast token counting), plus `aiofiles` ≥ 24.1.0 for async file operations, `loguru` ≥ 0.7.2 for structured logging, and `markdown` ≥ 3.5 for rendering results.

## Installing the Complete RAG Stack

Install all integrated tools with a single command:

```bash
pip install "agentbook[rag]"

```

This resolves the dependency tree defined in [`pyproject.toml`](https://github.com/bojieli/ai-agent-book/blob/main/pyproject.toml), ensuring compatibility between `sentence-transformers`, `chromadb`, and the ANN libraries.

## Practical RAG Implementation Examples

Below are executable patterns demonstrating how the bundled tools interact in a retrieval pipeline.

### Initializing a Persistent Chroma Vector Store

```python
from chromadb import Client
from chromadb.config import Settings
from sentence_transformers import SentenceTransformer

# Initialise Chroma with persistence

client = Client(Settings(chroma_db_impl="duckdb+parquet", persist_directory="./chroma_db"))
collection = client.get_or_create_collection(name="my_docs")

# Load embedding model from the huggingface ecosystem

embedder = SentenceTransformer("all-MiniLM-L6-v2")

def add_documents(docs: list[str]):
    vectors = embedder.encode(docs, show_progress_bar=False).tolist()
    collection.add(
        documents=docs,
        embeddings=vectors,
        ids=[str(i) for i in range(len(docs))]
    )

```

### Hybrid Retrieval with BM25 and FAISS

Combine sparse lexical search with dense vector similarity:

```python
from rank_bm25 import BM25Okapi
from sentence_transformers import SentenceTransformer
import faiss
import numpy as np

# Setup BM25 on tokenized corpus

tokenized_corpus = [doc.lower().split() for doc in docs]
bm25 = BM25Okapi(tokenized_corpus)

def bm25_search(query: str, top_k: int = 5):
    tokenized_query = query.lower().split()
    scores = bm25.get_scores(tokenized_query)
    top_idxs = np.argsort(scores)[::-1][:top_k]
    return [docs[i] for i in top_idxs]

# Setup FAISS index for dense retrieval

embedder = SentenceTransformer("all-MiniLM-L6-v2")
dim = embedder.get_sentence_embedding_dimension()
index = faiss.IndexFlatL2(dim)
index.add(np.array(embedder.encode(docs)).astype("float32"))

def dense_search(query: str, top_k: int = 5):
    q_vec = embedder.encode([query]).astype("float32")
    _, I = index.search(q_vec, top_k)
    return [docs[i] for i in I[0]]

# Merge results (simple union-rank)

def hybrid_search(query: str, top_k: int = 5):
    sparse = bm25_search(query, top_k*2)
    dense = dense_search(query, top_k*2)
    combined = list(dict.fromkeys(sparse + dense))[:top_k]
    return combined

```

### Production Logging with Loguru and Async I/O

```python
from loguru import logger
import aiofiles
import asyncio

logger.add("rag.log", rotation="10 MB")

async def store_result(query: str, results: list[str]):
    async with aiofiles.open("rag_results.md", "a") as f:
        await f.write(f"## Query: {query}\n")

        for r in results:
            await f.write(f"- {r}\n")
        await f.write("\n")
    logger.info("Stored RAG results for query: {}", query)

# Execute async pipeline

asyncio.run(store_result("What is retrieval-augmented generation?", 
                         hybrid_search("retrieval augmented generation")))

```

## Summary

- The `rag` extra in [`pyproject.toml`](https://github.com/bojieli/ai-agent-book/blob/main/pyproject.toml) (lines 146-168) defines a complete RAG ecosystem bundling vector databases, embedding models, and retrieval algorithms
- **ChromaDB** and **FAISS** provide persistent and in-memory vector storage respectively, while **Annoy** and **HNSWLIB** offer specialized ANN indexing
- **Sentence-Transformers** and **FlagEmbedding** handle dense embeddings, complemented by **Rank-BM25** for sparse lexical retrieval
- Utility packages (`loguru`, `aiofiles`, `agentbook[tokens]`) support production-grade logging, async operations, and token accounting

## Frequently Asked Questions

### What vector databases are included in the agentbook rag extra?

The `rag` extra includes **ChromaDB** (≥0.5.0) for persistent document storage using DuckDB+Parquet backends, and **FAISS-CPU** (≥1.7.4) for high-performance in-memory similarity search. Additionally, **Annoy** (≥1.17.3) and **HNSWLIB** (≥0.8) provide tree-based and graph-based approximate nearest neighbor indexes for specialized retrieval scenarios.

### How do I install the RAG dependencies for ai-agent-book?

Install the complete RAG stack by running `pip install "agentbook[rag]"`. This command installs all 12+ packages declared in the `rag` extra of [`pyproject.toml`](https://github.com/bojieli/ai-agent-book/blob/main/pyproject.toml), ensuring compatible versions between ChromaDB, Sentence-Transformers, FAISS, and the utility libraries.

### Can I use both dense and sparse retrieval methods together?

Yes. The bundled **Rank-BM25** library enables sparse lexical retrieval, while **Sentence-Transformers** generates dense embeddings for FAISS or ChromaDB. You can implement hybrid retrieval by running queries through both pipelines and merging results using rank fusion techniques, as shown in the code examples above.

### What embedding models are available in the rag extra?

The stack provides **Sentence-Transformers** (≥2.2.2) for general-purpose embeddings, **FlagEmbedding** (≥1.2.11) for BGE multilingual models, and **Hugging Face Hub** (≥0.34.0) for accessing the broader transformer ecosystem. These integrate with the **agentbook[tokens]** extra for accurate token counting when chunking documents.