What RAG and Vector Database Tools are Integrated in the agentbook `rag` Extra?

The rag optional dependency in bojieli/ai-agent-book bundles 12+ specialized packages across vector stores (ChromaDB, FAISS), embedding models (Sentence-Transformers, FlagEmbedding), sparse retrieval (BM25), ANN indexes (Annoy, HNSWLIB), and production utilities for token counting, logging, and async I/O.

The bojieli/ai-agent-book repository consolidates its entire Retrieval-Augmented Generation (RAG) stack into a single pip-installable extra. Defined in pyproject.toml at lines 146-168, this curated dependency group eliminates version conflicts by providing compatible vector search backends, embedding generators, and preprocessing libraries in one command.

Core Components of the rag Extra

The rag extra is organized into functional categories that cover the complete retrieval pipeline—from text tokenization to vector persistence.

Vector Stores and ANN Indexes

For vector storage and approximate nearest neighbor search, the extra includes:

  • ChromaDB (chromadb ≥ 0.5.0) — A persistent vector database with DuckDB+Parquet backend for production document storage
  • FAISS (faiss-cpu ≥ 1.7.4) — Facebook AI's high-performance library for in-memory similarity search on dense vectors
  • Annoy (annoy ≥ 1.17.3) — Spotify's tree-based ANN index for read-heavy workloads
  • HNSWLIB (hnswlib ≥ 0.8) — Header-only implementation of Hierarchical Navigable Small World graphs for fast approximate search

Embedding Models

To generate dense vector representations, the stack provides:

  • Sentence-Transformers (sentence-transformers ≥ 2.2.2) — Framework for state-of-the-art sentence and text embeddings
  • FlagEmbedding (FlagEmbedding ≥ 1.2.11) — BGE (BAAI General Embedding) model support for multilingual retrieval
  • Hugging Face Hub (huggingface-hub ≥ 0.34.0) — Client library for downloading and caching transformer models

Sparse Retrieval and Text Processing

For lexical retrieval and preprocessing:

  • Rank-BM25 (rank-bm25 ≥ 0.2.2) — Classical probabilistic retrieval for keyword-based search
  • Jieba (jieba ≥ 0.42.1) — Chinese text tokenization for multilingual RAG pipelines
  • NetworkX (networkx ≥ 3.2) — Graph-based text processing and analysis utilities
  • UMAP (umap-learn ≥ 0.5.4) — Dimensionality reduction for visualization and index optimization

Infrastructure and Utilities

Supporting infrastructure includes agentbook[tokens] (providing tiktoken for fast token counting), plus aiofiles ≥ 24.1.0 for async file operations, loguru ≥ 0.7.2 for structured logging, and markdown ≥ 3.5 for rendering results.

Installing the Complete RAG Stack

Install all integrated tools with a single command:

pip install "agentbook[rag]"

This resolves the dependency tree defined in pyproject.toml, ensuring compatibility between sentence-transformers, chromadb, and the ANN libraries.

Practical RAG Implementation Examples

Below are executable patterns demonstrating how the bundled tools interact in a retrieval pipeline.

Initializing a Persistent Chroma Vector Store

from chromadb import Client
from chromadb.config import Settings
from sentence_transformers import SentenceTransformer

# Initialise Chroma with persistence

client = Client(Settings(chroma_db_impl="duckdb+parquet", persist_directory="./chroma_db"))
collection = client.get_or_create_collection(name="my_docs")

# Load embedding model from the huggingface ecosystem

embedder = SentenceTransformer("all-MiniLM-L6-v2")

def add_documents(docs: list[str]):
    vectors = embedder.encode(docs, show_progress_bar=False).tolist()
    collection.add(
        documents=docs,
        embeddings=vectors,
        ids=[str(i) for i in range(len(docs))]
    )

Hybrid Retrieval with BM25 and FAISS

Combine sparse lexical search with dense vector similarity:

from rank_bm25 import BM25Okapi
from sentence_transformers import SentenceTransformer
import faiss
import numpy as np

# Setup BM25 on tokenized corpus

tokenized_corpus = [doc.lower().split() for doc in docs]
bm25 = BM25Okapi(tokenized_corpus)

def bm25_search(query: str, top_k: int = 5):
    tokenized_query = query.lower().split()
    scores = bm25.get_scores(tokenized_query)
    top_idxs = np.argsort(scores)[::-1][:top_k]
    return [docs[i] for i in top_idxs]

# Setup FAISS index for dense retrieval

embedder = SentenceTransformer("all-MiniLM-L6-v2")
dim = embedder.get_sentence_embedding_dimension()
index = faiss.IndexFlatL2(dim)
index.add(np.array(embedder.encode(docs)).astype("float32"))

def dense_search(query: str, top_k: int = 5):
    q_vec = embedder.encode([query]).astype("float32")
    _, I = index.search(q_vec, top_k)
    return [docs[i] for i in I[0]]

# Merge results (simple union-rank)

def hybrid_search(query: str, top_k: int = 5):
    sparse = bm25_search(query, top_k*2)
    dense = dense_search(query, top_k*2)
    combined = list(dict.fromkeys(sparse + dense))[:top_k]
    return combined

Production Logging with Loguru and Async I/O

from loguru import logger
import aiofiles
import asyncio

logger.add("rag.log", rotation="10 MB")

async def store_result(query: str, results: list[str]):
    async with aiofiles.open("rag_results.md", "a") as f:
        await f.write(f"## Query: {query}\n")

        for r in results:
            await f.write(f"- {r}\n")
        await f.write("\n")
    logger.info("Stored RAG results for query: {}", query)

# Execute async pipeline

asyncio.run(store_result("What is retrieval-augmented generation?", 
                         hybrid_search("retrieval augmented generation")))

Summary

  • The rag extra in pyproject.toml (lines 146-168) defines a complete RAG ecosystem bundling vector databases, embedding models, and retrieval algorithms
  • ChromaDB and FAISS provide persistent and in-memory vector storage respectively, while Annoy and HNSWLIB offer specialized ANN indexing
  • Sentence-Transformers and FlagEmbedding handle dense embeddings, complemented by Rank-BM25 for sparse lexical retrieval
  • Utility packages (loguru, aiofiles, agentbook[tokens]) support production-grade logging, async operations, and token accounting

Frequently Asked Questions

What vector databases are included in the agentbook rag extra?

The rag extra includes ChromaDB (≥0.5.0) for persistent document storage using DuckDB+Parquet backends, and FAISS-CPU (≥1.7.4) for high-performance in-memory similarity search. Additionally, Annoy (≥1.17.3) and HNSWLIB (≥0.8) provide tree-based and graph-based approximate nearest neighbor indexes for specialized retrieval scenarios.

How do I install the RAG dependencies for ai-agent-book?

Install the complete RAG stack by running pip install "agentbook[rag]". This command installs all 12+ packages declared in the rag extra of pyproject.toml, ensuring compatible versions between ChromaDB, Sentence-Transformers, FAISS, and the utility libraries.

Can I use both dense and sparse retrieval methods together?

Yes. The bundled Rank-BM25 library enables sparse lexical retrieval, while Sentence-Transformers generates dense embeddings for FAISS or ChromaDB. You can implement hybrid retrieval by running queries through both pipelines and merging results using rank fusion techniques, as shown in the code examples above.

What embedding models are available in the rag extra?

The stack provides Sentence-Transformers (≥2.2.2) for general-purpose embeddings, FlagEmbedding (≥1.2.11) for BGE multilingual models, and Hugging Face Hub (≥0.34.0) for accessing the broader transformer ecosystem. These integrate with the agentbook[tokens] extra for accurate token counting when chunking documents.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →