What RAG and Vector Database Tools are Integrated in the agentbook `rag` Extra?
The rag optional dependency in bojieli/ai-agent-book bundles 12+ specialized packages across vector stores (ChromaDB, FAISS), embedding models (Sentence-Transformers, FlagEmbedding), sparse retrieval (BM25), ANN indexes (Annoy, HNSWLIB), and production utilities for token counting, logging, and async I/O.
The bojieli/ai-agent-book repository consolidates its entire Retrieval-Augmented Generation (RAG) stack into a single pip-installable extra. Defined in pyproject.toml at lines 146-168, this curated dependency group eliminates version conflicts by providing compatible vector search backends, embedding generators, and preprocessing libraries in one command.
Core Components of the rag Extra
The rag extra is organized into functional categories that cover the complete retrieval pipeline—from text tokenization to vector persistence.
Vector Stores and ANN Indexes
For vector storage and approximate nearest neighbor search, the extra includes:
- ChromaDB (
chromadb≥ 0.5.0) — A persistent vector database with DuckDB+Parquet backend for production document storage - FAISS (
faiss-cpu≥ 1.7.4) — Facebook AI's high-performance library for in-memory similarity search on dense vectors - Annoy (
annoy≥ 1.17.3) — Spotify's tree-based ANN index for read-heavy workloads - HNSWLIB (
hnswlib≥ 0.8) — Header-only implementation of Hierarchical Navigable Small World graphs for fast approximate search
Embedding Models
To generate dense vector representations, the stack provides:
- Sentence-Transformers (
sentence-transformers≥ 2.2.2) — Framework for state-of-the-art sentence and text embeddings - FlagEmbedding (
FlagEmbedding≥ 1.2.11) — BGE (BAAI General Embedding) model support for multilingual retrieval - Hugging Face Hub (
huggingface-hub≥ 0.34.0) — Client library for downloading and caching transformer models
Sparse Retrieval and Text Processing
For lexical retrieval and preprocessing:
- Rank-BM25 (
rank-bm25≥ 0.2.2) — Classical probabilistic retrieval for keyword-based search - Jieba (
jieba≥ 0.42.1) — Chinese text tokenization for multilingual RAG pipelines - NetworkX (
networkx≥ 3.2) — Graph-based text processing and analysis utilities - UMAP (
umap-learn≥ 0.5.4) — Dimensionality reduction for visualization and index optimization
Infrastructure and Utilities
Supporting infrastructure includes agentbook[tokens] (providing tiktoken for fast token counting), plus aiofiles ≥ 24.1.0 for async file operations, loguru ≥ 0.7.2 for structured logging, and markdown ≥ 3.5 for rendering results.
Installing the Complete RAG Stack
Install all integrated tools with a single command:
pip install "agentbook[rag]"
This resolves the dependency tree defined in pyproject.toml, ensuring compatibility between sentence-transformers, chromadb, and the ANN libraries.
Practical RAG Implementation Examples
Below are executable patterns demonstrating how the bundled tools interact in a retrieval pipeline.
Initializing a Persistent Chroma Vector Store
from chromadb import Client
from chromadb.config import Settings
from sentence_transformers import SentenceTransformer
# Initialise Chroma with persistence
client = Client(Settings(chroma_db_impl="duckdb+parquet", persist_directory="./chroma_db"))
collection = client.get_or_create_collection(name="my_docs")
# Load embedding model from the huggingface ecosystem
embedder = SentenceTransformer("all-MiniLM-L6-v2")
def add_documents(docs: list[str]):
vectors = embedder.encode(docs, show_progress_bar=False).tolist()
collection.add(
documents=docs,
embeddings=vectors,
ids=[str(i) for i in range(len(docs))]
)
Hybrid Retrieval with BM25 and FAISS
Combine sparse lexical search with dense vector similarity:
from rank_bm25 import BM25Okapi
from sentence_transformers import SentenceTransformer
import faiss
import numpy as np
# Setup BM25 on tokenized corpus
tokenized_corpus = [doc.lower().split() for doc in docs]
bm25 = BM25Okapi(tokenized_corpus)
def bm25_search(query: str, top_k: int = 5):
tokenized_query = query.lower().split()
scores = bm25.get_scores(tokenized_query)
top_idxs = np.argsort(scores)[::-1][:top_k]
return [docs[i] for i in top_idxs]
# Setup FAISS index for dense retrieval
embedder = SentenceTransformer("all-MiniLM-L6-v2")
dim = embedder.get_sentence_embedding_dimension()
index = faiss.IndexFlatL2(dim)
index.add(np.array(embedder.encode(docs)).astype("float32"))
def dense_search(query: str, top_k: int = 5):
q_vec = embedder.encode([query]).astype("float32")
_, I = index.search(q_vec, top_k)
return [docs[i] for i in I[0]]
# Merge results (simple union-rank)
def hybrid_search(query: str, top_k: int = 5):
sparse = bm25_search(query, top_k*2)
dense = dense_search(query, top_k*2)
combined = list(dict.fromkeys(sparse + dense))[:top_k]
return combined
Production Logging with Loguru and Async I/O
from loguru import logger
import aiofiles
import asyncio
logger.add("rag.log", rotation="10 MB")
async def store_result(query: str, results: list[str]):
async with aiofiles.open("rag_results.md", "a") as f:
await f.write(f"## Query: {query}\n")
for r in results:
await f.write(f"- {r}\n")
await f.write("\n")
logger.info("Stored RAG results for query: {}", query)
# Execute async pipeline
asyncio.run(store_result("What is retrieval-augmented generation?",
hybrid_search("retrieval augmented generation")))
Summary
- The
ragextra inpyproject.toml(lines 146-168) defines a complete RAG ecosystem bundling vector databases, embedding models, and retrieval algorithms - ChromaDB and FAISS provide persistent and in-memory vector storage respectively, while Annoy and HNSWLIB offer specialized ANN indexing
- Sentence-Transformers and FlagEmbedding handle dense embeddings, complemented by Rank-BM25 for sparse lexical retrieval
- Utility packages (
loguru,aiofiles,agentbook[tokens]) support production-grade logging, async operations, and token accounting
Frequently Asked Questions
What vector databases are included in the agentbook rag extra?
The rag extra includes ChromaDB (≥0.5.0) for persistent document storage using DuckDB+Parquet backends, and FAISS-CPU (≥1.7.4) for high-performance in-memory similarity search. Additionally, Annoy (≥1.17.3) and HNSWLIB (≥0.8) provide tree-based and graph-based approximate nearest neighbor indexes for specialized retrieval scenarios.
How do I install the RAG dependencies for ai-agent-book?
Install the complete RAG stack by running pip install "agentbook[rag]". This command installs all 12+ packages declared in the rag extra of pyproject.toml, ensuring compatible versions between ChromaDB, Sentence-Transformers, FAISS, and the utility libraries.
Can I use both dense and sparse retrieval methods together?
Yes. The bundled Rank-BM25 library enables sparse lexical retrieval, while Sentence-Transformers generates dense embeddings for FAISS or ChromaDB. You can implement hybrid retrieval by running queries through both pipelines and merging results using rank fusion techniques, as shown in the code examples above.
What embedding models are available in the rag extra?
The stack provides Sentence-Transformers (≥2.2.2) for general-purpose embeddings, FlagEmbedding (≥1.2.11) for BGE multilingual models, and Hugging Face Hub (≥0.34.0) for accessing the broader transformer ecosystem. These integrate with the agentbook[tokens] extra for accurate token counting when chunking documents.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →