# Python Packages Used for Sentence Embeddings in the AI-Agent-Book Repository

> Discover the Python packages sentence-transformers and faiss powering sentence embeddings in the bojieli/ai-agent-book repository. Learn how they enable efficient indexing and searching.

- Repository: [Bojie Li/ai-agent-book](https://github.com/bojieli/ai-agent-book)
- Tags: tutorial
- Published: 2026-08-22

---

**The AI-Agent-Book repository uses `sentence-transformers` to generate sentence embeddings and `faiss` to index and search them efficiently.**

The `bojieli/ai-agent-book` project implements a semantic search system that converts text into dense vector representations. Understanding which Python packages handle sentence embeddings in this codebase reveals how the project transforms unstructured text into searchable vectors. The implementation separates embedding generation from vector storage for optimal performance.

## Sentence-Transformers for Embedding Generation

In [`chapter9/gaia-experience/knowledge_base.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter9/gaia-experience/knowledge_base.py), the repository imports the `SentenceTransformer` class from the `sentence-transformers` library at lines 12-15. This package serves as the primary engine for converting natural language sentences into high-dimensional numerical vectors suitable for similarity comparison.

### Loading the Pre-trained Model

The codebase utilizes the `all-MiniLM-L6-v2` model by default, which produces 384-dimensional embeddings. The `SentenceTransformer` class loads the model weights and provides the `encode()` method to transform input text into numpy arrays.

```python
from sentence_transformers import SentenceTransformer

# Load a pre-trained model (the default in the repo is "all-MiniLM-L6-v2")

model = SentenceTransformer('all-MiniLM-L6-v2')

sentences = [
    "The quick brown fox jumps over the lazy dog.",
    "Artificial intelligence is transforming many industries."
]

# Compute a 384-dimensional embedding for each sentence

embeddings = model.encode(sentences)

print(embeddings.shape)   # (2, 384)

print(embeddings[0][:5]) # first five values of the first embedding

```

## FAISS for Vector Storage and Retrieval

While `sentence-transformers` creates the embeddings, the repository employs `faiss` (Facebook AI Similarity Search) for efficient storage and querying. The import appears at lines 16-19 of [`knowledge_base.py`](https://github.com/bojieli/ai-agent-book/blob/main/knowledge_base.py). Unlike the embedding generator, `faiss` provides optimized approximate nearest neighbor search capabilities for high-dimensional vectors.

### Building the Search Index

The implementation creates an `IndexFlatL2` instance to store vectors using L2 (Euclidean) distance calculations. This separation of concerns allows the system to generate embeddings once and query them repeatedly without recomputation.

```python
import faiss
import numpy as np
from sentence_transformers import SentenceTransformer

model = SentenceTransformer('all-MiniLM-L6-v2')
sentences = ["first sentence", "second sentence", "third sentence"]
embeds = model.encode(sentences)               # → (3, 384)

dim = embeds.shape[1]                           # 384

index = faiss.IndexFlatL2(dim)                  # L2 distance index

index.add(embeds)                               # store vectors

# Query the index

query = model.encode(["search query"])
dist, idx = index.search(query, k=2)            # retrieve 2 nearest neighbours

print(idx)                                      # indices of closest sentences

```

## Key Implementation Files

The following files demonstrate how these Python packages work together for sentence embeddings:

- **[`chapter9/gaia-experience/knowledge_base.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter9/gaia-experience/knowledge_base.py)**: Contains the conditional imports of `SentenceTransformer` and `faiss`, along with the logic that builds, loads, and queries the embedding index.
- **[`tests/test_kb_topk_zero.py`](https://github.com/bojieli/ai-agent-book/blob/main/tests/test_kb_topk_zero.py)**: Unit tests that mock the embedding model and FAISS index, confirming behavior when dependencies are present.
- **[`tests/test_kb_faiss_negative_idx.py`](https://github.com/bojieli/ai-agent-book/blob/main/tests/test_kb_faiss_negative_idx.py)**: Additional test coverage for edge cases involving FAISS index operations and embedding dimension mismatches.

## Summary

- **`sentence-transformers`**: Generates 384-dimensional sentence embeddings using the `all-MiniLM-L6-v2` model, as implemented in [`knowledge_base.py`](https://github.com/bojieli/ai-agent-book/blob/main/knowledge_base.py).
- **`faiss`**: Provides vector indexing and similarity search capabilities, imported at lines 16-19 of the knowledge base module.
- **Separation of concerns**: The repository distinguishes between embedding generation (computationally expensive, done once) and vector search (performed frequently on stored embeddings).
- **Integration pattern**: The `KnowledgeBase` class combines these packages to enable semantic search over document collections.

## Frequently Asked Questions

### Which Python package generates sentence embeddings in the AI-Agent-Book repository?

The `sentence-transformers` library generates all sentence embeddings. Specifically, the `SentenceTransformer` class is imported in [`chapter9/gaia-experience/knowledge_base.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter9/gaia-experience/knowledge_base.py) and used to encode text into 384-dimensional vectors using the `all-MiniLM-L6-v2` model.

### Does the repository use FAISS to create embeddings?

No. `faiss` is used exclusively for storing and searching pre-generated embeddings. The package provides efficient vector indexing and approximate nearest neighbor search, but it does not generate embeddings itself. The embeddings are created by `sentence-transformers` before being added to the FAISS index.

### What embedding model does the knowledge base use by default?

The implementation defaults to the `all-MiniLM-L6-v2` model from the MiniLM family. This model produces 384-dimensional embeddings and balances computational efficiency with semantic accuracy, making it suitable for document retrieval tasks within the knowledge base system.

### How are these packages integrated in the codebase?

The `KnowledgeBase` class in [`knowledge_base.py`](https://github.com/bojieli/ai-agent-book/blob/main/knowledge_base.py) initializes the `SentenceTransformer` model to encode queries and documents, then stores the resulting vectors in a `faiss.IndexFlatL2`. When searching, it encodes the query using `sentence-transformers` and retrieves similar vectors using `faiss.search()`, returning the indices of the most relevant sentences.