# How to Build User Memory and Knowledge Bases with RAG: A Complete Implementation Guide

> Learn to build user memory and knowledge bases with RAG. Persist conversation history, index it, and use OpenAI function-calling for LLM access. A complete implementation guide.

- Repository: [Bojie Li/ai-agent-book](https://github.com/bojieli/ai-agent-book)
- Tags: how-to-guide
- Published: 2026-08-06

---

**Build user memory and knowledge bases with RAG by persisting conversation history as JSON chunks, indexing them with BM25 or dense retrieval, and exposing search tools to your LLM via OpenAI function-calling.**

Retrieval-Augmented Generation (RAG) transforms static LLMs into systems that remember. In the `bojieli/ai-agent-book` repository, a production-ready pipeline demonstrates how to build **user memory and knowledge bases with RAG**—enabling agents to recall past conversations, retrieve relevant context, and generate informed responses. This guide walks through the four-layer architecture: persistence, chunking, indexing, and tool exposure.

---

## Architecture Overview: Four Components Working Together

The RAG pipeline in this repository consists of tightly coupled components that handle the full lifecycle of user memory:

| Component | Purpose | Location |
|-----------|---------|----------|
| **User-memory store** | Serializes conversation history as JSON files | [`chapter3/user-memory/memory_manager.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter3/user-memory/memory_manager.py) |
| **Chunk model** | Represents conversation slices with metadata | Referenced in [`indexer.py`](https://github.com/bojieli/ai-agent-book/blob/main/indexer.py) lines 18-20 |
| **RAG indexer** | Builds BM25 or dense indexes, executes search | [`chapter3/agentic-rag-for-user-memory/indexer.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter3/agentic-rag-for-user-memory/indexer.py) |
| **Tool layer** | Exposes indexer to LLMs via function-calling | [`chapter3/agentic-rag-for-user-memory/tools.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter3/agentic-rag-for-user-memory/tools.py) |

---

## Persisting User Memory to Disk

Long-term memory requires durability. The `MemoryManager` class writes each user's conversation history to a dedicated JSON file that survives process restarts.

In [`chapter3/user-memory/memory_manager.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter3/user-memory/memory_manager.py), the storage path is constructed at initialization (line 81):

```python
self.memory_file = os.path.join(Config.MEMORY_STORAGE_DIR,
                                f"{user_id}_memory.json")

```

By default, `Config.MEMORY_STORAGE_DIR` points to `data/memories/` (defined in [`chapter3/user-memory/config.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter3/user-memory/config.py)). Each file follows the naming convention `{user_id}_memory.json`, making user isolation trivial.

---

## Chunking Conversations for Retrieval

Raw conversation logs are too coarse for effective retrieval. The pipeline transforms them into `ConversationChunk` objects—discrete, searchable units with rich metadata.

The `MemoryIndexer` maintains two parallel data structures (lines 28-29 of [`indexer.py`](https://github.com/bojieli/ai-agent-book/blob/main/indexer.py)):

```python
self.chunks: Dict[str, ConversationChunk]      # Original chunk objects

self.chunk_texts: Dict[str, str]                # Enriched texts for indexing

```

**Enrichment** adds searchable context to each chunk:

- Test-case ID and conversation ID (`Test Case: …`, `Conversation: …`)
- All metadata key/value pairs
- The raw chunk text
- Auto-generated semantic tags (financial, insurance, etc.) via `_generate_semantic_tags` (lines 49-91)

This enrichment ensures that retrieval matches on both explicit content and implicit categories.

---

## Indexing: Local BM25 or External Pipeline

The `MemoryIndexer` supports two backends, selected at startup (lines 42-51):

### Local BM25 Fallback

The `LocalBM25Backend` (lines 31-38) provides zero-dependency lexical search. It requires no network, no API keys, and no external services—ideal for development or air-gapped deployments.

### External Retrieval Pipeline

For production scale, the indexer connects to an HTTP service at `http://localhost:4242` (configurable). The service exposes three endpoints: `/index`, `/search`, and `/clear`. Health verification runs in `_check_retrieval_pipeline` (lines 59-71).

**Adding chunks** flows through `add_chunks` (lines 75-89), which updates both local storage and the chosen backend:

- Local: `_index_documents` (lines 93-103)
- Remote: `_index_documents` (lines 111-139)

After bulk ingestion, `build_indexes` (lines 144-172) synchronizes all stored chunks to the remote service.

---

## Searching: The Core RAG Operation

The `MemoryIndexer.search()` method (lines 73-99) implements the retrieval logic that powers user memory and knowledge bases with RAG:

```python
def search(self,
           query: str,
           top_k: int = 5,
           index_mode: IndexMode = None) -> List[SearchResult]:

```

**Two execution paths:**

- **Local mode** → `self.local_backend.search()` (lines 103-114), results wrapped in `SearchResult` objects (lines 115-121)
- **Remote mode** → POST to `/search` with `mode` parameter (`dense`, `sparse`, or `hybrid`), returns reranked results (lines 126-160)

The remote path includes a critical mapping step (lines 152-166): pipeline-generated `doc_id` values are resolved back to original `chunk_id` values via `self.doc_id_mapping`.

Each `SearchResult` contains:
- The `ConversationChunk` object
- BM25/dense/hybrid score
- Match type annotation (`local_bm25`, `dense`, `sparse`, `hybrid`)

---

## Exposing Memory to the LLM: The Tool Layer

Raw search results are useless without an interface. The `MemoryTools` class (lines 33-38 of [`tools.py`](https://github.com/bojieli/ai-agent-book/blob/main/tools.py)) wraps the indexer in OpenAI-compatible function definitions.

Three tools form the complete API:

| Tool | Purpose | Parameters |
|------|---------|------------|
| `search_memory` | Lexical/semantic search over all chunks | `query` (required, string) |
| `get_conversation_context` | Fetch N surrounding chunks for context | `chunk_id` (required), `context_size` (optional, default 2) |
| `get_full_conversation` | Return all chunks for a conversation/test pair | `conversation_id`, `test_id` (both required) |

The OpenAI function specifications are auto-generated by `get_tool_definitions()` and consumed by the agent in [`agentic-rag-for-user-memory/agent.py`](https://github.com/bojieli/ai-agent-book/blob/main/agentic-rag-for-user-memory/agent.py).

---

## Saving and Loading Indexes

Offline persistence enables pre-built knowledge bases. Two JSON files capture the complete index state:

- `{path}_chunks.json` — serialized `ConversationChunk` objects
- `{path}_texts.json` — pre-computed enriched texts

**Save** via `save_index(path)` (lines 94-115).  
**Load** via `load_index(path)` (lines 117-172), which:
1. Restores `self.chunks` dictionary
2. Recreates `ConversationChunk` objects
3. Regenerates `self.chunk_texts` if needed
4. Calls `self.build_indexes()` to synchronize the backend

---

## Complete End-to-End Example

This minimal script demonstrates the full pipeline for building user memory and knowledge bases with RAG:

```python

# 1️⃣  Build the indexer

from indexer import MemoryIndexer, IndexConfig
from tools import MemoryTools, get_tool_definitions

indexer = MemoryIndexer(IndexConfig())               # Uses local BM25 by default

# 2️⃣  Load pre-saved conversation chunks

indexer.load_index()                                # Reads *_chunks.json & *_texts.json

# 3️⃣  Initialize the tool layer

tools = MemoryTools(indexer)

# 4️⃣  Search the memory base

result = tools.search_memory(query="How many flights did I book last year?")
print(result.to_dict())

# 5️⃣  Fetch surrounding context for a hit

if result.success and result.data["results"]:
    first_chunk_id = result.data["results"][0]["chunk_id"]
    ctx = tools.get_conversation_context(chunk_id=first_chunk_id, context_size=3)
    print(ctx.to_dict())

```

The `ToolResult.to_dict()` output drops directly into LLM prompts or OpenAI `function_call` payloads.

---

## Extending the Pipeline

The repository includes alternative implementations for specialized use cases:

- **[`chapter3/structured-index/graphrag_indexer.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter3/structured-index/graphrag_indexer.py)** — GraphRAG-based structured indexing for rich semantic relationships
- **`chapter3/contextual-retrieval-for-user-memory/`** — Contextual indexer with evaluation tools and prompt templates

Swap `retrieval_backend` between `"local"` and `"pipeline"` to scale from laptop development to production clusters.

---

## Summary

- **Persistence**: `MemoryManager` writes per-user JSON files to `data/memories/{user_id}_memory.json`
- **Chunking**: `ConversationChunk` objects with enriched metadata enable precise retrieval
- **Indexing**: `MemoryIndexer` supports local BM25 or remote dense/sparse/hybrid pipelines
- **Search**: `search()` method returns scored `SearchResult` objects with full provenance
- **Integration**: `MemoryTools` exposes three OpenAI-compatible functions for agent consumption
- **Portability**: Save/load indexes as JSON for reproducible, shareable knowledge bases

---

## Frequently Asked Questions

### What file format stores user memory in this RAG system?

User memory persists as JSON files following the pattern `{user_id}_memory.json` in the directory specified by `Config.MEMORY_STORAGE_DIR` (default: `data/memories/`). The `MemoryManager` class in [`chapter3/user-memory/memory_manager.py`](https://github.com/bojieli/ai-agent-book/blob/main/chapter3/user-memory/memory_manager.py) handles all read/write operations via standard Python JSON serialization.

### Can this RAG pipeline run without external services?

Yes. The `LocalBM25Backend` (lines 31-38 of [`indexer.py`](https://github.com/bojieli/ai-agent-book/blob/main/indexer.py)) provides a pure-Python implementation requiring no network, API keys, or Docker containers. Set `retrieval_backend="local"` in `IndexConfig` for fully offline operation. The trade-off is lexical-only search versus the dense/hybrid modes available through the external pipeline.

### How does the system prevent losing context when retrieving individual chunks?

The `get_conversation_context` tool fetches configurable surrounding chunks (default: 2 before and after). Given a `chunk_id`, it retrieves the neighboring `ConversationChunk` objects from `self.chunks`, ensuring the LLM sees conversational flow rather than isolated snippets.

### What semantic modes does the external retrieval pipeline support?

The pipeline service at port 4242 accepts three index modes via the `/search` endpoint: `dense` (neural embeddings), `sparse` (inverted index), and `hybrid` (combined with reranking). These are passed through `MemoryIndexer.search()` as the `index_mode` parameter and translated to the JSON payload in lines 126-160.