How to Build User Memory and Knowledge Bases with RAG: A Complete Implementation Guide

Build user memory and knowledge bases with RAG by persisting conversation history as JSON chunks, indexing them with BM25 or dense retrieval, and exposing search tools to your LLM via OpenAI function-calling.

Retrieval-Augmented Generation (RAG) transforms static LLMs into systems that remember. In the bojieli/ai-agent-book repository, a production-ready pipeline demonstrates how to build user memory and knowledge bases with RAG—enabling agents to recall past conversations, retrieve relevant context, and generate informed responses. This guide walks through the four-layer architecture: persistence, chunking, indexing, and tool exposure.


Architecture Overview: Four Components Working Together

The RAG pipeline in this repository consists of tightly coupled components that handle the full lifecycle of user memory:

Component Purpose Location
User-memory store Serializes conversation history as JSON files chapter3/user-memory/memory_manager.py
Chunk model Represents conversation slices with metadata Referenced in indexer.py lines 18-20
RAG indexer Builds BM25 or dense indexes, executes search chapter3/agentic-rag-for-user-memory/indexer.py
Tool layer Exposes indexer to LLMs via function-calling chapter3/agentic-rag-for-user-memory/tools.py

Persisting User Memory to Disk

Long-term memory requires durability. The MemoryManager class writes each user's conversation history to a dedicated JSON file that survives process restarts.

In chapter3/user-memory/memory_manager.py, the storage path is constructed at initialization (line 81):

self.memory_file = os.path.join(Config.MEMORY_STORAGE_DIR,
                                f"{user_id}_memory.json")

By default, Config.MEMORY_STORAGE_DIR points to data/memories/ (defined in chapter3/user-memory/config.py). Each file follows the naming convention {user_id}_memory.json, making user isolation trivial.


Chunking Conversations for Retrieval

Raw conversation logs are too coarse for effective retrieval. The pipeline transforms them into ConversationChunk objects—discrete, searchable units with rich metadata.

The MemoryIndexer maintains two parallel data structures (lines 28-29 of indexer.py):

self.chunks: Dict[str, ConversationChunk]      # Original chunk objects

self.chunk_texts: Dict[str, str]                # Enriched texts for indexing

Enrichment adds searchable context to each chunk:

  • Test-case ID and conversation ID (Test Case: …, Conversation: …)
  • All metadata key/value pairs
  • The raw chunk text
  • Auto-generated semantic tags (financial, insurance, etc.) via _generate_semantic_tags (lines 49-91)

This enrichment ensures that retrieval matches on both explicit content and implicit categories.


Indexing: Local BM25 or External Pipeline

The MemoryIndexer supports two backends, selected at startup (lines 42-51):

Local BM25 Fallback

The LocalBM25Backend (lines 31-38) provides zero-dependency lexical search. It requires no network, no API keys, and no external services—ideal for development or air-gapped deployments.

External Retrieval Pipeline

For production scale, the indexer connects to an HTTP service at http://localhost:4242 (configurable). The service exposes three endpoints: /index, /search, and /clear. Health verification runs in _check_retrieval_pipeline (lines 59-71).

Adding chunks flows through add_chunks (lines 75-89), which updates both local storage and the chosen backend:

  • Local: _index_documents (lines 93-103)
  • Remote: _index_documents (lines 111-139)

After bulk ingestion, build_indexes (lines 144-172) synchronizes all stored chunks to the remote service.


Searching: The Core RAG Operation

The MemoryIndexer.search() method (lines 73-99) implements the retrieval logic that powers user memory and knowledge bases with RAG:

def search(self,
           query: str,
           top_k: int = 5,
           index_mode: IndexMode = None) -> List[SearchResult]:

Two execution paths:

  • Local mode → self.local_backend.search() (lines 103-114), results wrapped in SearchResult objects (lines 115-121)
  • Remote mode → POST to /search with mode parameter (dense, sparse, or hybrid), returns reranked results (lines 126-160)

The remote path includes a critical mapping step (lines 152-166): pipeline-generated doc_id values are resolved back to original chunk_id values via self.doc_id_mapping.

Each SearchResult contains:

  • The ConversationChunk object
  • BM25/dense/hybrid score
  • Match type annotation (local_bm25, dense, sparse, hybrid)

Exposing Memory to the LLM: The Tool Layer

Raw search results are useless without an interface. The MemoryTools class (lines 33-38 of tools.py) wraps the indexer in OpenAI-compatible function definitions.

Three tools form the complete API:

Tool Purpose Parameters
search_memory Lexical/semantic search over all chunks query (required, string)
get_conversation_context Fetch N surrounding chunks for context chunk_id (required), context_size (optional, default 2)
get_full_conversation Return all chunks for a conversation/test pair conversation_id, test_id (both required)

The OpenAI function specifications are auto-generated by get_tool_definitions() and consumed by the agent in agentic-rag-for-user-memory/agent.py.


Saving and Loading Indexes

Offline persistence enables pre-built knowledge bases. Two JSON files capture the complete index state:

  • {path}_chunks.json — serialized ConversationChunk objects
  • {path}_texts.json — pre-computed enriched texts

Save via save_index(path) (lines 94-115).
Load via load_index(path) (lines 117-172), which:

  1. Restores self.chunks dictionary
  2. Recreates ConversationChunk objects
  3. Regenerates self.chunk_texts if needed
  4. Calls self.build_indexes() to synchronize the backend

Complete End-to-End Example

This minimal script demonstrates the full pipeline for building user memory and knowledge bases with RAG:


# 1️⃣  Build the indexer

from indexer import MemoryIndexer, IndexConfig
from tools import MemoryTools, get_tool_definitions

indexer = MemoryIndexer(IndexConfig())               # Uses local BM25 by default

# 2️⃣  Load pre-saved conversation chunks

indexer.load_index()                                # Reads *_chunks.json & *_texts.json

# 3️⃣  Initialize the tool layer

tools = MemoryTools(indexer)

# 4️⃣  Search the memory base

result = tools.search_memory(query="How many flights did I book last year?")
print(result.to_dict())

# 5️⃣  Fetch surrounding context for a hit

if result.success and result.data["results"]:
    first_chunk_id = result.data["results"][0]["chunk_id"]
    ctx = tools.get_conversation_context(chunk_id=first_chunk_id, context_size=3)
    print(ctx.to_dict())

The ToolResult.to_dict() output drops directly into LLM prompts or OpenAI function_call payloads.


Extending the Pipeline

The repository includes alternative implementations for specialized use cases:

  • chapter3/structured-index/graphrag_indexer.py — GraphRAG-based structured indexing for rich semantic relationships
  • chapter3/contextual-retrieval-for-user-memory/ — Contextual indexer with evaluation tools and prompt templates

Swap retrieval_backend between "local" and "pipeline" to scale from laptop development to production clusters.


Summary

  • Persistence: MemoryManager writes per-user JSON files to data/memories/{user_id}_memory.json
  • Chunking: ConversationChunk objects with enriched metadata enable precise retrieval
  • Indexing: MemoryIndexer supports local BM25 or remote dense/sparse/hybrid pipelines
  • Search: search() method returns scored SearchResult objects with full provenance
  • Integration: MemoryTools exposes three OpenAI-compatible functions for agent consumption
  • Portability: Save/load indexes as JSON for reproducible, shareable knowledge bases

Frequently Asked Questions

What file format stores user memory in this RAG system?

User memory persists as JSON files following the pattern {user_id}_memory.json in the directory specified by Config.MEMORY_STORAGE_DIR (default: data/memories/). The MemoryManager class in chapter3/user-memory/memory_manager.py handles all read/write operations via standard Python JSON serialization.

Can this RAG pipeline run without external services?

Yes. The LocalBM25Backend (lines 31-38 of indexer.py) provides a pure-Python implementation requiring no network, API keys, or Docker containers. Set retrieval_backend="local" in IndexConfig for fully offline operation. The trade-off is lexical-only search versus the dense/hybrid modes available through the external pipeline.

How does the system prevent losing context when retrieving individual chunks?

The get_conversation_context tool fetches configurable surrounding chunks (default: 2 before and after). Given a chunk_id, it retrieves the neighboring ConversationChunk objects from self.chunks, ensuring the LLM sees conversational flow rather than isolated snippets.

What semantic modes does the external retrieval pipeline support?

The pipeline service at port 4242 accepts three index modes via the /search endpoint: dense (neural embeddings), sparse (inverted index), and hybrid (combined with reranking). These are passed through MemoryIndexer.search() as the index_mode parameter and translated to the JSON payload in lines 126-160.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →