How to Build User Memory and Knowledge Bases with RAG: A Complete Implementation Guide
Build user memory and knowledge bases with RAG by persisting conversation history as JSON chunks, indexing them with BM25 or dense retrieval, and exposing search tools to your LLM via OpenAI function-calling.
Retrieval-Augmented Generation (RAG) transforms static LLMs into systems that remember. In the bojieli/ai-agent-book repository, a production-ready pipeline demonstrates how to build user memory and knowledge bases with RAG—enabling agents to recall past conversations, retrieve relevant context, and generate informed responses. This guide walks through the four-layer architecture: persistence, chunking, indexing, and tool exposure.
Architecture Overview: Four Components Working Together
The RAG pipeline in this repository consists of tightly coupled components that handle the full lifecycle of user memory:
| Component | Purpose | Location |
|---|---|---|
| User-memory store | Serializes conversation history as JSON files | chapter3/user-memory/memory_manager.py |
| Chunk model | Represents conversation slices with metadata | Referenced in indexer.py lines 18-20 |
| RAG indexer | Builds BM25 or dense indexes, executes search | chapter3/agentic-rag-for-user-memory/indexer.py |
| Tool layer | Exposes indexer to LLMs via function-calling | chapter3/agentic-rag-for-user-memory/tools.py |
Persisting User Memory to Disk
Long-term memory requires durability. The MemoryManager class writes each user's conversation history to a dedicated JSON file that survives process restarts.
In chapter3/user-memory/memory_manager.py, the storage path is constructed at initialization (line 81):
self.memory_file = os.path.join(Config.MEMORY_STORAGE_DIR,
f"{user_id}_memory.json")
By default, Config.MEMORY_STORAGE_DIR points to data/memories/ (defined in chapter3/user-memory/config.py). Each file follows the naming convention {user_id}_memory.json, making user isolation trivial.
Chunking Conversations for Retrieval
Raw conversation logs are too coarse for effective retrieval. The pipeline transforms them into ConversationChunk objects—discrete, searchable units with rich metadata.
The MemoryIndexer maintains two parallel data structures (lines 28-29 of indexer.py):
self.chunks: Dict[str, ConversationChunk] # Original chunk objects
self.chunk_texts: Dict[str, str] # Enriched texts for indexing
Enrichment adds searchable context to each chunk:
- Test-case ID and conversation ID (
Test Case: …,Conversation: …) - All metadata key/value pairs
- The raw chunk text
- Auto-generated semantic tags (financial, insurance, etc.) via
_generate_semantic_tags(lines 49-91)
This enrichment ensures that retrieval matches on both explicit content and implicit categories.
Indexing: Local BM25 or External Pipeline
The MemoryIndexer supports two backends, selected at startup (lines 42-51):
Local BM25 Fallback
The LocalBM25Backend (lines 31-38) provides zero-dependency lexical search. It requires no network, no API keys, and no external services—ideal for development or air-gapped deployments.
External Retrieval Pipeline
For production scale, the indexer connects to an HTTP service at http://localhost:4242 (configurable). The service exposes three endpoints: /index, /search, and /clear. Health verification runs in _check_retrieval_pipeline (lines 59-71).
Adding chunks flows through add_chunks (lines 75-89), which updates both local storage and the chosen backend:
- Local:
_index_documents(lines 93-103) - Remote:
_index_documents(lines 111-139)
After bulk ingestion, build_indexes (lines 144-172) synchronizes all stored chunks to the remote service.
Searching: The Core RAG Operation
The MemoryIndexer.search() method (lines 73-99) implements the retrieval logic that powers user memory and knowledge bases with RAG:
def search(self,
query: str,
top_k: int = 5,
index_mode: IndexMode = None) -> List[SearchResult]:
Two execution paths:
- Local mode →
self.local_backend.search()(lines 103-114), results wrapped inSearchResultobjects (lines 115-121) - Remote mode → POST to
/searchwithmodeparameter (dense,sparse, orhybrid), returns reranked results (lines 126-160)
The remote path includes a critical mapping step (lines 152-166): pipeline-generated doc_id values are resolved back to original chunk_id values via self.doc_id_mapping.
Each SearchResult contains:
- The
ConversationChunkobject - BM25/dense/hybrid score
- Match type annotation (
local_bm25,dense,sparse,hybrid)
Exposing Memory to the LLM: The Tool Layer
Raw search results are useless without an interface. The MemoryTools class (lines 33-38 of tools.py) wraps the indexer in OpenAI-compatible function definitions.
Three tools form the complete API:
| Tool | Purpose | Parameters |
|---|---|---|
search_memory |
Lexical/semantic search over all chunks | query (required, string) |
get_conversation_context |
Fetch N surrounding chunks for context | chunk_id (required), context_size (optional, default 2) |
get_full_conversation |
Return all chunks for a conversation/test pair | conversation_id, test_id (both required) |
The OpenAI function specifications are auto-generated by get_tool_definitions() and consumed by the agent in agentic-rag-for-user-memory/agent.py.
Saving and Loading Indexes
Offline persistence enables pre-built knowledge bases. Two JSON files capture the complete index state:
{path}_chunks.json— serializedConversationChunkobjects{path}_texts.json— pre-computed enriched texts
Save via save_index(path) (lines 94-115).
Load via load_index(path) (lines 117-172), which:
- Restores
self.chunksdictionary - Recreates
ConversationChunkobjects - Regenerates
self.chunk_textsif needed - Calls
self.build_indexes()to synchronize the backend
Complete End-to-End Example
This minimal script demonstrates the full pipeline for building user memory and knowledge bases with RAG:
# 1️⃣ Build the indexer
from indexer import MemoryIndexer, IndexConfig
from tools import MemoryTools, get_tool_definitions
indexer = MemoryIndexer(IndexConfig()) # Uses local BM25 by default
# 2️⃣ Load pre-saved conversation chunks
indexer.load_index() # Reads *_chunks.json & *_texts.json
# 3️⃣ Initialize the tool layer
tools = MemoryTools(indexer)
# 4️⃣ Search the memory base
result = tools.search_memory(query="How many flights did I book last year?")
print(result.to_dict())
# 5️⃣ Fetch surrounding context for a hit
if result.success and result.data["results"]:
first_chunk_id = result.data["results"][0]["chunk_id"]
ctx = tools.get_conversation_context(chunk_id=first_chunk_id, context_size=3)
print(ctx.to_dict())
The ToolResult.to_dict() output drops directly into LLM prompts or OpenAI function_call payloads.
Extending the Pipeline
The repository includes alternative implementations for specialized use cases:
chapter3/structured-index/graphrag_indexer.py— GraphRAG-based structured indexing for rich semantic relationshipschapter3/contextual-retrieval-for-user-memory/— Contextual indexer with evaluation tools and prompt templates
Swap retrieval_backend between "local" and "pipeline" to scale from laptop development to production clusters.
Summary
- Persistence:
MemoryManagerwrites per-user JSON files todata/memories/{user_id}_memory.json - Chunking:
ConversationChunkobjects with enriched metadata enable precise retrieval - Indexing:
MemoryIndexersupports local BM25 or remote dense/sparse/hybrid pipelines - Search:
search()method returns scoredSearchResultobjects with full provenance - Integration:
MemoryToolsexposes three OpenAI-compatible functions for agent consumption - Portability: Save/load indexes as JSON for reproducible, shareable knowledge bases
Frequently Asked Questions
What file format stores user memory in this RAG system?
User memory persists as JSON files following the pattern {user_id}_memory.json in the directory specified by Config.MEMORY_STORAGE_DIR (default: data/memories/). The MemoryManager class in chapter3/user-memory/memory_manager.py handles all read/write operations via standard Python JSON serialization.
Can this RAG pipeline run without external services?
Yes. The LocalBM25Backend (lines 31-38 of indexer.py) provides a pure-Python implementation requiring no network, API keys, or Docker containers. Set retrieval_backend="local" in IndexConfig for fully offline operation. The trade-off is lexical-only search versus the dense/hybrid modes available through the external pipeline.
How does the system prevent losing context when retrieving individual chunks?
The get_conversation_context tool fetches configurable surrounding chunks (default: 2 before and after). Given a chunk_id, it retrieves the neighboring ConversationChunk objects from self.chunks, ensuring the LLM sees conversational flow rather than isolated snippets.
What semantic modes does the external retrieval pipeline support?
The pipeline service at port 4242 accepts three index modes via the /search endpoint: dense (neural embeddings), sparse (inverted index), and hybrid (combined with reranking). These are passed through MemoryIndexer.search() as the index_mode parameter and translated to the JSON payload in lines 126-160.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →