What Is the Role of the MemorySystem in Nemori? A Deep Dive into the Core Orchestrator
The MemorySystem serves as the central orchestration layer that manages Nemori’s long-term memory by buffering messages, generating episodes, extracting semantic knowledge, and maintaining vector and lexical indices for fast retrieval.
The MemorySystem class in the nemori-ai/nemori repository acts as the primary entry point for all memory-related operations. It coordinates between message buffers, storage backends, search indices, and language models to transform raw conversational data into structured, searchable memories while enforcing multi-tenant isolation through owner_id scoping.
Core Responsibilities of the MemorySystem in Nemori
Message Buffering and Episode Generation
The MemorySystem manages transient conversation data through the MessageBufferManager. When you call add_messages() in src/core/memory_system.py (line 652), the system accumulates incoming messages in a per-user buffer and applies internal heuristics to detect conversation boundaries.
Once thresholds are met, the system invokes _create_episode_from_buffer() to instantiate Episode objects via the EpisodeGenerator interface (line 629). This conversion transforms raw message lists into persistent, timestamped conversation chunks stored in src/storage/episode_storage.py.
Semantic Memory Extraction
Beyond raw episode storage, the MemorySystem elevates conversations into high-level knowledge through semantic memory extraction. The _schedule_semantic_generation() method (line 1446) queues background tasks that run LLM-driven analysis on new episodes to extract SemanticMemory objects—condensed facts, preferences, and insights independent of the original conversation flow.
This process runs asynchronously via the TaskManager (src/services/task_manager.py) using a dedicated thread pool controlled by semantic_generation_workers in MemoryConfig.
Indexing and Search Orchestration
The MemorySystem abstracts Nemori’s dual-index architecture through unified search methods. It maintains:
- Vector indices (Chroma) for semantic similarity search
- Lexical indices (BM25) for keyword-based retrieval
The search() method (line 1024) in src/core/memory_system.py serves as the primary query interface, delegating to _search_episodes_by_method() or _search_semantic_by_method() based on the search_method parameter ("vector" or "bm25"). The system automatically keeps indices synchronized with new episodes through background reindexing tasks.
Performance Optimization and Caching
To minimize redundant LLM and embedding calls, the MemorySystem implements sophisticated caching via PerUserCache and SemanticEmbeddingCache (src/services/cache.py). Vector embeddings and search results are cached per owner_id, ensuring multi-tenant isolation while maximizing throughput.
Parallelism is controlled through ThreadPoolExecutor instances configured via MemoryConfig (src/config.py), with separate worker pools for semantic generation, indexing, and search operations. The PerformanceOptimizer class monitors latency and success rates through MetricsReporter, emitting lifecycle events on the global EventBus.
How the MemorySystem Works: Code Examples
# 1️⃣ Initialise the memory system (uses defaults or a custom config)
from nemori.src.core.memory_system import MemorySystem
mem = MemorySystem() # reads OPENAI_API_KEY from env, loads defaults
# 2️⃣ Add a series of messages for a particular user
messages = [
{"role": "user", "content": "What's the weather today?"},
{"role": "assistant", "content": "It's sunny in San Francisco."},
{"role": "user", "content": "Remind me to buy sunscreen."},
]
mem.add_messages(owner_id="alice@example.com", messages=messages)
# 3️⃣ Force episode creation (useful for testing)
mem.force_episode_creation(owner_id="alice@example.com")
# 4️⃣ Search across both episodes and semantic memories
results = mem.search(
owner_id="alice@example.com",
query="weather forecast",
top_k=5,
search_method="vector", # "vector" or "bm25"
)
print(results) # → dict with `episodes` and `semantic` hits
# 5️⃣ Get system statistics (cache size, number of episodes, etc.)
stats = mem.get_stats(owner_id="alice@example.com")
print(stats)
# 6️⃣ Clean up a user’s data when they leave the app
mem.delete_user_data(owner_id="alice@example.com")
Key implementation details:
- The
MemorySystemis instantiated once per application lifecycle (or per request in serverless environments). add_messages()triggers automatic buffering; thresholds are defined inMemoryConfig.search()abstracts the underlying Chroma and BM25 implementations insrc/search/chroma_search.pyandsrc/search/bm25_search.py.- All operations require an
owner_idparameter, ensuring strict data isolation between users.
Key Files and Architecture
| File | Role |
|---|---|
src/core/memory_system.py |
Core orchestrator class – the entry point for all memory-related actions. |
src/config.py |
MemoryConfig dataclass that drives the system’s defaults and limits. |
src/core/message_buffer.py |
Implements MessageBufferManager, handling message accumulation and boundary logic. |
src/storage/episode_storage.py |
Persistence layer for Episode objects (filesystem or in-memory). |
src/storage/semantic_storage.py |
Persistence for SemanticMemory records. |
src/search/chroma_search.py & src/search/bm25_search.py |
Vector (Chroma) and lexical (BM25) search implementations used by the Memory System. |
src/services/providers.py |
Supplies concrete implementations for LLM, embedding, repositories, and indices. |
src/services/task_manager.py |
Manages background tasks (semantic generation, index rebuilding). |
src/models/episode.py & src/models/semantic.py |
Data models representing episodes and extracted semantic memories. |
src/services/cache.py |
Implements per-user caching for embeddings and search results. |
These files constitute the full “memory stack” in Nemori, with MemorySystem acting as the glue that ties them all together.
Summary
- The MemorySystem is the central orchestrator for Nemori’s long-term memory, implemented in
src/core/memory_system.py. - It buffers messages via
MessageBufferManagerand converts them intoEpisodeobjects using configurable heuristics. - It extracts semantic knowledge by running background LLM tasks that generate
SemanticMemoryrecords from episodes. - It maintains dual search indices (Chroma for vectors, BM25 for lexical) and provides a unified
search()interface. - It enforces multi-tenant isolation through
owner_idscoping and optimizes performance via per-user caching and thread-pool parallelism.
Frequently Asked Questions
What is the primary function of the MemorySystem in Nemori?
The primary function is to act as a high-performance orchestration layer that transforms raw conversational messages into structured, searchable long-term memory. It coordinates message buffering, episode generation, semantic extraction, and dual-index search (vector and lexical) while maintaining strict user isolation through owner_id scoping.
How does the MemorySystem handle multi-tenant data isolation?
Every method in the MemorySystem requires an owner_id parameter (typically a user email or UUID) that scopes all data operations to a specific user. The system uses PerUserCache instances and separate storage paths per owner in episode_storage.py and semantic_storage.py, ensuring that queries, embeddings, and generated episodes never leak between users.
What indexing methods does the MemorySystem support?
The MemorySystem supports two primary search methods: vector similarity search using ChromaDB for semantic retrieval, and lexical search using BM25 for keyword-based matching. The search() method accepts a search_method parameter ("vector" or "bm25") and delegates to _search_episodes_by_method() or _search_semantic_by_method() to query the appropriate index.
Where is the MemorySystem class defined in the Nemori codebase?
The MemorySystem class is defined in src/core/memory_system.py. This file contains the main orchestration logic, including methods for message buffering (add_messages), episode generation (_create_episode_from_buffer), semantic extraction (_schedule_semantic_generation), and unified search (search). Configuration defaults are imported from src/config.py via the MemoryConfig dataclass.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →