How Nemori's Architecture Handles Long-Term Memory for LLM Workflows: Episodes, Vectors, and Search
Nemori's architecture handles long-term memory for LLM workflows by persisting conversation episodes in a relational database while asynchronously generating vector embeddings for semantic search, enabling agents to retrieve relevant historical context through hybrid vector and keyword indexing.
Nemori (nemori-ai/nemori) implements a robust persistence layer that allows LLM-driven agents to maintain context across unlimited sessions. The system combines chronological episode storage with semantic vector indexing to create a searchable, persistent knowledge base that survives server restarts and scales from local deployments to cloud infrastructure.
The Dual-Store Data Model
Nemori's memory system centers on two complementary structures managed by the MemorySystem class (src/core/memory_system.py): Episodes and Semantic Memories.
Episodes: Chronological Conversation Chunks
The Episode model (src/models/episode.py) represents discrete chunks of conversation history containing messages, timestamps, and user identifiers. These records persist in a relational database through EpisodeStorage (src/storage/episode_storage.py), providing durable, ACID-compliant storage for conversation history.
Semantic Memories: Vector Encodings for Similarity Search
The SemanticMemory model (src/models/semantic.py) stores vector-encoded representations of episodes or message subsets. Generated asynchronously via the EmbeddingClient (src/utils/embedding_client.py), these vectors enable cosine-similarity search and reside in a vector database (e.g., Chroma) managed by SemanticStorage (src/storage/semantic_storage.py).
Buffering and Episode Creation
Incoming messages first accumulate in a message buffer (src/core/message_buffer.py). When the buffer reaches the configurable BUFFER_MAX_MESSAGES threshold or a timeout expires, Nemori triggers _create_episode_from_buffer to persist the buffered content as a new Episode. Immediately following creation, the system schedules asynchronous semantic generation via _schedule_semantic_generation, which computes embeddings without blocking the main workflow.
Loading and Indexing Strategies
When initializing a user session, MemorySystem.load_user_data_and_indices executes a three-phase initialization:
- Fetch episodes from
EpisodeStoragefor the specific user. - Load semantic vectors from
SemanticStorageinto memory. - Build hybrid indices:
- Vector index (FAISS/Chroma via
src/search/chroma_search.py) for semantic similarity. - BM25 index (via
src/search/bm25_search.py) for keyword-based retrieval.
- Vector index (FAISS/Chroma via
These indices cache per-user and support forced rebuilding through rebuild_user_indices when data staleness requires a refresh.
Search and Retrieval Implementation
The public search method abstracts retrieval complexity, accepting parameters for query, top_k results, and search_method ("vector" or "bm25"). Internally, _search_episodes_by_method routes queries:
-
Vector search embeds the query using
EmbeddingClient, performs cosine-similarity lookup against the semantic index, and filters duplicates via_is_duplicate_semantic_memory. -
BM25 search tokenizes the query and executes inverted-index lookups through the BM25 implementation for exact keyword matching.
Both methods return Episode objects enriched with matched SemanticMemory snippets, allowing downstream LLM prompts to inject relevant historical context.
Lifecycle Management and Consistency
Asynchronous Processing
Semantic generation runs as background tasks managed by TaskManager (src/services/task_manager.py). Status monitoring is available through get_semantic_generation_status and wait_for_semantic_generation, ensuring agents can verify embedding completion before dependent operations.
Deletion and Cleanup
The delete_episode method removes episodes with optional cascading deletion of semantic vectors (cascade_semantic=True). For complete user data removal, delete_user_data wipes both relational and vector stores atomically.
Thread Safety
All memory operations acquire a per-user re-entrant lock (_get_user_processing_lock), guaranteeing consistency when concurrent agents manipulate the same user's memory store.
Configuration Tuning
Memory behavior is controlled via src/config.py settings including:
BUFFER_MAX_MESSAGES: Trigger threshold for episode creation.VECTOR_SEARCH_TOP_K: Default semantic result count.FORCE_REBUILD_INDICES: Flag to mandate full re-indexing on next load.
These parameters adapt the subsystem for lightweight local runs or high-throughput cloud deployments.
Code Examples
Persisting New Conversations
from nemori import MemorySystem
mem = MemorySystem()
user_id = "user-123"
# Buffer messages for automatic episode creation
mem.add_messages(
owner_id=user_id,
messages=[
{"role": "user", "content": "Explain quantum computing"},
{"role": "assistant", "content": "Quantum computing uses qubits..."}
]
)
# Force immediate persistence if needed
mem.force_episode_creation(owner_id=user_id)
Retrieving Context with Vector Search
from nemori import MemorySystem
mem = MemorySystem()
user_id = "user-123"
# Search historical episodes semantically
results = mem.search(
owner_id=user_id,
query="quantum entanglement applications",
top_k=3,
search_method="vector"
)
for episode in results:
print(f"Episode {episode.id}: {episode.semantic_memories[0].text}")
Complete User Data Removal
from nemori import MemorySystem
mem = MemorySystem()
mem.delete_user_data(owner_id="user-123")
Summary
- Nemori implements a dual-store architecture combining relational episode storage (
EpisodeStorage) with vector semantic memories (SemanticStorage) for comprehensive long-term retention. - The
MemorySystemclass orchestrates asynchronous embedding generation and hybrid indexing (vector + BM25) to support both similarity and keyword retrieval. - Per-user locking and configurable buffering ensure thread-safe, high-performance operation across concurrent LLM workflows.
- All persistence mechanisms are abstracted behind simple APIs like
add_messages,search, anddelete_user_data, enabling rapid integration into agent architectures.
Frequently Asked Questions
How does Nemori decide when to create a new episode from buffered messages?
Nemori triggers episode creation when the message buffer reaches the BUFFER_MAX_MESSAGES threshold configured in src/config.py or when a timeout expires. The _create_episode_from_buffer method then persists the buffered content and schedules asynchronous semantic vector generation.
What embedding models does Nemori use for semantic memory?
Nemori supports configurable embedding models through the EmbeddingClient utility (src/utils/embedding_client.py). The specific model is deployment-dependent and generates vectors stored in the vector database for cosine-similarity searches.
Can Nemori retrieve memories using both semantic similarity and exact keywords?
Yes. The MemorySystem.search method supports dual retrieval strategies: "vector" for semantic similarity search using FAISS/Chroma indices, and "bm25" for keyword-based retrieval using the BM25 inverted index implementation in src/search/bm25_search.py.
How does Nemori handle concurrent access to a single user's memory?
All memory operations acquire a per-user re-entrant lock via _get_user_processing_lock in MemorySystem. This mechanism guarantees thread safety when multiple concurrent agents or workflows access and modify the same user's episodic and semantic data.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →