Core Components of the Nemori Memory System: Architecture Guide

The Nemori memory system core components comprise a centralized MemoryConfig, a MemorySystem orchestrator, per-user MessageBufferManager locking, dual-stage EpisodeGenerator and SemanticGenerator processors, repository-backed persistence, and hybrid BM25/Chroma search indices.

The nemori-ai/nemori repository delivers a modular, high-performance pipeline for transforming conversational messages into searchable episodic and semantic memories. Understanding these Nemori memory system core components enables developers to customize storage backends, tune retrieval parameters, and extend the generation workflow without modifying internal logic.

Configuration Layer: Centralized System Control

All tunable parameters—model names, buffer limits, caching strategies, and storage backends—are centralized in MemoryConfig within [src/config.py](https://github.com/nemori-ai/nemori/blob/main/src/config.py). This singleton configuration object propagates settings throughout the stack, allowing runtime swaps between in-memory and filesystem storage, or between Chroma and lexical indices, without code changes.

The MemorySystem Orchestrator

The MemorySystem class in [src/core/memory_system.py](https://github.com/nemori-ai/nemori/blob/main/src/core/memory_system.py) serves as the primary entry point and workflow conductor. It coordinates message buffering, episode generation, semantic extraction, indexing, and search operations while maintaining thread safety for concurrent multi-user access. When instantiated, it lazily initializes providers, repositories, and indices according to the global MemoryConfig.

Data Ingestion Pipeline

Per-User Message Buffering

Incoming messages first land in the MessageBufferManager defined in [src/core/message_buffer.py](https://github.com/nemori-ai/nemori/blob/main/src/core/message_buffer.py). This component maintains a rolling window of messages per user ID, guarded by threading.RLock to ensure parallel requests for different users never block each other while preventing race conditions within a single user's buffer.

Episodic Memory Generation

When the buffer reaches a configured threshold, MemorySystem invokes the EpisodeGenerator from [src/generation/episode_generator.py](https://github.com/nemori-ai/nemori/blob/main/src/generation/episode_generator.py). This module converts message batches into Episode objects containing titles, condensed content, and timestamps, forming the episodic memory layer.

Semantic Memory Extraction

The pipeline then triggers SemanticGenerator in [src/generation/semantic_generator.py](https://github.com/nemori-ai/nemori/blob/main/src/generation/semantic_generator.py) (optionally leveraging PredictionCorrectionEngine) to distill higher-level semantic memories—such as user knowledge, profiles, and experiences—from one or many episodes. These SemanticMemory objects capture abstracted facts independent of specific conversational turns.

Storage and Persistence Layer

Repository Pattern

Concrete storage implementations reside in [src/infrastructure/repositories.py](https://github.com/nemori-ai/nemori/blob/main/src/infrastructure/repositories.py). The EpisodeStorageRepository and SemanticStorageRepository abstract whether data persists on disk or in memory, providing a clean interface for the core engine while hiding backend-specific details.

Search and Retrieval Infrastructure

Dual-Mode Indexing

Fast retrieval relies on two index types managed in [src/infrastructure/indices.py](https://github.com/nemori-ai/nemori/blob/main/src/infrastructure/indices.py):

  • Bm25Index: Provides lexical search capabilities using the BM25 algorithm.
  • ChromaVectorIndex: Handles dense vector similarity search via ChromaDB.

Both implement abstract LexicalIndex and VectorIndex interfaces, permitting in-memory fallbacks for testing environments.

UnifiedSearchEngine

The UnifiedSearchEngine in [src/search/unified_search.py](https://github.com/nemori-ai/nemori/blob/main/src/search/unified_search.py) orchestrates hybrid queries, combining BM25 lexical scores with vector similarity to return ranked results across episodic and semantic stores. Developers can also invoke BM25Search or ChromaSearchEngine directly for pure lexical or vector retrieval.

Auxiliary Services

Supporting infrastructure in [src/services/providers.py](https://github.com/nemori-ai/nemori/blob/main/src/services/providers.py) includes:

  • DefaultProviders: Lazy factories that wire default implementations together based on MemoryConfig.
  • EventBus: Decoupled publish/subscribe mechanism triggering semantic generation when episode_created events fire.
  • MetricsReporter: Records operation timings for performance monitoring.
  • PerUserCache and SemanticEmbeddingCache: In-memory caches for episodes, semantic memories, and LLM embeddings to reduce redundant computation.

Immutable Data Models

Cross-layer communication relies on immutable structures defined in [src/models/message.py](https://github.com/nemori-ai/nemori/blob/main/src/models/message.py), [src/models/episode.py](https://github.com/nemori-ai/nemori/blob/main/src/models/episode.py), and [src/models/semantic.py](https://github.com/nemori-ai/nemori/blob/main/src/models/semantic.py). The Message, Episode, and SemanticMemory classes enforce data integrity as messages flow from raw input through generation to final storage.

Implementation Examples

Initialize the memory system with default configuration:

from nemori.core.memory_system import MemorySystem

mem = MemorySystem()  # loads OpenAI key from env, sets up providers

Submit a batch of conversational messages for processing:

messages = [
    {"role": "user", "content": "I love hiking in the Alps."},
    {"role": "assistant", "content": "That sounds amazing!"},
    {"role": "user", "content": "Last summer I climbed Mont Blanc."}
]

result = mem.add_messages(owner_id="user123", messages=messages)
print(result["episodes_created"])  # shows generated episodes (titles, ids, etc.)

Execute hybrid search across both memory types:

search_results = mem.search_all(
    user_id="user123",
    query="What outdoor activities do I enjoy?",
    top_k_episodes=5,
    top_k_semantic=5,
    search_method="hybrid"
)

print("Episodes:", search_results["episodic"])
print("Semantic:", search_results["semantic"])

Access low-level indices directly for custom pipelines:

vector_idx = mem._vector_index    # ChromaVectorIndex by default

lexical_idx = mem._lexical_index  # Bm25Index by default

vector_hits = vector_idx.search_episodes("user123", "hiking", top_k=3)
lexical_hits = lexical_idx.search_episodes("user123", "hiking", top_k=3)

Summary

  • Configuration: MemoryConfig in src/config.py centralizes all tunable parameters.
  • Orchestration: MemorySystem in src/core/memory_system.py coordinates the entire workflow.
  • Buffering: MessageBufferManager provides thread-safe, per-user message queues.
  • Generation: EpisodeGenerator and SemanticGenerator create episodic and semantic memories respectively.
  • Storage: Repository pattern in src/infrastructure/repositories.py abstracts persistence.
  • Search: Dual indices (Bm25Index, ChromaVectorIndex) and UnifiedSearchEngine enable hybrid retrieval.
  • Services: EventBus, MetricsReporter, and caching layers optimize performance and decouple components.

Frequently Asked Questions

What distinguishes episodic from semantic memories in Nemori?

Episodic memories are time-bound records of specific conversations generated by EpisodeGenerator, containing titles and timestamps. Semantic memories are abstracted knowledge entities extracted by SemanticGenerator that capture facts, user preferences, and learned concepts independent of the original dialogue sequence.

How does Nemori ensure thread safety during message ingestion?

The MessageBufferManager utilizes threading.RLock to guard per-user buffers, as implemented in src/core/message_buffer.py. This design allows concurrent access across different users while serializing access for any single user ID, preventing race conditions during message addition or buffer flushing.

Can storage backends be swapped without modifying source code?

Yes. The repository pattern and provider system allow backend substitution through MemoryConfig alone. Developers can switch between in-memory and filesystem storage, or between Chroma and in-memory vector indices, by adjusting configuration parameters in src/config.py without touching the core logic in src/core/memory_system.py.

What search methods does the UnifiedSearchEngine support?

The UnifiedSearchEngine supports three query modes: hybrid (combining BM25 lexical and vector similarity scores), pure-vector (dense embedding search via Chroma), and pure-lexical (BM25 keyword search). These modes are selectable via the search_method parameter in MemorySystem.search_all().

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →