How to Initialize and Use the Nemori MemorySystem: Complete Guide

Initialize the Nemori MemorySystem by instantiating NemoriMemory with a MemoryConfig, then use add_messages(), flush(), and search() to manage conversational memory.

The Nemori MemorySystem serves as the central orchestrator for the nemori-ai/nemori library, combining message buffering, episode generation, and hybrid search capabilities into a thread-safe interface. This guide covers the initialization patterns, core operations, and extension points available in the MemorySystem class.

What Is the Nemori MemorySystem?

The MemorySystem class in src/core/memory_system.py acts as a façade that orchestrates all sub-services required for persistent conversational memory. It integrates:

  • Message Buffer: Temporary storage for raw conversation turns before episode generation
  • Episode Generation: Logic that converts message buffers into structured episodic memories
  • Semantic Extraction: Background processes that distill semantic memories from episodes
  • Hybrid Search: Combined vector (ChromaDB) and lexical (BM25) search indices
  • Caching Layer: Per-user caches and performance optimizers to reduce redundant LLM calls

The system handles concurrent user-level locks and maintains statistics for monitoring, making it suitable for multi-user LLM applications.

Prerequisites and Configuration

Setting Up MemoryConfig

Before initializing the system, configure the MemoryConfig class from src/config.py. This dataclass holds all tunable defaults, including storage paths, model names, buffer limits, and cache settings.

from nemori.config import MemoryConfig

cfg = MemoryConfig(
    buffer_size_max=50,              # Messages before auto-flush

    storage_path="./nemori_data",    # Local filesystem storage

    openai_api_key="sk-...",         # Or set OPENAI_API_KEY env var

    embedding_model="text-embedding-3-small",
    llm_model="gpt-4o-mini"
)

The configuration validates required environment variables automatically. If openai_api_key is not provided, the system attempts to read OPENAI_API_KEY from the environment.

Initializing the Nemori MemorySystem

Quick Start with the NemoriMemory Façade

For most applications, use the NemoriMemory façade from src/api/facade.py. This public API builds a MemorySystem internally and forwards calls to it, providing a clean one-liner initialization.

from nemori import NemoriMemory

# Assumes OPENAI_API_KEY is set in environment

with NemoriMemory() as mem:
    mem.add_messages("alice", [
        {"role": "user", "content": "What's the weather in London?"},
        {"role": "assistant", "content": "It's rainy today."},
    ])
    mem.flush("alice")
    results = mem.search("alice", "weather")

The NemoriMemory class acts as a context manager, ensuring proper cleanup of resources and caches when exiting the with block.

Direct MemorySystem Initialization

For advanced use cases requiring fine-grained control, instantiate MemorySystem directly from src/core/memory_system.py. This approach allows you to inject custom clients and repositories.

from nemori.core.memory_system import MemorySystem
from nemori.config import MemoryConfig
from nemori.utils import LLMClient, EmbeddingClient

cfg = MemoryConfig(storage_path="./custom_storage")
llm = LLMClient(api_key="sk-...", model="gpt-4o-mini")
embed = EmbeddingClient(api_key="sk-...", model="text-embedding-3-small")

core = MemorySystem(
    config=cfg,
    llm_client=llm,
    embedding_client=embed
)

core.add_messages("bob", [{"role": "user", "content": "Tell me a joke."}])
core.flush("bob")

The MemorySystem.__init__ method lazily creates required providers when any component is missing, then wires them together via self.llm_client, self.embedding_client, and self._episode_repository.

Dependency Injection for Testing

The architecture supports complete dependency injection, enabling unit tests with mock repositories. Pass custom repository objects to MemorySystem to bypass filesystem or database requirements.

from nemori.core.memory_system import MemorySystem
from nemori.config import MemoryConfig

class InMemoryEpisodeRepo:
    def __init__(self): self.store = {}
    def list_by_user(self, uid): return self.store.get(uid, [])
    def save(self, ep):
        self.store.setdefault(ep.owner_id, []).append(ep)
        return ep.episode_id

cfg = MemoryConfig()
mem = MemorySystem(
    config=cfg,
    episode_repository=InMemoryEpisodeRepo(),
    semantic_repository=InMemoryEpisodeRepo()
)

mem.add_messages("test", [{"role":"user","content":"Hello"}])
mem.flush("test")
assert len(mem.search("test", "Hello")) > 0

When injecting repositories, the MemorySystem will still auto-create LLMClient and EmbeddingClient unless you also provide mocks for those parameters.

Core Operations and Usage Patterns

Adding Messages and Managing Buffers

Use add_messages(user_id, messages) to buffer conversation turns. The system accumulates messages until a boundary condition (buffer size or batch threshold) triggers automatic episode generation.

mem.add_messages("user_123", [
    {"role": "system", "content": "You are a helpful assistant."},
    {"role": "user", "content": "Explain quantum computing."}
])

Creating Episodes with Flush

Call flush(user_id) to immediately convert the current message buffer into a structured episode. This is essential at conversation boundaries or when you need persisted context before searching.

mem.flush("user_123")

Searching Memory

The search(user_id, query, top_k_episodes=5) method performs hybrid retrieval across episodic and semantic memories. By default, it uses the configured search method (hybrid), combining vector similarity from src/search/chroma_search.py and lexical matching from src/search/bm25_search.py.

results = mem.search("user_123", "quantum computing concepts")

Monitoring Statistics

Use stats(user_id) to retrieve counters for processed messages, generated episodes, search queries, and cache hits. This aids in debugging and performance tuning.

print(mem.stats("user_123"))

Deleting Memories

Remove specific episodes or semantic memories using delete_episode(user_id, episode_id) or delete_semantic_memory(user_id, mem_id). The system automatically updates indices and invalidates relevant caches.

mem.delete_episode("user_123", "ep_456")

Summary

  • Initialize the Nemori MemorySystem using NemoriMemory() for quick starts or MemorySystem() for advanced control with dependency injection.
  • Configure behavior via MemoryConfig in src/config.py, setting buffer sizes, storage paths, and model names.
  • Buffer messages with add_messages(), then persist episodes using flush() or automatic thresholds.
  • Search across episodic and semantic memories via search(), which leverages ChromaDB vector indices and BM25 lexical indices in src/search/.
  • Extend the system by injecting custom repositories, LLM clients, or embedding clients through the MemorySystem constructor in src/core/memory_system.py.

Frequently Asked Questions

Do I need OpenAI API keys to use Nemori MemorySystem?

Yes, by default the system requires an OpenAI API key to power the LLMClient and EmbeddingClient used for semantic memory extraction and vector embeddings. You can provide the key via the openai_api_key parameter in MemoryConfig or set the OPENAI_API_KEY environment variable. For offline or test environments, you can inject mock clients that implement the same interface without requiring API access.

Can I use Nemori MemorySystem without persistent storage?

Yes, you can configure the system to use in-memory repositories instead of filesystem storage. Instantiate MemorySystem with custom repository objects that implement the storage interface, such as dictionaries or lists held in memory. This approach is useful for unit testing or ephemeral conversational contexts where persistence across restarts is not required. The DefaultProviders factory in src/services/providers.py will use filesystem storage only if you do not override the repository parameters.

How do I switch from ChromaDB to another vector database?

To replace ChromaDB with an alternative vector store, implement a custom class that matches the interface used by VectorIndex in src/search/chroma_search.py, then pass your implementation to the MemorySystem constructor via the appropriate parameter (typically vector_index or through a custom provider factory). The system uses dependency injection throughout, so as long as your replacement implements the search and add methods expected by the orchestrator, the Nemori MemorySystem will use it for all vector-based retrieval operations.

What is the difference between NemoriMemory and MemorySystem?

NemoriMemory is the high-level façade located in src/api/facade.py that provides a simplified, ergonomic API for most users. It internally constructs and manages a MemorySystem instance and forwards method calls to it. MemorySystem in src/core/memory_system.py is the low-level orchestrator that directly manages all sub-services, repositories, and indices. Use NemoriMemory for standard applications where you want automatic resource management via context managers, and use MemorySystem directly when you need fine-grained control over dependency injection or want to customize internal components like LLM clients or storage backends.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →