How NemoriMemory Facade Works: A Complete Guide to Nemori's Python Memory API

NemoriMemory is a thin Python facade that wraps Nemori's MemorySystem engine, providing a single-line API for adding messages, searching episodic and semantic memory, and managing memory lifecycle while delegating all heavy processing to the underlying core.

The NemoriMemory class in the nemori-ai/nemori repository serves as the primary entry point for developers integrating Nemori's persistent memory capabilities into Python applications. This high-level facade abstracts the complexity of the MemorySystem engine located in src/core/memory_system.py, offering a clean interface for message buffering, episode generation, and hybrid vector search while maintaining full access to dependency injection hooks for custom LLM and embedding clients.

What Is the NemoriMemory Facade?

NemoriMemory functions as a structural facade that exposes a simplified, testable public surface while orchestrating the full memory pipeline. Rather than managing message buffers, vector indices, and semantic extraction directly, the facade constructs or receives a MemorySystem instance and forwards all operations to it.

The facade handles four primary responsibilities:

  • Dependency injection for custom configuration, LLM clients, and embedding providers
  • Context-manager lifecycle support for clean resource cleanup
  • Method delegation to the core engine for all memory operations
  • Convenience constructors for environment-based configuration

Core Architecture and Responsibilities

Dependency Injection and Initialization

The __init__ method in src/api/facade.py (lines 20-34) implements a flexible constructor that accepts optional pre-configured components or creates defaults from MemoryConfig:

class NemoriMemory:
    def __init__(self,
                 config: Optional[MemoryConfig] = None,
                 *,
                 memory_system: Optional[MemorySystem] = None,
                 llm_client: Optional[LLMClient] = None,
                 embedding_client: Optional[EmbeddingClient] = None):
        self.config = config or MemoryConfig()
        self._memory_system = memory_system or MemorySystem(
            config=self.config,
            language=self.config.language,
            llm_client=llm_client,
            embedding_client=embedding_client,
        )

This design enables unit testing by allowing injection of mock MemorySystem instances while production code can rely on the default instantiation. The MemorySystem class itself, defined in src/core/memory_system.py (lines 35-63), orchestrates the heavy processing including parallel indexing and semantic generation.

Context Manager Lifecycle

The facade implements standard Python context manager protocols in src/api/facade.py (lines 39-48) to ensure proper resource cleanup:

def close(self) -> None:
    self._memory_system.__exit__(None, None, None)

def __enter__(self) -> "NemoriMemory":
    return self

def __exit__(self, exc_type, exc_val, exc_tb) -> None:
    self.close()

When exiting a with block or calling close(), the facade triggers MemorySystem.__exit__, which gracefully shuts down thread pools, clears internal caches, and releases external resources such as database connections or HTTP clients.

Key Operations and Method Delegation

Adding Messages and Buffering

The add_messages method in src/api/facade.py (lines 52-84) forwards raw message dictionaries to the core engine:

def add_messages(self, user_id: str, messages: Sequence[Dict[str, Any]]) -> Dict[str, Any]:
    return self._memory_system.add_messages(user_id, list(messages))

Inside MemorySystem.add_messages (src/core/memory_system.py, lines 52-23), messages are transformed into Message objects, buffered per user, and automatically converted into episodes when batch segmentation triggers (based on size, time, or custom logic) are reached. The facade remains unaware of these implementation details.

Forcing Episode Creation with flush()

To immediately create an episode from buffered messages without waiting for automatic triggers, the facade exposes flush():

def flush(self, user_id: str) -> Optional[Dict[str, Any]]:
    return self._memory_system.force_episode_creation(user_id)

This maps directly to MemorySystem.force_episode_creation, useful for explicit checkpointing before critical operations or application shutdown.

Semantic Memory Generation

Semantic extraction runs asynchronously in a background thread pool. The facade provides a blocking wait mechanism:

def wait_for_semantic(self, user_id: str, timeout: float = 30.0) -> bool:
    return self._memory_system.wait_for_semantic_generation(user_id, timeout=timeout)

This ensures that all pending semantic memory tasks for a specific user complete before proceeding, with configurable timeout protection.

The search method delegates to MemorySystem.search_all for parallel hybrid retrieval across episodic (BM25) and semantic (vector) indices:

def search(self,
           user_id: str,
           query: str,
           *,
           top_k_episodes: Optional[int] = None,
           top_k_semantic: Optional[int] = None,
           search_method: str = "hybrid") -> Dict[str, List[Dict[str, Any]]]:
    return self._memory_system.search_all(
        user_id=user_id,
        query=query,
        top_k_episodes=top_k_episodes,
        top_k_semantic=top_k_semantic,
        search_method=search_method,
    )

The facade aggregates results into a dictionary containing episodic and semantic lists, abstracting the complex parallel cache lookups and index optimizations performed by the core engine.

Memory Management and Statistics

Diagnostic and CRUD operations follow the same delegation pattern:

def stats(self, user_id: Optional[str] = None) -> Dict[str, Any]:
    return self._memory_system.get_stats(user_id)

def delete_episode(self, user_id: str, episode_id: str, *, cascade_semantic: bool = True) -> Dict[str, Any]:
    return self._memory_system.delete_episode(user_id, episode_id, cascade_semantic)

def delete_semantic_memory(self, user_id: str, memory_id: str) -> Dict[str, Any]:
    return self._memory_system.delete_semantic_memory(user_id, memory_id)

These methods provide unified access to storage statistics and selective deletion capabilities without exposing underlying storage implementation details.

Practical Code Examples

Synchronous Usage Pattern

from nemori.api.facade import NemoriMemory

# One-line creation using environment variables

mem = NemoriMemory.from_env()

# Add a batch of chat messages

mem.add_messages(
    user_id="alice",
    messages=[
        {"role": "user", "content": "Hello!"},
        {"role": "assistant", "content": "Hi Alice, how can I help?"}
    ],
)

# Force immediate episode creation from buffer

mem.flush("alice")

# Wait for background semantic extraction to complete

mem.wait_for_semantic("alice")

# Perform hybrid search across episodic and semantic memory

results = mem.search(
    user_id="alice",
    query="What did we talk about yesterday?",
    top_k_episodes=5,
    top_k_semantic=5,
)

print(results["episodic"])   # BM25-ranked episode list

print(results["semantic"])   # Vector similarity results

Async Usage with asearch()

import asyncio
from nemori.api.facade import NemoriMemory

async def demo():
    # Context manager works with async patterns

    async with NemoriMemory.from_env() as mem:
        await mem.add_messages("bob", [
            {"role": "user", "content": "Tell me a joke."},
        ])
        
        # Non-blocking search using executor

        results = await mem.asearch(
            user_id="bob",
            query="joke",
            top_k_episodes=3,
            top_k_semantic=3,
        )
        print(results)

asyncio.run(demo())

Resource Cleanup and Diagnostics

mem = NemoriMemory()

# Retrieve global or per-user statistics

print(mem.stats())               # Global counters

print(mem.stats("alice"))        # User-specific metrics

# Delete specific episodes with optional cascade

mem.delete_episode("alice", "ep_123", cascade_semantic=True)

# Explicit cleanup if not using context manager

mem.close()

Key Source Files

File Purpose
src/api/facade.py Public NemoriMemory facade class – wiring, context manager, thin method wrappers.
src/core/memory_system.py Core engine handling buffering, episode/semantic generation, indexing, caching, and metrics.
src/config.py MemoryConfig reads environment variables and provides defaults for the system.
src/utils/llm_client.py Wrapper around LLM APIs used by MemorySystem.
src/utils/embedding_client.py Wrapper for embedding models – supplies vectors for semantic memory.

Summary

  • NemoriMemory acts as a thin facade over the MemorySystem engine in src/core/memory_system.py, exposing a simplified Python API while hiding complex buffering, indexing, and semantic extraction logic.
  • The facade supports dependency injection for custom LLM clients, embedding providers, and configuration, making it fully testable and adaptable to different deployment environments.
  • All heavy operations—including parallel hybrid search across BM25 and vector indices, asynchronous semantic memory generation, and episode batching—are delegated to MemorySystem, with NemoriMemory merely forwarding calls.
  • The class implements context manager protocols (__enter__/__exit__) for automatic resource cleanup, and provides both synchronous (search) and asynchronous (asearch) interfaces for flexible integration.

Frequently Asked Questions

What is the difference between NemoriMemory and MemorySystem?

NemoriMemory is the high-level facade located in src/api/facade.py that provides a simplified, user-facing API for memory operations. MemorySystem, defined in src/core/memory_system.py, is the internal engine that performs all heavy processing including message buffering, episode generation, semantic extraction, and hybrid search indexing. The facade delegates all method calls to the engine while handling initialization and lifecycle management.

How does NemoriMemory handle async operations?

While the underlying MemorySystem operates synchronously, NemoriMemory provides an asearch() method that wraps the synchronous search logic in asyncio.get_running_loop().run_in_executor(). This allows non-blocking search operations in async applications. For other operations like add_messages, the facade provides standard synchronous methods that return immediately while background semantic processing occurs in MemorySystem's internal thread pools.

Can I inject custom LLM or embedding clients into NemoriMemory?

Yes. The __init__ method in src/api/facade.py accepts optional llm_client and embedding_client parameters, allowing you to supply custom implementations of the LLMClient and EmbeddingClient interfaces. This dependency injection pattern enables unit testing with mock clients and integration with proprietary or self-hosted models rather than the default providers configured through MemoryConfig.

Where does NemoriMemory store the actual memory data?

NemoriMemory itself does not handle storage directly; it delegates persistence to the injected MemorySystem. According to the source architecture in src/core/memory_system.py, the engine manages episodic storage (typically BM25-indexed text) and semantic storage (vector embeddings) through internal indexing systems. The specific storage backend—whether local disk, SQLite, or external vector databases—is determined by the MemoryConfig settings passed during initialization or loaded via from_env().

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →