# How NemoriMemory Facade Works: A Complete Guide to Nemori's Python Memory API

> Discover how NemoriMemory's Python facade simplifies adding messages, searching memory, and managing lifecycle by delegating heavy processing to the core engine. Learn more.

- Repository: [Nemori AI/nemori](https://github.com/nemori-ai/nemori)
- Tags: how-to-guide
- Published: 2026-03-08

---

**NemoriMemory is a thin Python facade that wraps Nemori's MemorySystem engine, providing a single-line API for adding messages, searching episodic and semantic memory, and managing memory lifecycle while delegating all heavy processing to the underlying core.**

The `NemoriMemory` class in the `nemori-ai/nemori` repository serves as the primary entry point for developers integrating Nemori's persistent memory capabilities into Python applications. This high-level facade abstracts the complexity of the `MemorySystem` engine located in [`src/core/memory_system.py`](https://github.com/nemori-ai/nemori/blob/main/src/core/memory_system.py), offering a clean interface for message buffering, episode generation, and hybrid vector search while maintaining full access to dependency injection hooks for custom LLM and embedding clients.

## What Is the NemoriMemory Facade?

`NemoriMemory` functions as a **structural facade** that exposes a simplified, testable public surface while orchestrating the full memory pipeline. Rather than managing message buffers, vector indices, and semantic extraction directly, the facade constructs or receives a `MemorySystem` instance and forwards all operations to it.

The facade handles four primary responsibilities:
- **Dependency injection** for custom configuration, LLM clients, and embedding providers
- **Context-manager lifecycle** support for clean resource cleanup
- **Method delegation** to the core engine for all memory operations
- **Convenience constructors** for environment-based configuration

## Core Architecture and Responsibilities

### Dependency Injection and Initialization

The `__init__` method in [`src/api/facade.py`](https://github.com/nemori-ai/nemori/blob/main/src/api/facade.py) (lines 20-34) implements a flexible constructor that accepts optional pre-configured components or creates defaults from `MemoryConfig`:

```python
class NemoriMemory:
    def __init__(self,
                 config: Optional[MemoryConfig] = None,
                 *,
                 memory_system: Optional[MemorySystem] = None,
                 llm_client: Optional[LLMClient] = None,
                 embedding_client: Optional[EmbeddingClient] = None):
        self.config = config or MemoryConfig()
        self._memory_system = memory_system or MemorySystem(
            config=self.config,
            language=self.config.language,
            llm_client=llm_client,
            embedding_client=embedding_client,
        )

```

This design enables **unit testing** by allowing injection of mock `MemorySystem` instances while production code can rely on the default instantiation. The `MemorySystem` class itself, defined in [`src/core/memory_system.py`](https://github.com/nemori-ai/nemori/blob/main/src/core/memory_system.py) (lines 35-63), orchestrates the heavy processing including parallel indexing and semantic generation.

### Context Manager Lifecycle

The facade implements standard Python context manager protocols in [`src/api/facade.py`](https://github.com/nemori-ai/nemori/blob/main/src/api/facade.py) (lines 39-48) to ensure proper resource cleanup:

```python
def close(self) -> None:
    self._memory_system.__exit__(None, None, None)

def __enter__(self) -> "NemoriMemory":
    return self

def __exit__(self, exc_type, exc_val, exc_tb) -> None:
    self.close()

```

When exiting a `with` block or calling `close()`, the facade triggers `MemorySystem.__exit__`, which gracefully shuts down thread pools, clears internal caches, and releases external resources such as database connections or HTTP clients.

## Key Operations and Method Delegation

### Adding Messages and Buffering

The `add_messages` method in [`src/api/facade.py`](https://github.com/nemori-ai/nemori/blob/main/src/api/facade.py) (lines 52-84) forwards raw message dictionaries to the core engine:

```python
def add_messages(self, user_id: str, messages: Sequence[Dict[str, Any]]) -> Dict[str, Any]:
    return self._memory_system.add_messages(user_id, list(messages))

```

Inside `MemorySystem.add_messages` ([`src/core/memory_system.py`](https://github.com/nemori-ai/nemori/blob/main/src/core/memory_system.py), lines 52-23), messages are transformed into `Message` objects, buffered per user, and automatically converted into **episodes** when batch segmentation triggers (based on size, time, or custom logic) are reached. The facade remains unaware of these implementation details.

### Forcing Episode Creation with flush()

To immediately create an episode from buffered messages without waiting for automatic triggers, the facade exposes `flush()`:

```python
def flush(self, user_id: str) -> Optional[Dict[str, Any]]:
    return self._memory_system.force_episode_creation(user_id)

```

This maps directly to `MemorySystem.force_episode_creation`, useful for explicit checkpointing before critical operations or application shutdown.

### Semantic Memory Generation

Semantic extraction runs asynchronously in a background thread pool. The facade provides a blocking wait mechanism:

```python
def wait_for_semantic(self, user_id: str, timeout: float = 30.0) -> bool:
    return self._memory_system.wait_for_semantic_generation(user_id, timeout=timeout)

```

This ensures that all pending semantic memory tasks for a specific user complete before proceeding, with configurable timeout protection.

### Hybrid Memory Search

The `search` method delegates to `MemorySystem.search_all` for parallel hybrid retrieval across episodic (BM25) and semantic (vector) indices:

```python
def search(self,
           user_id: str,
           query: str,
           *,
           top_k_episodes: Optional[int] = None,
           top_k_semantic: Optional[int] = None,
           search_method: str = "hybrid") -> Dict[str, List[Dict[str, Any]]]:
    return self._memory_system.search_all(
        user_id=user_id,
        query=query,
        top_k_episodes=top_k_episodes,
        top_k_semantic=top_k_semantic,
        search_method=search_method,
    )

```

The facade aggregates results into a dictionary containing `episodic` and `semantic` lists, abstracting the complex parallel cache lookups and index optimizations performed by the core engine.

### Memory Management and Statistics

Diagnostic and CRUD operations follow the same delegation pattern:

```python
def stats(self, user_id: Optional[str] = None) -> Dict[str, Any]:
    return self._memory_system.get_stats(user_id)

def delete_episode(self, user_id: str, episode_id: str, *, cascade_semantic: bool = True) -> Dict[str, Any]:
    return self._memory_system.delete_episode(user_id, episode_id, cascade_semantic)

def delete_semantic_memory(self, user_id: str, memory_id: str) -> Dict[str, Any]:
    return self._memory_system.delete_semantic_memory(user_id, memory_id)

```

These methods provide unified access to storage statistics and selective deletion capabilities without exposing underlying storage implementation details.

## Practical Code Examples

### Synchronous Usage Pattern

```python
from nemori.api.facade import NemoriMemory

# One-line creation using environment variables

mem = NemoriMemory.from_env()

# Add a batch of chat messages

mem.add_messages(
    user_id="alice",
    messages=[
        {"role": "user", "content": "Hello!"},
        {"role": "assistant", "content": "Hi Alice, how can I help?"}
    ],
)

# Force immediate episode creation from buffer

mem.flush("alice")

# Wait for background semantic extraction to complete

mem.wait_for_semantic("alice")

# Perform hybrid search across episodic and semantic memory

results = mem.search(
    user_id="alice",
    query="What did we talk about yesterday?",
    top_k_episodes=5,
    top_k_semantic=5,
)

print(results["episodic"])   # BM25-ranked episode list

print(results["semantic"])   # Vector similarity results

```

### Async Usage with asearch()

```python
import asyncio
from nemori.api.facade import NemoriMemory

async def demo():
    # Context manager works with async patterns

    async with NemoriMemory.from_env() as mem:
        await mem.add_messages("bob", [
            {"role": "user", "content": "Tell me a joke."},
        ])
        
        # Non-blocking search using executor

        results = await mem.asearch(
            user_id="bob",
            query="joke",
            top_k_episodes=3,
            top_k_semantic=3,
        )
        print(results)

asyncio.run(demo())

```

### Resource Cleanup and Diagnostics

```python
mem = NemoriMemory()

# Retrieve global or per-user statistics

print(mem.stats())               # Global counters

print(mem.stats("alice"))        # User-specific metrics

# Delete specific episodes with optional cascade

mem.delete_episode("alice", "ep_123", cascade_semantic=True)

# Explicit cleanup if not using context manager

mem.close()

```

## Key Source Files

| File | Purpose |
|------|---------|
| [`src/api/facade.py`](https://github.com/nemori-ai/nemori/blob/main/src/api/facade.py) | Public **NemoriMemory** facade class – wiring, context manager, thin method wrappers. |
| [`src/core/memory_system.py`](https://github.com/nemori-ai/nemori/blob/main/src/core/memory_system.py) | Core engine handling buffering, episode/semantic generation, indexing, caching, and metrics. |
| [`src/config.py`](https://github.com/nemori-ai/nemori/blob/main/src/config.py) | `MemoryConfig` reads environment variables and provides defaults for the system. |
| [`src/utils/llm_client.py`](https://github.com/nemori-ai/nemori/blob/main/src/utils/llm_client.py) | Wrapper around LLM APIs used by `MemorySystem`. |
| [`src/utils/embedding_client.py`](https://github.com/nemori-ai/nemori/blob/main/src/utils/embedding_client.py) | Wrapper for embedding models – supplies vectors for semantic memory. |

## Summary

- **NemoriMemory** acts as a thin facade over the `MemorySystem` engine in [`src/core/memory_system.py`](https://github.com/nemori-ai/nemori/blob/main/src/core/memory_system.py), exposing a simplified Python API while hiding complex buffering, indexing, and semantic extraction logic.
- The facade supports **dependency injection** for custom LLM clients, embedding providers, and configuration, making it fully testable and adaptable to different deployment environments.
- All heavy operations—including **parallel hybrid search** across BM25 and vector indices, **asynchronous semantic memory generation**, and **episode batching**—are delegated to `MemorySystem`, with `NemoriMemory` merely forwarding calls.
- The class implements **context manager protocols** (`__enter__`/`__exit__`) for automatic resource cleanup, and provides both synchronous (`search`) and asynchronous (`asearch`) interfaces for flexible integration.

## Frequently Asked Questions

### What is the difference between NemoriMemory and MemorySystem?

**NemoriMemory** is the high-level facade located in [`src/api/facade.py`](https://github.com/nemori-ai/nemori/blob/main/src/api/facade.py) that provides a simplified, user-facing API for memory operations. **MemorySystem**, defined in [`src/core/memory_system.py`](https://github.com/nemori-ai/nemori/blob/main/src/core/memory_system.py), is the internal engine that performs all heavy processing including message buffering, episode generation, semantic extraction, and hybrid search indexing. The facade delegates all method calls to the engine while handling initialization and lifecycle management.

### How does NemoriMemory handle async operations?

While the underlying `MemorySystem` operates synchronously, `NemoriMemory` provides an `asearch()` method that wraps the synchronous search logic in `asyncio.get_running_loop().run_in_executor()`. This allows non-blocking search operations in async applications. For other operations like `add_messages`, the facade provides standard synchronous methods that return immediately while background semantic processing occurs in `MemorySystem`'s internal thread pools.

### Can I inject custom LLM or embedding clients into NemoriMemory?

Yes. The `__init__` method in [`src/api/facade.py`](https://github.com/nemori-ai/nemori/blob/main/src/api/facade.py) accepts optional `llm_client` and `embedding_client` parameters, allowing you to supply custom implementations of the `LLMClient` and `EmbeddingClient` interfaces. This dependency injection pattern enables unit testing with mock clients and integration with proprietary or self-hosted models rather than the default providers configured through `MemoryConfig`.

### Where does NemoriMemory store the actual memory data?

`NemoriMemory` itself does not handle storage directly; it delegates persistence to the injected `MemorySystem`. According to the source architecture in [`src/core/memory_system.py`](https://github.com/nemori-ai/nemori/blob/main/src/core/memory_system.py), the engine manages episodic storage (typically BM25-indexed text) and semantic storage (vector embeddings) through internal indexing systems. The specific storage backend—whether local disk, SQLite, or external vector databases—is determined by the `MemoryConfig` settings passed during initialization or loaded via `from_env()`.