Nemori Repository Architecture: 10 Key Files for Understanding the Codebase

The Nemori repository organizes its Python codebase into ten modular layers—from configuration and utilities to storage, services, and a public API façade—with critical files like src/config.py, src/core/memory_system.py, and src/api/facade.py serving as the architectural backbone.

Nemori is an open-source memory system for conversational AI, structured as a clean, modular Python package. Understanding the key files in the Nemori repository reveals how the system isolates concerns between vector storage, LLM generation, and retrieval strategies, making it easy to swap implementations without touching core logic.

Configuration: The Central Settings Hub

Start with src/config.py, the single source of truth for environment variables, model parameters, vector-store settings, and feature flags. All components import this module to ensure consistent configuration across the system.

Utilities: Vendor-Agnostic API Wrappers

The src/utils/ directory contains injectable helpers that keep core logic free of vendor-specific details.

LLM and Embedding Clients

Storage Abstractions: Database Independence

The src/storage/ layer isolates the system from specific database engines through abstract interfaces.

Core Storage Files

Service Layer: Orchestration and Glue

Located in src/services/, these components coordinate between utilities, storage, and generation pipelines.

Search Strategies: Pluggable Retrieval

The src/search/ package implements multiple retrieval algorithms behind a common interface:

Each file can be swapped via configuration to experiment with different retrieval mechanisms without modifying consumer code.

Domain Models: Core Data Structures

The src/models/ directory defines the data structures passed between storage and generation:

Infrastructure Adapters: Concrete Implementations

The src/infrastructure/ directory bridges abstractions to actual persistence technologies:

Generation Pipeline: LLM Orchestration

The src/generation/ directory handles the end-to-end creation of responses.

The pipeline is configurable via src/config.py; you can swap any stage without modifying other code.

Core Memory System: Session Orchestration

The src/core/ directory contains the primary runtime engine.

  • src/core/memory_system.py – Central orchestrator coordinating retrieval, generation, and storage for conversational sessions. This is the primary entry point for client code.
  • src/core/message_buffer.py – Maintains the sliding window of recent messages for context truncation.

Public API: The Application Entry Point

src/api/facade.py exposes core functionality (run_query, add_episode) as simple function calls or REST-style endpoints. This is the file you import in downstream applications.

Code Examples

Running a Query Through the High-Level Facade

from nemori.api.facade import run_query
from nemori.config import Settings

# Settings are loaded from environment variables / .env file

settings = Settings()

# Ask the system a question; the facade handles retrieval, generation, and storage

answer = run_query(
    query="What are the main advantages of using vector search over keyword search?",
    settings=settings,
)

print("Answer:", answer)

This call traverses semantic_generator → episode_generator → memory_system and persists the resulting episode via episode_storage.

Manually Assembling the Generation Pipeline

from nemori.core.memory_system import MemorySystem
from nemori.services.providers import get_llm_client, get_embedding_client
from nemori.config import Settings

settings = Settings()
llm = get_llm_client(settings)
embedder = get_embedding_client(settings)

mem = MemorySystem(settings, llm_client=llm, embedding_client=embedder)

# Retrieve relevant context manually

context = mem.retrieve("vector databases benefits")

# Generate a response

episode = mem.generate(context, user_query="Explain the benefits.")
print(episode.messages[-1].content)   # final assistant message

This demonstrates direct interaction with the core components without the façade shortcut.

Batch Indexing of a Document Collection

from nemori.generation.batch_segmenter import BatchSegmenter
from nemori.storage.semantic_storage import SemanticStorage
from nemori.services.providers import get_embedding_client
from nemori.config import Settings

settings = Settings()
embedder = get_embedding_client(settings)
semantic_store = SemanticStorage(settings, embedder)

segmenter = BatchSegmenter(chunk_size=512, overlap=64)

documents = ["..."]  # list of raw text strings

for doc in documents:
    chunks = segmenter.segment(doc)
    semantic_store.add_chunks(chunks)  # stores vectors for later BM25/Chroma search

This creates a searchable vector index that semantic_generator will later query.

Summary

Frequently Asked Questions

What is the entry point for using Nemori in my application?

Import run_query from src/api/facade.py. This function wraps the entire retrieval and generation pipeline, handling everything from semantic search to episode persistence without requiring manual orchestration of the core components.

How does Nemori handle different LLM providers?

The repository uses src/services/providers.py as a factory for dependency injection, returning configured instances of src/utils/llm_client.py and src/utils/embedding_client.py. This keeps vendor-specific API logic isolated from the core memory system.

Can I swap the search algorithm without changing the core logic?

Yes. The src/search/ directory implements a common interface across BM25, Chroma vector search, and unified ranking. You can configure which implementation to use via src/config.py, allowing experiments with different retrieval mechanisms without modifying src/core/memory_system.py.

Where does Nemori store conversation history?

Conversation history is managed by src/storage/episode_storage.py, which implements the abstract interface defined in src/storage/base_storage.py. Episodes—collections of messages—are persisted through this layer, while the specific database technology (SQL, NoSQL, or file-system) is handled by src/infrastructure/repositories.py.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →