Nemori Repository Architecture: 10 Key Files for Understanding the Codebase
The Nemori repository organizes its Python codebase into ten modular layers—from configuration and utilities to storage, services, and a public API façade—with critical files like src/config.py, src/core/memory_system.py, and src/api/facade.py serving as the architectural backbone.
Nemori is an open-source memory system for conversational AI, structured as a clean, modular Python package. Understanding the key files in the Nemori repository reveals how the system isolates concerns between vector storage, LLM generation, and retrieval strategies, making it easy to swap implementations without touching core logic.
Configuration: The Central Settings Hub
Start with src/config.py, the single source of truth for environment variables, model parameters, vector-store settings, and feature flags. All components import this module to ensure consistent configuration across the system.
Utilities: Vendor-Agnostic API Wrappers
The src/utils/ directory contains injectable helpers that keep core logic free of vendor-specific details.
LLM and Embedding Clients
src/utils/llm_client.py– Wraps language-model APIs (OpenAI, Anthropic) with retry logic and token usage tracking.src/utils/embedding_client.py– Handles embedding service interactions for semantic search.src/utils/token_counter.py– Budgets LLM calls by counting tokens.src/utils/performance.py– Profiles latency and throughput of critical sections.
Storage Abstractions: Database Independence
The src/storage/ layer isolates the system from specific database engines through abstract interfaces.
Core Storage Files
src/storage/base_storage.py– Abstract base class defining CRUD operations for persisted entities.src/storage/episode_storage.py– Concrete implementation for Episode objects (message groupings).src/storage/semantic_storage.py– Manages vector-store interactions and wraps the embedding client.
Service Layer: Orchestration and Glue
Located in src/services/, these components coordinate between utilities, storage, and generation pipelines.
src/services/task_manager.py– Orchestrates asynchronous jobs like batch generation and vector indexing.src/services/providers.py– Factory functions that return configured instances of utilities for dependency injection.src/services/metrics.py– Collects runtime statistics including latency, token usage, and success rates.src/services/event_bus.py– Publish/subscribe mechanism for domain events (episode_created, search_completed).src/services/cache.py– In-memory caching for expensive lookups like recent embeddings.
Search Strategies: Pluggable Retrieval
The src/search/ package implements multiple retrieval algorithms behind a common interface:
src/search/bm25_search.py– Classic BM25 over tokenized text.src/search/chroma_search.py– Vector similarity using Chroma.src/search/original_message_search.py– Keyword search over raw messages.src/search/episode_original_message_search.py– Episode-level keyword search.src/search/unified_search.py– Combines multiple back-ends into a single ranking pipeline.
Each file can be swapped via configuration to experiment with different retrieval mechanisms without modifying consumer code.
Domain Models: Core Data Structures
The src/models/ directory defines the data structures passed between storage and generation:
src/models/message.py– Represents single user/system messages (content, role, timestamps).src/models/episode.py– Collections of messages with metadata for generation and evaluation.src/models/semantic.py– Embedding metadata including vectors and distance scores.
Infrastructure Adapters: Concrete Implementations
The src/infrastructure/ directory bridges abstractions to actual persistence technologies:
src/infrastructure/repositories.py– Implementsbase_storagecontracts for SQL, NoSQL, or file-system back-ends.src/infrastructure/indices.py– Manages vector-index structures like Chroma collections.
Generation Pipeline: LLM Orchestration
The src/generation/ directory handles the end-to-end creation of responses.
src/generation/semantic_generator.py– Converts queries to dense embeddings and fetches relevant passages.src/generation/episode_generator.py– Orchestrates LLM calls to produce new episodes from retrieved context.src/generation/episode_merger.py– Merges candidate episodes into coherent final outputs.src/generation/batch_segmenter.py– Chunks large corpora for batch processing.src/generation/prediction_correction_engine.py– Post-processes LLM outputs for hallucination detection.
The pipeline is configurable via src/config.py; you can swap any stage without modifying other code.
Core Memory System: Session Orchestration
The src/core/ directory contains the primary runtime engine.
src/core/memory_system.py– Central orchestrator coordinating retrieval, generation, and storage for conversational sessions. This is the primary entry point for client code.src/core/message_buffer.py– Maintains the sliding window of recent messages for context truncation.
Public API: The Application Entry Point
src/api/facade.py exposes core functionality (run_query, add_episode) as simple function calls or REST-style endpoints. This is the file you import in downstream applications.
Code Examples
Running a Query Through the High-Level Facade
from nemori.api.facade import run_query
from nemori.config import Settings
# Settings are loaded from environment variables / .env file
settings = Settings()
# Ask the system a question; the facade handles retrieval, generation, and storage
answer = run_query(
query="What are the main advantages of using vector search over keyword search?",
settings=settings,
)
print("Answer:", answer)
This call traverses semantic_generator → episode_generator → memory_system and persists the resulting episode via episode_storage.
Manually Assembling the Generation Pipeline
from nemori.core.memory_system import MemorySystem
from nemori.services.providers import get_llm_client, get_embedding_client
from nemori.config import Settings
settings = Settings()
llm = get_llm_client(settings)
embedder = get_embedding_client(settings)
mem = MemorySystem(settings, llm_client=llm, embedding_client=embedder)
# Retrieve relevant context manually
context = mem.retrieve("vector databases benefits")
# Generate a response
episode = mem.generate(context, user_query="Explain the benefits.")
print(episode.messages[-1].content) # final assistant message
This demonstrates direct interaction with the core components without the façade shortcut.
Batch Indexing of a Document Collection
from nemori.generation.batch_segmenter import BatchSegmenter
from nemori.storage.semantic_storage import SemanticStorage
from nemori.services.providers import get_embedding_client
from nemori.config import Settings
settings = Settings()
embedder = get_embedding_client(settings)
semantic_store = SemanticStorage(settings, embedder)
segmenter = BatchSegmenter(chunk_size=512, overlap=64)
documents = ["..."] # list of raw text strings
for doc in documents:
chunks = segmenter.segment(doc)
semantic_store.add_chunks(chunks) # stores vectors for later BM25/Chroma search
This creates a searchable vector index that semantic_generator will later query.
Summary
src/config.pyserves as the central configuration hub for all environment variables and model settings.- The storage layer (
base_storage.py,episode_storage.py,semantic_storage.py) abstracts database operations to keep the system technology-agnostic. src/core/memory_system.pyis the primary orchestrator, whilesrc/api/facade.pyprovides the public interface.- Search strategies are modular and swappable via the
src/search/directory, supporting BM25, Chroma vector search, and unified ranking. - The generation pipeline (
episode_generator.py,semantic_generator.py) is highly configurable and handles everything from retrieval to hallucination correction.
Frequently Asked Questions
What is the entry point for using Nemori in my application?
Import run_query from src/api/facade.py. This function wraps the entire retrieval and generation pipeline, handling everything from semantic search to episode persistence without requiring manual orchestration of the core components.
How does Nemori handle different LLM providers?
The repository uses src/services/providers.py as a factory for dependency injection, returning configured instances of src/utils/llm_client.py and src/utils/embedding_client.py. This keeps vendor-specific API logic isolated from the core memory system.
Can I swap the search algorithm without changing the core logic?
Yes. The src/search/ directory implements a common interface across BM25, Chroma vector search, and unified ranking. You can configure which implementation to use via src/config.py, allowing experiments with different retrieval mechanisms without modifying src/core/memory_system.py.
Where does Nemori store conversation history?
Conversation history is managed by src/storage/episode_storage.py, which implements the abstract interface defined in src/storage/base_storage.py. Episodes—collections of messages—are persisted through this layer, while the specific database technology (SQL, NoSQL, or file-system) is handled by src/infrastructure/repositories.py.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →