How Nemori Implements Semantic Memory: A Deep Dive into the Vector-Indexed Knowledge Store
Nemori implements semantic memory as a high-performance, persistent, vector-indexed knowledge store that extracts facts from conversations using LLM-based generation, stores them in JSONL files with in-memory indexes, and makes them searchable via ChromaDB and BM25.
The nemori-ai/nemori repository provides a production-grade implementation of semantic memory designed for long-term knowledge retention across user sessions. This article examines how Nemori structures, generates, persists, and retrieves semantic facts using a multi-layered architecture that balances durability with search performance.
Architecture Overview
Nemori’s semantic memory system consists of four tightly-coupled layers that handle everything from data structure definition to vector indexing:
| Layer | Responsibility | Core modules |
|---|---|---|
| Data model | Defines the structure of a semantic fact (content, source episodes, timestamps, etc.) | src/models/semantic.py |
| Storage | Persistent JSONL files + in-memory indexes for fast reads/writes | src/storage/semantic_storage.py |
| Generation | Turns new episodes into semantic facts, either per-episode extraction or a prediction-correction pipeline | src/generation/semantic_generator.py |
| Runtime orchestration | Asynchronously schedules generation, caches embeddings, deduplicates, and updates vector & lexical indexes | src/core/memory_system.py, src/services/task_manager.py, src/services/cache.py |
The Data Model
At the core of the system is the SemanticMemory dataclass defined in src/models/semantic.py. This immutable structure represents a single factual statement extracted from conversation history:
@dataclass
class SemanticMemory:
content: str # The factual statement
knowledge_type: str # E.g. "knowledge"
user_id: str # Owner
created_at: datetime = field(default_factory=datetime.now)
memory_id: str = field(default_factory=lambda: str(uuid.uuid4()))
source_episodes: List[str] = field(default_factory=list) # Episodes that produced this fact
confidence: float = 0.8
metadata: Dict[str, Any] = field(default_factory=dict)
Every semantic fact maintains a list of source_episodes that trace back to the original conversation segments, enabling full auditability of how knowledge was derived.
Persistent Storage Layer
The SemanticStorage class in src/storage/semantic_storage.py handles durability using a simple but effective strategy:
- One JSONL file per user: Data persists to
<user>_semantic.jsonlon disk - In-memory indexes: Maintains
memory_id → file_pathanduser_id → [memory_ids]maps for O(1) lookups - Thread safety: All reads/writes are protected by a per-user
RLockto guarantee consistency during concurrent access
Key persistence methods include:
save_semantic_memory(memory)– Appends a new line to the JSONL file and updates indexeslist_user_items(user_id)– Loads the file, sorts bycreated_at, and returns ordered factsdelete(user_id, memory_id)– Rewrites the file without the target line
Semantic Generation Pipeline
The SemanticGenerator in src/generation/semantic_generator.py transforms raw conversation episodes into structured facts. The system supports two distinct generation modes controlled via MemoryConfig:
Per-Episode Direct Extraction
When extract_semantic_per_episode=True, every newly created episode immediately invokes generate_semantic_memories([episode]). The LLM receives the original messages (formatted via _format_episodes_from_original_messages) and returns a JSON object containing separate lists for user_profile, experience, knowledge, and other. Each item is converted to a SemanticMemory object.
Prediction-Correction Pipeline
When enable_prediction_correction=True, the system employs a more sophisticated approach. The new episode is compared against all existing episodes and current semantic memories. The PredictionCorrectionEngine (instantiated inside SemanticGenerator) executes a two-step "predict-then-correct" process that can merge similar facts, split compound statements, or refine existing knowledge based on new evidence.
Both paths ultimately call _convert_to_semantic_memories, which creates SemanticMemory objects while preserving the earliest episode timestamp for chronological consistency.
Runtime Orchestration and Async Processing
The MemorySystem class in src/core/memory_system.py orchestrates the entire lifecycle of semantic memory generation and indexing.
When messages are ingested via add_messages, the system buffers them and eventually creates an episode via _create_episode_from_messages. Upon saving the episode, _schedule_semantic_generation is invoked:
def _schedule_semantic_generation(self, owner_id: str, new_episode: Episode):
task_id = f"{owner_id}_{new_episode.episode_id}"
future = self.semantic_task_manager.submit(
self._async_generate_semantic_memories, owner_id, new_episode)
self._semantic_generation_futures[task_id] = {"future": future, ...}
future.add_done_callback(lambda f: self._on_semantic_generation_complete(task_id, f))
The SemanticTaskManager (defined in src/services/task_manager.py) wraps a ThreadPoolExecutor with retry logic to handle transient failures.
The async generation pipeline (_async_generate_semantic_memories) performs the following:
- Loads cached episodes and semantic memories via
PerUserCache - Retrieves or computes embeddings for existing memories using
SemanticEmbeddingCacheto avoid duplicate model calls - Invokes
self.semantic_generator.check_and_generate_semantic_memories(...) - For each new
SemanticMemory:- Embeds content via
self.embedding_client.embed_text - Checks for duplicates using
_is_duplicate_semantic_memory - Persists via
self._semantic_repository.save - Indexes in ChromaDB (
self.search_engine.add_semantic_memory) and BM25 lexical index
- Embeds content via
This background processing ensures the main request path remains fast while maintaining a rich, searchable knowledge base.
Vector and Lexical Indexing
Nemori employs a dual-index strategy for semantic memory retrieval:
- ChromaDB (
src/search/chroma_search.py): Stores embeddings in per-user collections, enabling semantic similarity search. Embeddings are cached inSemanticEmbeddingCacheto minimize redundant computation. - BM25 (
src/search/bm25_search.py): Provides fast lexical search capabilities for both episodes and semantic memories, complementing the vector-based semantic search.
Both indexes are updated immediately after a semantic memory is persisted, ensuring new facts are immediately discoverable.
Configuration and Feature Flags
All semantic memory behaviors are controlled via MemoryConfig in src/config.py:
enable_semantic_memory # Master switch for the entire subsystem
extract_semantic_per_episode # Enable direct per-episode extraction mode
enable_prediction_correction # Enable the advanced prediction-correction pipeline
semantic_generation_workers # Thread pool size for async generation tasks
semantic_cache_ttl # Time-to-live for semantic memory cache entries
These flags allow operators to tune the trade-off between processing latency and knowledge extraction sophistication.
Code Examples
Adding Messages and Automatically Generating Semantic Memory
from nemori import MemorySystem
# Initialise the system (uses defaults from env or config)
mem = MemorySystem()
# Simulate a user conversation
owner = "user_123"
messages = [
{"role": "user", "content": "I love hiking in the Alps."},
{"role": "assistant", "content": "That sounds great!"},
{"role": "user", "content": "I usually go in September."}
]
# Add messages – this will create an episode and schedule semantic generation
result = mem.add_messages(owner_id=owner, messages=messages)
print(result["episodes_created"][0]["title"]) # → Generated episode title
# Semantic memories will be generated asynchronously; you can wait for tasks if needed:
# mem._semantic_generation_futures contains the futures.
All heavy lifting—including LLM calls, embedding computation, deduplication, and indexing—happens in the background via MemorySystem.add_messages → _schedule_semantic_generation.
Retrieving a User’s Semantic Memories
# Direct repository access (fast, reads from JSONL)
semantic_repo = mem._semantic_repository # type: SemanticRepository
memories = semantic_repo.list_by_user("user_123")
for mem in memories[:5]:
print(f"{mem.created_at.date()}: {mem.content}")
Alternatively, use the vector search for semantic retrieval:
results = mem.search(
owner_id="user_123",
query="What activities does the user enjoy?",
memory_types=["semantic"]
)
for r in results:
print(r["content"])
Source: MemorySystem.search combines vector and lexical search capabilities.
Manually Triggering Semantic Generation
# Assume you already have an Episode object `ep`
mem._async_generate_semantic_memories(owner_id="user_123", new_episode=ep)
This method is useful for testing or backfilling historical data without going through the standard message ingestion flow.
Summary
- Nemori implements semantic memory through a four-layer architecture spanning data models, JSONL persistence, LLM-based generation, and async runtime orchestration.
- Immutable facts are stored as
SemanticMemoryobjects insrc/models/semantic.py, tracking content, source episodes, and confidence scores. - Dual persistence strategy uses per-user JSONL files (
src/storage/semantic_storage.py) for durability and in-memory indexes for fast access, protected byRLockfor thread safety. - Two extraction modes are available: per-episode direct extraction (
extract_semantic_per_episode) and an advanced prediction-correction pipeline (enable_prediction_correction) that merges and refines facts across episodes. - Async processing via
MemorySystem._schedule_semantic_generationandSemanticTaskManagerkeeps the main request path fast while handling LLM calls, embeddings, and indexing in background threads. - Dual indexing through ChromaDB (vector similarity) and BM25 (lexical search) ensures semantic memories are immediately discoverable after persistence.
Frequently Asked Questions
What is the difference between Nemori’s semantic memory and episodic memory?
Episodic memory stores raw conversation history as distinct episodes (sequences of messages), while semantic memory extracts generalized facts and knowledge from those episodes into immutable SemanticMemory objects. The semantic layer distills long-term knowledge (e.g., "User likes hiking in September") from the temporal episode stream, making it searchable independently of the original conversation context.
How does Nemori prevent duplicate semantic memories?
Nemori implements deduplication in MemorySystem._async_generate_semantic_memories using the _is_duplicate_semantic_memory method. Before persisting a new SemanticMemory, the system compares its embedding against existing memories using vector similarity. If a sufficiently similar memory exists (based on configurable thresholds), the new fact is discarded or merged with the existing entry to prevent knowledge base bloat.
Can I use Nemori’s semantic memory without the prediction-correction pipeline?
Yes. The prediction-correction pipeline is optional and controlled by the enable_prediction_correction flag in MemoryConfig. If disabled, Nemori defaults to per-episode direct extraction (extract_semantic_per_episode=True), where each new episode immediately triggers generate_semantic_memories([episode]) without cross-referencing existing knowledge. This mode is faster and simpler but less sophisticated at merging related facts across multiple episodes.
What storage backend does Nemori use for semantic memory vectors?
Nemori uses ChromaDB as the primary vector storage backend, implemented in src/search/chroma_search.py. Semantic memories are stored in per-user collections with their embeddings, enabling efficient similarity search. The system also maintains a BM25 lexical index (src/search/bm25_search.py) for keyword-based retrieval, providing a hybrid search capability that combines semantic similarity with exact text matching.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →