How Nemori Implements Semantic Memory: A Deep Dive into the Vector-Indexed Knowledge Store

Nemori implements semantic memory as a high-performance, persistent, vector-indexed knowledge store that extracts facts from conversations using LLM-based generation, stores them in JSONL files with in-memory indexes, and makes them searchable via ChromaDB and BM25.

The nemori-ai/nemori repository provides a production-grade implementation of semantic memory designed for long-term knowledge retention across user sessions. This article examines how Nemori structures, generates, persists, and retrieves semantic facts using a multi-layered architecture that balances durability with search performance.

Architecture Overview

Nemori’s semantic memory system consists of four tightly-coupled layers that handle everything from data structure definition to vector indexing:

Layer Responsibility Core modules
Data model Defines the structure of a semantic fact (content, source episodes, timestamps, etc.) src/models/semantic.py
Storage Persistent JSONL files + in-memory indexes for fast reads/writes src/storage/semantic_storage.py
Generation Turns new episodes into semantic facts, either per-episode extraction or a prediction-correction pipeline src/generation/semantic_generator.py
Runtime orchestration Asynchronously schedules generation, caches embeddings, deduplicates, and updates vector & lexical indexes src/core/memory_system.py, src/services/task_manager.py, src/services/cache.py

The Data Model

At the core of the system is the SemanticMemory dataclass defined in src/models/semantic.py. This immutable structure represents a single factual statement extracted from conversation history:

@dataclass
class SemanticMemory:
    content: str                # The factual statement

    knowledge_type: str         # E.g. "knowledge"

    user_id: str                # Owner

    created_at: datetime = field(default_factory=datetime.now)
    memory_id: str = field(default_factory=lambda: str(uuid.uuid4()))
    source_episodes: List[str] = field(default_factory=list)  # Episodes that produced this fact

    confidence: float = 0.8
    metadata: Dict[str, Any] = field(default_factory=dict)

Every semantic fact maintains a list of source_episodes that trace back to the original conversation segments, enabling full auditability of how knowledge was derived.

Persistent Storage Layer

The SemanticStorage class in src/storage/semantic_storage.py handles durability using a simple but effective strategy:

  • One JSONL file per user: Data persists to <user>_semantic.jsonl on disk
  • In-memory indexes: Maintains memory_id → file_path and user_id → [memory_ids] maps for O(1) lookups
  • Thread safety: All reads/writes are protected by a per-user RLock to guarantee consistency during concurrent access

Key persistence methods include:

  • save_semantic_memory(memory) – Appends a new line to the JSONL file and updates indexes
  • list_user_items(user_id) – Loads the file, sorts by created_at, and returns ordered facts
  • delete(user_id, memory_id) – Rewrites the file without the target line

Semantic Generation Pipeline

The SemanticGenerator in src/generation/semantic_generator.py transforms raw conversation episodes into structured facts. The system supports two distinct generation modes controlled via MemoryConfig:

Per-Episode Direct Extraction

When extract_semantic_per_episode=True, every newly created episode immediately invokes generate_semantic_memories([episode]). The LLM receives the original messages (formatted via _format_episodes_from_original_messages) and returns a JSON object containing separate lists for user_profile, experience, knowledge, and other. Each item is converted to a SemanticMemory object.

Prediction-Correction Pipeline

When enable_prediction_correction=True, the system employs a more sophisticated approach. The new episode is compared against all existing episodes and current semantic memories. The PredictionCorrectionEngine (instantiated inside SemanticGenerator) executes a two-step "predict-then-correct" process that can merge similar facts, split compound statements, or refine existing knowledge based on new evidence.

Both paths ultimately call _convert_to_semantic_memories, which creates SemanticMemory objects while preserving the earliest episode timestamp for chronological consistency.

Runtime Orchestration and Async Processing

The MemorySystem class in src/core/memory_system.py orchestrates the entire lifecycle of semantic memory generation and indexing.

When messages are ingested via add_messages, the system buffers them and eventually creates an episode via _create_episode_from_messages. Upon saving the episode, _schedule_semantic_generation is invoked:

def _schedule_semantic_generation(self, owner_id: str, new_episode: Episode):
    task_id = f"{owner_id}_{new_episode.episode_id}"
    future = self.semantic_task_manager.submit(
        self._async_generate_semantic_memories, owner_id, new_episode)
    self._semantic_generation_futures[task_id] = {"future": future, ...}
    future.add_done_callback(lambda f: self._on_semantic_generation_complete(task_id, f))

The SemanticTaskManager (defined in src/services/task_manager.py) wraps a ThreadPoolExecutor with retry logic to handle transient failures.

The async generation pipeline (_async_generate_semantic_memories) performs the following:

  1. Loads cached episodes and semantic memories via PerUserCache
  2. Retrieves or computes embeddings for existing memories using SemanticEmbeddingCache to avoid duplicate model calls
  3. Invokes self.semantic_generator.check_and_generate_semantic_memories(...)
  4. For each new SemanticMemory:
    • Embeds content via self.embedding_client.embed_text
    • Checks for duplicates using _is_duplicate_semantic_memory
    • Persists via self._semantic_repository.save
    • Indexes in ChromaDB (self.search_engine.add_semantic_memory) and BM25 lexical index

This background processing ensures the main request path remains fast while maintaining a rich, searchable knowledge base.

Vector and Lexical Indexing

Nemori employs a dual-index strategy for semantic memory retrieval:

  • ChromaDB (src/search/chroma_search.py): Stores embeddings in per-user collections, enabling semantic similarity search. Embeddings are cached in SemanticEmbeddingCache to minimize redundant computation.
  • BM25 (src/search/bm25_search.py): Provides fast lexical search capabilities for both episodes and semantic memories, complementing the vector-based semantic search.

Both indexes are updated immediately after a semantic memory is persisted, ensuring new facts are immediately discoverable.

Configuration and Feature Flags

All semantic memory behaviors are controlled via MemoryConfig in src/config.py:

enable_semantic_memory          # Master switch for the entire subsystem

extract_semantic_per_episode    # Enable direct per-episode extraction mode

enable_prediction_correction    # Enable the advanced prediction-correction pipeline

semantic_generation_workers     # Thread pool size for async generation tasks

semantic_cache_ttl              # Time-to-live for semantic memory cache entries

These flags allow operators to tune the trade-off between processing latency and knowledge extraction sophistication.

Code Examples

Adding Messages and Automatically Generating Semantic Memory

from nemori import MemorySystem

# Initialise the system (uses defaults from env or config)

mem = MemorySystem()

# Simulate a user conversation

owner = "user_123"
messages = [
    {"role": "user", "content": "I love hiking in the Alps."},
    {"role": "assistant", "content": "That sounds great!"},
    {"role": "user", "content": "I usually go in September."}
]

# Add messages – this will create an episode and schedule semantic generation

result = mem.add_messages(owner_id=owner, messages=messages)

print(result["episodes_created"][0]["title"])   # → Generated episode title

# Semantic memories will be generated asynchronously; you can wait for tasks if needed:

# mem._semantic_generation_futures contains the futures.

All heavy lifting—including LLM calls, embedding computation, deduplication, and indexing—happens in the background via MemorySystem.add_messages → _schedule_semantic_generation.

Retrieving a User’s Semantic Memories


# Direct repository access (fast, reads from JSONL)

semantic_repo = mem._semantic_repository   # type: SemanticRepository

memories = semantic_repo.list_by_user("user_123")

for mem in memories[:5]:
    print(f"{mem.created_at.date()}: {mem.content}")

Alternatively, use the vector search for semantic retrieval:

results = mem.search(
    owner_id="user_123",
    query="What activities does the user enjoy?",
    memory_types=["semantic"]
)

for r in results:
    print(r["content"])

Source: MemorySystem.search combines vector and lexical search capabilities.

Manually Triggering Semantic Generation


# Assume you already have an Episode object `ep`

mem._async_generate_semantic_memories(owner_id="user_123", new_episode=ep)

This method is useful for testing or backfilling historical data without going through the standard message ingestion flow.

Summary

  • Nemori implements semantic memory through a four-layer architecture spanning data models, JSONL persistence, LLM-based generation, and async runtime orchestration.
  • Immutable facts are stored as SemanticMemory objects in src/models/semantic.py, tracking content, source episodes, and confidence scores.
  • Dual persistence strategy uses per-user JSONL files (src/storage/semantic_storage.py) for durability and in-memory indexes for fast access, protected by RLock for thread safety.
  • Two extraction modes are available: per-episode direct extraction (extract_semantic_per_episode) and an advanced prediction-correction pipeline (enable_prediction_correction) that merges and refines facts across episodes.
  • Async processing via MemorySystem._schedule_semantic_generation and SemanticTaskManager keeps the main request path fast while handling LLM calls, embeddings, and indexing in background threads.
  • Dual indexing through ChromaDB (vector similarity) and BM25 (lexical search) ensures semantic memories are immediately discoverable after persistence.

Frequently Asked Questions

What is the difference between Nemori’s semantic memory and episodic memory?

Episodic memory stores raw conversation history as distinct episodes (sequences of messages), while semantic memory extracts generalized facts and knowledge from those episodes into immutable SemanticMemory objects. The semantic layer distills long-term knowledge (e.g., "User likes hiking in September") from the temporal episode stream, making it searchable independently of the original conversation context.

How does Nemori prevent duplicate semantic memories?

Nemori implements deduplication in MemorySystem._async_generate_semantic_memories using the _is_duplicate_semantic_memory method. Before persisting a new SemanticMemory, the system compares its embedding against existing memories using vector similarity. If a sufficiently similar memory exists (based on configurable thresholds), the new fact is discarded or merged with the existing entry to prevent knowledge base bloat.

Can I use Nemori’s semantic memory without the prediction-correction pipeline?

Yes. The prediction-correction pipeline is optional and controlled by the enable_prediction_correction flag in MemoryConfig. If disabled, Nemori defaults to per-episode direct extraction (extract_semantic_per_episode=True), where each new episode immediately triggers generate_semantic_memories([episode]) without cross-referencing existing knowledge. This mode is faster and simpler but less sophisticated at merging related facts across multiple episodes.

What storage backend does Nemori use for semantic memory vectors?

Nemori uses ChromaDB as the primary vector storage backend, implemented in src/search/chroma_search.py. Semantic memories are stored in per-user collections with their embeddings, enabling efficient similarity search. The system also maintains a BM25 lexical index (src/search/bm25_search.py) for keyword-based retrieval, providing a hybrid search capability that combines semantic similarity with exact text matching.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →