How Nemori Ensures Semantic Coherence in Generated Episodes: A Technical Deep Dive

Nemori guarantees semantic coherence through a multi-stage pipeline combining topic-aware segmentation, LLM-driven validation, vector-based similarity search, and intelligent content merging.

The nemori-ai/nemori repository implements a sophisticated memory management system that transforms raw conversation logs into semantically coherent episodes. By enforcing strict narrative continuity at every stage—from initial segmentation to final storage—Nemori ensures that each generated episode represents a single, logically consistent event rather than fragmented or redundant memories.

Topic-Aware Batch Segmentation

Detecting Coherent Boundaries with LLM Prompts

The coherence pipeline begins in src/generation/prompts.py with the BATCH_SEGMENTATION_PROMPT. This prompt instructs the LLM to analyze conversation batches and identify natural episode boundaries by detecting topic changes, intent transitions, temporal markers, and structural signals.

According to the source code at lines 44-55, the prompt explicitly prioritizes "topical coherence over strict chronological order," allowing the system to group related messages even when they are separated by brief interruptions. This prevents semantically related content from being split across multiple episodes merely due to time gaps.

from nemori.src.generation.prompts import PromptTemplates

messages = [...]  # List of raw message dicts

prompt = PromptTemplates.get_batch_segmentation_prompt(
    count=len(messages),
    messages=PromptTemplates.format_conversation(messages)
)

# Send `prompt` to an LLM → receive JSON with episode indices & topics

Narrative Flow Enforcement During Generation

Timestamp Normalization and Causal Preservation

Once segments are identified, the EPISODE_GENERATION_PROMPT (lines 11-39 in src/generation/prompts.py) mandates that the LLM produce third-person narratives preserving causal relationships and chronological order. The prompt requires precise timestamp handling to maintain temporal consistency.

The EpisodeGenerator._determine_episode_timestamp method in src/generation/episode_generator.py (lines 86-100) enforces this by normalizing timestamps across all messages in an episode. It calculates the appropriate episode timestamp and falls back to the earliest message timestamp when necessary, ensuring every episode has a consistent temporal anchor that reflects its actual position in the conversation history.

from nemori.src.generation.episode_generator import EpisodeGenerator
from nemori.src.utils.llm_client import LLMClient
from nemori.src.config import MemoryConfig

llm = LLMClient(...)
config = MemoryConfig(...)
gen = EpisodeGenerator(llm, config)

episode = gen.generate_episode(
    user_id="user-123",
    messages=segment_messages,
    boundary_reason="topic change"
)

Finding Similar Episodes with ChromaSearchEngine

Before storage, Nemori prevents semantic redundancy through vector-based similarity analysis. The EpisodeMerger._search_similar_episodes method in src/generation/episode_merger.py (lines 54-61) utilizes ChromaSearchEngine to perform a vector search on the episode content.

This search identifies existing episodes that are semantically close to the newly generated one, preventing the storage of duplicate or overlapping memories that would fragment the user's conversation history.

from nemori.src.generation.episode_merger import EpisodeMerger
from nemori.src.storage.episode_storage import EpisodeStorage
from nemori.src.search import ChromaSearchEngine
from nemori.src.utils.embedding_client import EmbeddingClient

merger = EpisodeMerger(
    llm_client=llm,
    embedding_client=EmbeddingClient(),
    config=config,
    episode_storage=EpisodeStorage(),
    vector_search=ChromaSearchEngine()
)

merged, final_ep, old_id = merger.check_and_merge(new_episode=episode)
if merged:
    # replace old episode with `final_ep` in storage

    pass

LLM-Driven Merge Decisions and Content Synthesis

The Merge Decision Pipeline

When similar episodes are detected, EpisodeMerger._decide_merge in src/generation/episode_merger.py feeds the new episode and candidate matches into the MERGE_DECISION_PROMPT (defined in src/generation/prompts.py, lines 12-32). The LLM evaluates whether the episodes describe the same event based on temporal overlap and content continuity.

Synthesizing Coherent Merged Content

If merging is required, EpisodeMerger._merge_contents invokes the MERGE_CONTENT_PROMPT (lines 49-66 in src/generation/prompts.py). This prompt explicitly instructs the LLM to "maintain chronological flow," "preserve all important details," and "write in third-person narrative style." The resulting merged episode inherits the earliest timestamp from its constituents, creating a unified narrative that eliminates redundancy while preserving semantic completeness.

Validation and Quality Assurance

Before any episode reaches persistent storage, EpisodeGenerator.validate_episode in src/generation/episode_generator.py (lines 27-44) performs rigorous validation. This method verifies that each episode contains non-empty title and content fields, an associated user ID, a valid timestamp, and that the number of original messages matches the stored count.

Only episodes passing these structural and semantic checks are persisted, ensuring that the memory store maintains high-quality, coherent narratives without malformed or incomplete entries.

Summary

  • Topic-aware segmentation uses BATCH_SEGMENTATION_PROMPT to identify natural episode boundaries based on topical coherence rather than just chronology.
  • Narrative enforcement through EPISODE_GENERATION_PROMPT and _determine_episode_timestamp ensures causal relationships and consistent temporal anchors.
  • Vector similarity search via ChromaSearchEngine prevents semantic duplication by identifying overlapping content before storage.
  • LLM-driven merging using MERGE_DECISION_PROMPT and MERGE_CONTENT_PROMPT synthesizes redundant episodes into unified, chronologically coherent narratives.
  • Structural validation through validate_episode guarantees that only complete, well-formed episodes enter the persistent memory store.

Frequently Asked Questions

How does Nemori handle topic transitions within a conversation?

Nemori uses the BATCH_SEGMENTATION_PROMPT in src/generation/prompts.py to analyze conversation batches for topic changes, intent transitions, and structural signals. The system prioritizes topical coherence over strict chronological order, ensuring that messages discussing the same subject remain grouped together even if separated by brief interruptions.

What happens when two episodes describe the same event?

When semantic similarity search detects potential duplicates, the EpisodeMerger._decide_merge method evaluates the episodes using the MERGE_DECISION_PROMPT. If the LLM determines they describe the same event based on temporal overlap and content continuity, EpisodeMerger._merge_contents synthesizes them into a single episode using the MERGE_CONTENT_PROMPT, preserving chronological flow and all important details.

How does timestamp normalization prevent temporal inconsistencies?

The EpisodeGenerator._determine_episode_timestamp method in src/generation/episode_generator.py normalizes timestamps across all messages in an episode and falls back to the earliest message timestamp when necessary. This ensures every episode has a consistent temporal anchor that accurately reflects its position in the conversation history, preventing misordering in the memory store.

Can the semantic coherence thresholds be customized?

Yes, coherence thresholds and model settings are configurable through MemoryConfig in src/config.py. This configuration object controls parameters for the EpisodeGenerator, EpisodeMerger, and vector search components, allowing developers to adjust sensitivity for topic detection, similarity thresholds for merging, and validation requirements based on their specific use case requirements.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →