How the Nemori Semantic Generator Extracts Knowledge from Conversations

The Nemori Semantic Generator converts raw conversational episodes into high-value, persistent semantic memories by formatting dialogue context into specialized LLM prompts, extracting structured facts via prediction-correction or direct extraction modes, and wrapping results in SemanticMemory objects for long-term retrieval.

The open-source nemori-ai/nemori project implements a sophisticated memory architecture where the Semantic Generator serves as the critical bridge between transient conversations and durable knowledge. This component analyzes dialogue episodes—whether single interactions or batched histories—to identify facts worth remembering long after a chat ends. Understanding how the Semantic Generator extracts knowledge reveals how Nemori builds its persistent understanding of users over time.

Semantic Knowledge Extraction Pipeline

The extraction process follows a deterministic pipeline defined in src/generation/semantic_generator.py, transforming unstructured text into queryable memory records through seven distinct stages.

Entry Point and Mode Selection

The orchestration begins at SemanticGenerator.check_and_generate_semantic_memories (line 57), which inspects the MemoryConfig to decide between two extraction strategies. If enable_prediction_correction is True (the default), the system delegates to the PredictionCorrectionEngine; otherwise, it proceeds with direct per-episode or batch extraction based on the extract_semantic_per_episode flag.

Episode Formatting for Context Windows

Before LLM processing, raw episodes require structured formatting to fit context windows. For batch operations, the generator uses PromptTemplates.format_episodes_for_semantic to compress multiple interactions into concise summaries. When processing single episodes with original message preservation, it calls _format_episodes_from_original_messages (line 34) to maintain the full conversational turn structure including timestamps and role metadata from the original_messages field.

LLM Prompting and Structured Generation

The formatted text injects into PromptTemplates.SEMANTIC_GENERATION_PROMPT (defined in src/generation/prompts.py at line 59), which explicitly instructs the LLM to output high-value, persistent facts as a JSON list of statements. The generator then invokes LLMClient.generate_json_response via the internal _generate_semantic_with_llm method (line 74), sending the prompt to the configured model (e.g., GPT-4) and receiving structured JSON payloads.

Statement Extraction and Memory Construction

Depending on the active mode, the system handles LLM responses differently. In prediction-correction mode, PredictionCorrectionEngine.learn_from_episode_simplified executes a "predict → compare → extract" workflow to refine statements. In direct per-episode mode, the generator flattens top-level JSON keys (user_profile, experience, knowledge, other) into a unified statements list. Finally, _convert_to_semantic_memories (line 70) instantiates each statement as a SemanticMemory model, copying the earliest episode timestamp to preserve temporal ordering.

Two Extraction Modes: Prediction-Correction vs. Per-Episode

Nemori supports dual extraction strategies controlled by configuration flags in src/config.py, allowing developers to balance extraction depth against computational cost.

Prediction-Correction Mode (default when enable_prediction_correction is True) runs a two-step validation where the engine predicts user attributes based on existing memories, compares these predictions against actual episode content, and extracts only the discrepancies as new knowledge. This reduces noise by focusing strictly on novel information.

Per-Episode Direct Extraction activates when extract_semantic_per_episode is True, bypassing the prediction step to extract all candidate facts from individual episodes using a richer schema that categorizes memories into distinct semantic types (user_profile, experience, etc.).

Configuration Flags That Control Extraction

Three boolean flags in MemoryConfig govern the Semantic Generator's behavior, as referenced in SemanticGenerator.get_generator_stats:

  • enable_prediction_correction: Toggles the two-step prediction-correction workflow. Default: True.
  • extract_semantic_per_episode: Enables direct single-episode extraction with multi-type categorization. Default: False.
  • enable_semantic_memory: Master switch for the entire semantic memory subsystem. Default: True.

Implementation Example: Generating Semantic Memories

The following Python example demonstrates initializing the generator and extracting memories from a new episode:

from datetime import datetime
from nemori.src.utils import LLMClient, EmbeddingClient
from nemori.src.generation.semantic_generator import SemanticGenerator
from nemori.src.config import MemoryConfig
from nemori.src.models import Episode

# Initialise required components

llm = LLMClient(model="gpt-4o-mini")
emb = EmbeddingClient(model="text-embedding-3-large")
cfg = MemoryConfig()  # loads default flags

# Build the generator

generator = SemanticGenerator(
    llm_client=llm,
    embedding_client=emb,
    config=cfg
)

# Prepare a new episode

new_ep = Episode(
    user_id="user-123",
    title="Planning a weekend hike",
    content="I want to hike this weekend. Which trail are you thinking of?",
    original_messages=[
        {"role": "user", "content": "I want to hike this weekend.", "metadata": {"timestamp": "2024-03-01T10:12:00"}},
        {"role": "assistant", "content": "Which trail are you thinking of?", "metadata": {"timestamp": "2024-03-01T10:12:05"}}
    ],
    created_at=datetime.utcnow(),
    timestamp=datetime.utcnow()
)

# Extract semantic memories

semantic_memories = generator.check_and_generate_semantic_memories(
    user_id="user-123",
    new_episode=new_ep,
    existing_episodes=[],
    existing_semantic_memories=None
)

# Inspect results

for mem in semantic_memories:
    print(f"[{mem.knowledge_type}] {mem.content}")

Under the hood, because extract_semantic_per_episode defaults to False, this invokes generate_semantic_memories, formats the episode using original messages via _format_episodes_from_original_messages, and wraps LLM outputs in SemanticMemory objects via _convert_to_semantic_memories.

Summary

Frequently Asked Questions

What is the difference between prediction-correction and per-episode extraction in Nemori?

Prediction-correction extraction (default) uses PredictionCorrectionEngine.learn_from_episode_simplified to validate new facts against existing knowledge before extraction, reducing redundancy. Per-episode extraction processes each conversation independently through direct LLM prompting without the prediction validation step, categorizing outputs into user_profile, experience, knowledge, and other buckets.

How does the Semantic Generator format conversations for the LLM?

The generator selects formatting based on configuration. For batch processing, it compresses multiple episodes using PromptTemplates.format_episodes_for_semantic. For single episodes requiring full fidelity, _format_episodes_from_original_messages preserves original message roles, content, and timestamps from the original_messages field.

Where are the semantic generation prompts defined in the Nemori codebase?

Prompt templates reside in src/generation/prompts.py. Specifically, PromptTemplates.SEMANTIC_GENERATION_PROMPT (line 59) contains the instructions that guide the LLM to extract high-value, persistent facts and return them as structured JSON statements.

Can I disable the Semantic Generator while keeping other Nemori memory features active?

Yes. Set enable_semantic_memory to False in your MemoryConfig. This master switch disables the entire semantic memory subsystem—including the Semantic Generator—while preserving episodic memory and other non-semantic components.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →