# How the Nemori Semantic Generator Extracts Knowledge from Conversations

> Discover how the Nemori Semantic Generator transforms conversations into structured knowledge using LLM prompts and advanced extraction modes for long-term memory.

- Repository: [Nemori AI/nemori](https://github.com/nemori-ai/nemori)
- Tags: deep-dive
- Published: 2026-03-08

---

**The Nemori Semantic Generator converts raw conversational episodes into high-value, persistent semantic memories by formatting dialogue context into specialized LLM prompts, extracting structured facts via prediction-correction or direct extraction modes, and wrapping results in SemanticMemory objects for long-term retrieval.**

The open-source nemori-ai/nemori project implements a sophisticated memory architecture where the Semantic Generator serves as the critical bridge between transient conversations and durable knowledge. This component analyzes dialogue episodes—whether single interactions or batched histories—to identify facts worth remembering long after a chat ends. Understanding how the Semantic Generator extracts knowledge reveals how Nemori builds its persistent understanding of users over time.

## Semantic Knowledge Extraction Pipeline

The extraction process follows a deterministic pipeline defined in [`src/generation/semantic_generator.py`](https://github.com/nemori-ai/nemori/blob/main/src/generation/semantic_generator.py), transforming unstructured text into queryable memory records through seven distinct stages.

### Entry Point and Mode Selection

The orchestration begins at `SemanticGenerator.check_and_generate_semantic_memories` (line 57), which inspects the `MemoryConfig` to decide between two extraction strategies. If `enable_prediction_correction` is **True** (the default), the system delegates to the `PredictionCorrectionEngine`; otherwise, it proceeds with direct per-episode or batch extraction based on the `extract_semantic_per_episode` flag.

### Episode Formatting for Context Windows

Before LLM processing, raw episodes require structured formatting to fit context windows. For batch operations, the generator uses `PromptTemplates.format_episodes_for_semantic` to compress multiple interactions into concise summaries. When processing single episodes with original message preservation, it calls `_format_episodes_from_original_messages` (line 34) to maintain the full conversational turn structure including timestamps and role metadata from the `original_messages` field.

### LLM Prompting and Structured Generation

The formatted text injects into `PromptTemplates.SEMANTIC_GENERATION_PROMPT` (defined in [`src/generation/prompts.py`](https://github.com/nemori-ai/nemori/blob/main/src/generation/prompts.py) at line 59), which explicitly instructs the LLM to output high-value, persistent facts as a JSON list of statements. The generator then invokes `LLMClient.generate_json_response` via the internal `_generate_semantic_with_llm` method (line 74), sending the prompt to the configured model (e.g., GPT-4) and receiving structured JSON payloads.

### Statement Extraction and Memory Construction

Depending on the active mode, the system handles LLM responses differently. In prediction-correction mode, `PredictionCorrectionEngine.learn_from_episode_simplified` executes a "predict → compare → extract" workflow to refine statements. In direct per-episode mode, the generator flattens top-level JSON keys (`user_profile`, `experience`, `knowledge`, `other`) into a unified statements list. Finally, `_convert_to_semantic_memories` (line 70) instantiates each statement as a `SemanticMemory` model, copying the earliest episode timestamp to preserve temporal ordering.

## Two Extraction Modes: Prediction-Correction vs. Per-Episode

Nemori supports dual extraction strategies controlled by configuration flags in [`src/config.py`](https://github.com/nemori-ai/nemori/blob/main/src/config.py), allowing developers to balance extraction depth against computational cost.

**Prediction-Correction Mode** (default when `enable_prediction_correction` is **True**) runs a two-step validation where the engine predicts user attributes based on existing memories, compares these predictions against actual episode content, and extracts only the discrepancies as new knowledge. This reduces noise by focusing strictly on novel information.

**Per-Episode Direct Extraction** activates when `extract_semantic_per_episode` is **True**, bypassing the prediction step to extract all candidate facts from individual episodes using a richer schema that categorizes memories into distinct semantic types (`user_profile`, `experience`, etc.).

## Configuration Flags That Control Extraction

Three boolean flags in `MemoryConfig` govern the Semantic Generator's behavior, as referenced in `SemanticGenerator.get_generator_stats`:

- **enable_prediction_correction**: Toggles the two-step prediction-correction workflow. Default: `True`.
- **extract_semantic_per_episode**: Enables direct single-episode extraction with multi-type categorization. Default: `False`.
- **enable_semantic_memory**: Master switch for the entire semantic memory subsystem. Default: `True`.

## Implementation Example: Generating Semantic Memories

The following Python example demonstrates initializing the generator and extracting memories from a new episode:

```python
from datetime import datetime
from nemori.src.utils import LLMClient, EmbeddingClient
from nemori.src.generation.semantic_generator import SemanticGenerator
from nemori.src.config import MemoryConfig
from nemori.src.models import Episode

# Initialise required components

llm = LLMClient(model="gpt-4o-mini")
emb = EmbeddingClient(model="text-embedding-3-large")
cfg = MemoryConfig()  # loads default flags

# Build the generator

generator = SemanticGenerator(
    llm_client=llm,
    embedding_client=emb,
    config=cfg
)

# Prepare a new episode

new_ep = Episode(
    user_id="user-123",
    title="Planning a weekend hike",
    content="I want to hike this weekend. Which trail are you thinking of?",
    original_messages=[
        {"role": "user", "content": "I want to hike this weekend.", "metadata": {"timestamp": "2024-03-01T10:12:00"}},
        {"role": "assistant", "content": "Which trail are you thinking of?", "metadata": {"timestamp": "2024-03-01T10:12:05"}}
    ],
    created_at=datetime.utcnow(),
    timestamp=datetime.utcnow()
)

# Extract semantic memories

semantic_memories = generator.check_and_generate_semantic_memories(
    user_id="user-123",
    new_episode=new_ep,
    existing_episodes=[],
    existing_semantic_memories=None
)

# Inspect results

for mem in semantic_memories:
    print(f"[{mem.knowledge_type}] {mem.content}")

```

Under the hood, because `extract_semantic_per_episode` defaults to `False`, this invokes `generate_semantic_memories`, formats the episode using original messages via `_format_episodes_from_original_messages`, and wraps LLM outputs in `SemanticMemory` objects via `_convert_to_semantic_memories`.

## Summary

- The Semantic Generator entry point at `check_and_generate_semantic_memories` in [`src/generation/semantic_generator.py`](https://github.com/nemori-ai/nemori/blob/main/src/generation/semantic_generator.py) routes between prediction-correction and direct extraction modes based on `MemoryConfig` flags.
- Episode formatting varies by batch size: `PromptTemplates.format_episodes_for_semantic` for batches, `_format_episodes_from_original_messages` for single-episode preservation.
- LLM calls use `SEMANTIC_GENERATION_PROMPT` from [`src/generation/prompts.py`](https://github.com/nemori-ai/nemori/blob/main/src/generation/prompts.py) to enforce JSON output of high-value facts, processed by `_generate_semantic_with_llm`.
- Extracted statements become `SemanticMemory` instances defined in [`src/models/semantic.py`](https://github.com/nemori-ai/nemori/blob/main/src/models/semantic.py), maintaining temporal metadata from source episodes.
- Core implementation files include [`src/generation/semantic_generator.py`](https://github.com/nemori-ai/nemori/blob/main/src/generation/semantic_generator.py), [`src/generation/prompts.py`](https://github.com/nemori-ai/nemori/blob/main/src/generation/prompts.py), [`src/generation/prediction_correction_engine.py`](https://github.com/nemori-ai/nemori/blob/main/src/generation/prediction_correction_engine.py), and [`src/config.py`](https://github.com/nemori-ai/nemori/blob/main/src/config.py).

## Frequently Asked Questions

### What is the difference between prediction-correction and per-episode extraction in Nemori?

Prediction-correction extraction (default) uses `PredictionCorrectionEngine.learn_from_episode_simplified` to validate new facts against existing knowledge before extraction, reducing redundancy. Per-episode extraction processes each conversation independently through direct LLM prompting without the prediction validation step, categorizing outputs into `user_profile`, `experience`, `knowledge`, and `other` buckets.

### How does the Semantic Generator format conversations for the LLM?

The generator selects formatting based on configuration. For batch processing, it compresses multiple episodes using `PromptTemplates.format_episodes_for_semantic`. For single episodes requiring full fidelity, `_format_episodes_from_original_messages` preserves original message roles, content, and timestamps from the `original_messages` field.

### Where are the semantic generation prompts defined in the Nemori codebase?

Prompt templates reside in [`src/generation/prompts.py`](https://github.com/nemori-ai/nemori/blob/main/src/generation/prompts.py). Specifically, `PromptTemplates.SEMANTIC_GENERATION_PROMPT` (line 59) contains the instructions that guide the LLM to extract high-value, persistent facts and return them as structured JSON statements.

### Can I disable the Semantic Generator while keeping other Nemori memory features active?

Yes. Set `enable_semantic_memory` to `False` in your `MemoryConfig`. This master switch disables the entire semantic memory subsystem—including the Semantic Generator—while preserving episodic memory and other non-semantic components.