How the Nemori Episode Generator Creates Episodic Memories: A Complete Technical Guide
The Nemori Episode Generator converts conversation buffers into structured episodic memories by formatting raw messages into LLM prompts, synthesizing titles and summaries via an LLM client, validating and sanitizing the output, resolving timestamps from either LLM extraction or message metadata, and encapsulating the result in an immutable Episode dataclass with full error fallback support.
The nemori-ai/nemori repository implements an intelligent memory architecture that compresses conversational histories into searchable, timestamped knowledge units. At the heart of this system lies the Episode Generator, a pipeline defined in src/generation/episode_generator.py that transforms raw Message objects into structured Episode dataclasses through a 10-step synthesis and validation process.
Message Collection and Prompt Engineering
The pipeline begins by collecting raw Message objects—containing role, content, and timestamp—from the user’s conversation buffer. According to the source code in src/generation/prompts.py, the PromptTemplates.format_conversation method (lines 6-28) converts this list into a single formatted string that preserves temporal context by including timestamps when available.
This formatted dialogue is then injected into the EPISODE_GENERATION_PROMPT template (lines 11-25 in the same file), which combines the conversation history with the boundary-detection reason—explaining why this particular segment was isolated from the broader chat—to guide the LLM’s synthesis.
LLM Synthesis and Response Validation
With the prompt constructed, EpisodeGenerator._generate_episode_with_llm (implemented in src/generation/episode_generator.py lines 92-121) dispatches the request to the configured LLMClient. The system expects a JSON response containing at minimum a title and content field, with an optional timestamp.
Upon receiving the LLM output, the generator performs strict validation and sanitization (lines 124-158). Title and content strings are stripped of extraneous whitespace, default values are applied to missing mandatory fields, and any timestamp string present is parsed into a Python datetime object to ensure type consistency across the storage layer.
Timestamp Resolution and Episode Instantiation
Timestamp assignment follows a hierarchical fallback strategy defined in lines 186-199 of episode_generator.py:
- Primary: Use the timestamp extracted from the LLM response if it parses as a valid
datetime. - Secondary: Fall back to the earliest message timestamp within the conversation segment, normalized to naïve UTC.
- Tertiary: Use the current system time if no message timestamps exist.
Once resolved, the generator instantiates an immutable Episode dataclass (lines 74-82), encapsulating the title, synthesized content, references to original messages, message count, boundary reason, and the resolved timestamp.
Error Handling and Batch Processing
Robustness is ensured through _create_fallback_episode (referenced at lines 88-90 and implemented at lines 108-166), which triggers automatically when LLM calls raise exceptions. This method constructs a minimal but well-formed episode using raw message counts and available timestamps, ensuring the pipeline never returns malformed data.
For production workloads, batch_generate_episodes (lines 82-115) provides an optimized loop over pre-segmented message batches. The method invokes generate_episode for each segment while applying individual fallback handling per batch, preventing a single failure from corrupting an entire conversation history.
The module also exposes validate_episode (lines 166-196) for post-hoc consistency checks and get_generator_stats (lines 198-207) to report configuration state and model metadata for observability.
Practical Implementation: Code Examples
Generating a Single Episodic Memory
The following example demonstrates the complete workflow from message preparation to episode generation:
from nemori.src.utils.llm_client import LLMClient
from nemori.src.config import MemoryConfig
from nemori.src.generation.episode_generator import EpisodeGenerator
from nemori.src.models.message import Message
from datetime import datetime
# Initialize configuration and LLM client
config = MemoryConfig(
llm_model="gpt-4o-mini",
episode_min_messages=5,
episode_max_messages=50,
)
llm = LLMClient(model_name=config.llm_model)
# Build message list from conversation buffer
messages = [
Message(role="user", content="Hey, I want to plan a hike next weekend.",
timestamp=datetime(2024, 3, 14, 15, 0)),
Message(role="assistant", content="Great! Where would you like to go?",
timestamp=datetime(2024, 3, 14, 15, 1)),
Message(role="user", content="Mount Rainier, aiming for sunrise.",
timestamp=datetime(2024, 3, 14, 15, 2)),
]
# Create generator and produce episode
gen = EpisodeGenerator(llm_client=llm, config=config)
episode = gen.generate_episode(
user_id="user-123",
messages=messages,
boundary_reason="topic shift to outdoor planning"
)
print(f"Title: {episode.title}")
print(f"Timestamp: {episode.timestamp}")
print(f"Content: {episode.content[:200]}...")
Processing Multiple Segments in Batch
For segmented conversations, use the batch interface:
# segments is a list of Message lists from the BatchSegmenter
# reasons aligns with each segment's boundary detection
episodes = gen.batch_generate_episodes(
user_id="user-123",
message_batches=segments,
boundary_reasons=reasons,
)
print(f"Successfully generated {len(episodes)} episodic memories.")
Core Source Files and Architecture
| File | Purpose | Key Components |
|---|---|---|
src/models/message.py |
Defines input data structure | Message dataclass (role, content, timestamp) |
src/models/episode.py |
Defines output data structure | Episode dataclass with metadata |
src/generation/prompts.py |
Template management | PromptTemplates.format_conversation, EPISODE_GENERATION_PROMPT |
src/utils/llm_client.py |
LLM abstraction layer | LLMClient.generate_json_response |
src/generation/episode_generator.py |
Core orchestration logic | EpisodeGenerator.generate_episode, _generate_episode_with_llm, _create_fallback_episode |
src/config.py |
Configuration management | MemoryConfig dataclass |
Summary
- The Episode Generator follows a strict 10-step pipeline from raw messages to validated episodic memories.
- Prompt engineering in
src/generation/prompts.pystructures conversations with temporal context and boundary reasons. - LLM synthesis occurs through
_generate_episode_with_llm, expecting JSON outputs withtitle,content, and optionaltimestamp. - Validation logic sanitizes whitespace, applies defaults, and parses temporal data before instantiation.
- Timestamp resolution prioritizes LLM-extracted dates, falls back to earliest message time (naïve UTC), then system time.
- Fault tolerance is guaranteed by
_create_fallback_episode, ensuring every message batch produces a validEpisodeeven during LLM failures. - Batch processing via
batch_generate_episodeshandles segmented conversations with per-segment error isolation.
Frequently Asked Questions
What JSON schema does the LLM need to return for episode generation?
The LLM must return a JSON object containing at minimum a title string and content string. An optional timestamp field may be included as a string; if present, the Episode Generator parses it into a datetime object. This contract is enforced in EpisodeGenerator._generate_episode_with_llm (lines 92-121) and the subsequent validation block (lines 124-158).
How does Nemori handle generation failures or malformed LLM output?
Any exception during the LLM call or validation phase triggers _create_fallback_episode (lines 88-90, 108-166). This method constructs a minimal but valid Episode using the raw message count and available timestamps, ensuring the pipeline never emits null or corrupted memory objects even when the LLM returns invalid JSON or times out.
Which timestamp is assigned when the LLM omits the timestamp field?
If the LLM response lacks a timestamp or contains an unparsable value, the generator selects the earliest timestamp from the input Message objects, normalizing it to naïve UTC (lines 186-199). If the message list contains no timestamps, the system falls back to the current system time as the final default.
Does the Episode Generator support batch processing of conversation segments?
Yes. The batch_generate_episodes method (lines 82-115) accepts a list of pre-segmented message batches and corresponding boundary reasons. It iterates through each batch, calling generate_episode individually and applying fallback handling per segment. This allows efficient bulk creation of episodic memories from long conversations split by topic detection algorithms.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →