Nemori's Storage Backends and Data Formats: A Complete Technical Guide
Nemori uses two specialized storage backends—SemanticStorage and EpisodeStorage—that persist data as JSONL files with in-memory indexing and per-user locking for thread-safe concurrent access.
Nemori is an open-source memory system for AI applications that separates persistent data into semantic and episodic storage layers. Understanding Nemori's storage backends and data formats is essential for developers integrating the system or optimizing performance for high-concurrency scenarios.
Storage Architecture Overview
Nemori's persistence layer is built around an abstract base class that defines the contract for all storage operations. Both concrete implementations share common infrastructure for indexing, locking, and file I/O.
BaseStorage Interface
All storage backends inherit from BaseStorage defined in src/storage/base_storage.py. This abstract class establishes the public API:
save(item) → str– Persist an object and return its generated IDload(item_id) → Optional[Any]– Retrieve an object by its unique identifierdelete(item_id) → bool– Remove a specific object from storagelist_user_items(user_id) → List[Any]– Enumerate all objects belonging to a userdelete_user_data(user_id) → bool– Atomically wipe all data for a specific userget_user_stats(user_id) → Dict[str, Any]– Return usage statistics without loading full records
SemanticStorage Backend
SemanticStorage handles vector-like semantic memories used for deduplication and similarity search. It is implemented in src/storage/semantic_storage.py.
Data Format and File Structure
Semantic memories persist as JSON Lines (JSONL) files, with one file per user:
- File naming:
<user_id>_semantic.jsonl - Location:
{storage_path}/semantic/ - Format: Each line is an independent JSON object representing a
SemanticMemory
{"memory_id":"c123...","user_id":"u42","content":"Time: 2024-11-06 12:34:56 This is a test statement","created_at":"2024-11-06T12:34:56.789Z"}
{"memory_id":"d456...","user_id":"u42","content":"...","created_at":"2024-11-06T12:35:10.123Z"}
In-Memory Indexing
To avoid scanning entire files during lookups, SemanticStorage maintains two indexes:
self._memory_index: Dict[str, str]– Maps memory IDs to their file paths for O(1) location resolutionself._user_index: Dict[str, List[str]]– Maps user IDs to their list of memory IDs for fast enumeration
Concurrency Control
Thread safety is achieved through per-user reentrant locks (threading.RLock):
- Locks are stored in
self._user_file_lockswith user ID as the key - Each write operation acquires the specific user's lock before appending to their JSONL file
- This allows concurrent operations across different users while serializing access per user
EpisodeStorage Backend
EpisodeStorage manages episodic memories—complete conversation episodes with message lists. It is implemented in src/storage/episode_storage.py.
Data Formats and Migration
Episode storage supports two formats with automatic migration:
Primary Format: JSONL
- File naming:
<user_id>_episodes.jsonl - Location:
{storage_path}/episodes/ - Structure: Append-only lines of JSON episode objects
Legacy Format: JSON
- File naming:
<user_id>.json - Status: Deprecated but supported for backward compatibility
- Migration: Automatically converted to JSONL on first read via
_load_episodes_from_jsonand_load_episodes_from_jsonlmethods
Caching Strategy
To reduce disk I/O, EpisodeStorage implements a TTL-based cache:
- Cache duration defaults to 5 minutes
- Stores the most recent episode list per user
- Subsequent
list_user_itemscalls within the TTL window return cached data without file system access
Indexing and Concurrency
Similar to semantic storage, but with additional infrastructure:
Indexes:
self._episode_index: Dict[str, str]– Episode ID to file path mappingself._user_index: Dict[str, List[str]]– User ID to episode IDs list
Locking:
- Per-user
RLockinstances inself._user_file_locks - Global
self._locks_managerlock protects the lock dictionary itself during concurrent lock creation
Working with Nemori Storage
Instantiating Storage Backends
from nemori.src.storage.semantic_storage import SemanticStorage
from nemori.src.storage.episode_storage import EpisodeStorage
# Initialize with base storage path
semantic_store = SemanticStorage(storage_path="/data/nemori")
episode_store = EpisodeStorage(storage_path="/data/nemori")
Both classes automatically create required subdirectories (semantic/ and episodes/).
Saving Semantic Memories
from nemori.src.models.semantic import SemanticMemory
from datetime import datetime
mem = SemanticMemory(
memory_id="mem-001",
user_id="user-123",
content="Time: 2024-11-06 12:34:56 This is a test statement",
created_at=datetime.utcnow(),
)
mem_id = semantic_store.save(mem) # Returns: "mem-001"
The object is appended to user-123_semantic.jsonl and the in-memory index updates atomically.
Loading and Querying
# Load specific memory by ID
loaded = semantic_store.load("mem-001")
print(loaded.content)
# List all episodes for a user
episodes = episode_store.list_user_items("user-123")
print(f"{len(episodes)} episodes stored")
If episode data is cached (within 5-minute TTL), list_user_items returns cached results without disk I/O.
Deleting User Data
# Remove all data for a specific user
semantic_store.delete_user_data("user-123")
episode_store.delete_user_data("user-123")
Both calls remove the JSONL files, clear caches, and drop in-memory indexes for the specified user.
Summary
- Nemori's storage backends consist of
SemanticStorageandEpisodeStorage, both implementing theBaseStorageinterface defined insrc/storage/base_storage.py. - Data formats use JSONL (JSON Lines) as the primary persistence format, with legacy JSON support for backward compatibility in episode storage.
- Performance optimizations include O(1) in-memory indexes mapping IDs to file paths, append-only JSONL writes, and TTL-based caching (5-minute default) for episodic data.
- Concurrency is handled via per-user
threading.RLockinstances, allowing parallel operations across different users while ensuring atomic access per user file.
Frequently Asked Questions
What file formats does Nemori use for persistent storage?
Nemori primarily uses JSON Lines (JSONL) format, where each line represents a single memory object. Episode storage also supports legacy JSON array files for backward compatibility, automatically migrating them to JSONL on first read. Semantic storage exclusively uses JSONL via files named <user_id>_semantic.jsonl.
How does Nemori handle concurrent write operations?
Nemori implements per-user reentrant locks (threading.RLock) stored in self._user_file_locks. When a thread writes to a user's semantic or episode file, it acquires that specific user's lock, allowing concurrent operations across different users while serializing access per user. Episode storage adds a global self._locks_manager lock to protect lock dictionary creation.
What is the difference between SemanticStorage and EpisodeStorage?
SemanticStorage manages vector-like semantic memories for deduplication and similarity search, storing data in <user_id>_semantic.jsonl with simple ID-to-path indexing. EpisodeStorage handles full conversation episodes with message lists, uses both JSONL and legacy JSON formats, implements TTL-based caching (5-minute default), and manages more complex indexing for conversation retrieval.
How does Nemori optimize read performance for large datasets?
Nemori employs three optimization strategies: O(1) in-memory indexes that map memory/episode IDs directly to file paths without scanning; JSONL append-only format enabling streaming reads without loading entire files into memory; and TTL-based caching for episodic data that returns cached results for 5 minutes, eliminating repeated disk I/O for frequent queries.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →