Nemori Caching Strategies: How the AI Memory System Optimizes Performance with Four Thread-Safe Layers
Nemori employs four complementary caching strategies—per-user TTL cache, semantic-embedding cache, sharded LRU cache, and episode-storage cache—to eliminate redundant disk I/O, embedding computation, and repeated function calls.
The nemori-ai/nemori repository implements a sophisticated, multi-layered caching architecture designed to keep expensive operations fast and thread-safe. These caching strategies work in concert to minimize latency across user sessions, semantic memory retrieval, and episodic data access. Below is a detailed breakdown of each layer, its implementation, and how they integrate within the system.
Four Caching Strategies in Nemori
1. Per-User TTL Cache
The per-user TTL cache stores arbitrary objects—such as loaded episodes and semantic memories—for a configurable duration. This prevents redundant data fetching when the same user interacts with the system repeatedly within a short window.
In src/services/cache.py, the PerUserCache class provides thread-safe access via a global lock. Each entry automatically expires after a time-to-live (TTL) period, defaulting to 600 seconds.
from nemori.services.cache import PerUserCache
# Create a cache that expires entries after 10 minutes
user_cache = PerUserCache(ttl_seconds=600)
# Store a value for a specific user
user_cache.put(user_id="alice", value={"last_seen": "2026‑03‑08"})
# Retrieve it later (returns None if expired or missing)
session = user_cache.get("alice")
Source: src/services/cache.py – class definition starts at line 12.
2. Semantic-Embedding Cache
Computing vector embeddings for semantic memories is computationally expensive. The semantic-embedding cache avoids recomputation by storing the mapping between memory_id and its pre-calculated embedding vector.
Implemented in src/services/cache.py as SemanticEmbeddingCache, this layer uses a separate lock per user to minimize contention. The cache maintains an in-memory dictionary mapping memory_id → embedding.
from nemori.services.cache import SemanticEmbeddingCache
embed_cache = SemanticEmbeddingCache()
# After computing an embedding for a memory
embed_cache.set(user_id="bob", memory_id="mem123", embedding=[0.1, 0.2, …])
# Later fetch without recomputation
cached = embed_cache.get(user_id="bob", memory_id="mem123")
Source: src/services/cache.py – embedding cache definition lines 43‑66.
3. Sharded LRU Cache
For generic function memoization across the system—such as search results, loading flags, and semantic memory lookups—Nemori uses a sharded LRU cache. This design distributes entries across multiple shards, each with its own lock, to reduce lock contention in high-concurrency scenarios.
Located in src/utils/performance.py, the OptimizedLRUCache (aliased as LRUCache) supports configurable TTL per entry and automatic LRU eviction when a shard exceeds its size limit. It also exposes statistics including size, hit-rate, and shard distribution.
from nemori.utils.performance import PerformanceOptimizer
optimizer = PerformanceOptimizer(
cache_size=2000,
cache_ttl=1800,
max_workers=8,
num_cache_shards=32,
)
def heavy_computation(x, y):
# some expensive work...
return x ** y
# First call – runs the function and stores the result
result1 = optimizer.cached_call(heavy_computation, "heavy_computation", 2, 30)
# Second call – returns instantly from the sharded LRU cache
result2 = optimizer.cached_call(heavy_computation, "heavy_computation", 2, 30)
Source: src/utils/performance.py – OptimizedLRUCache (lines 29‑41) and cached_call (lines 61‑78).
4. Episode-Storage Internal Cache
The episode-storage cache operates at the persistence layer to reduce filesystem reads when loading episodic data for a user. This is particularly effective when the same user session repeatedly accesses recent episodes.
Implemented in src/storage/episode_storage.py, this cache uses a timestamp-based TTL of five minutes and an RLock to protect both the cache dictionary and its timestamps. It tracks cache hit and miss counters for performance monitoring.
# Inside Nemori, you typically just call the storage API:
episodes = memory_system.storage["episode"].load(owner_id="alice")
# The storage layer will hit its internal cache if the data was loaded recently.
Source: src/storage/episode_storage.py – TTL and lock setup lines 57‑66.
How Nemori Caching Layers Work Together
The caching strategies are not isolated; they are orchestrated through the MemorySystem class in src/core/memory_system.py to create a cohesive performance optimization pipeline.
-
Initialization:
MemorySysteminstantiates aPerformanceOptimizerwith the sharded LRU cache, configuring it viaconfig.cache_sizeandconfig.cache_ttl_seconds(lines 137‑140). -
Memoization of Heavy Operations: The same optimizer memoizes expensive calls:
- Data-loading flags (
user_data_loaded_*) usingcache.containsandcache.put(lines 288‑295, 346‑353). - Search results (
search_*) (lines 1055‑1061). - Semantic-memory loading (
semantic_*) (lines 1510‑1521).
- Data-loading flags (
-
Per-User Storage:
MemorySystemholdsPerUserCacheinstances for:- Semantic memories (
semantic_memory_cache) (lines 177‑179). - Episodes (
episode_cache) (lines 183‑185).
- Semantic memories (
-
Embedding Optimization:
SemanticEmbeddingCacheis created once perMemorySystemand accessed during the generation pipeline. Its per-user lock ensures that concurrent threads never corrupt the same user’s embedding dictionary (lines 51‑56 insrc/services/cache.py).
Configuration and Monitoring
Nemori exposes caching behavior through src/config.py, allowing operators to tune performance without modifying code. Key configuration options include:
enable_cache– Global toggle for caching functionality.cache_size– Maximum entries per sharded LRU cache shard.cache_ttl_seconds– Default expiration time for cached entries.num_cache_shards– Concurrency level for the LRU cache (16‑32 recommended for high-load scenarios).
The sharded LRU cache and episode storage cache both expose statistics (hit rates, size, shard distribution) enabling runtime performance analysis and bottleneck identification.
Summary
Nemori’s caching architecture combines four specialized strategies to minimize latency and computational overhead:
- Per-User TTL Cache – Stores user-specific objects like episodes and memories with configurable expiration, protected by a global thread-safe lock.
- Semantic-Embedding Cache – Eliminates redundant vector embedding computation using per-user locks to minimize contention.
- Sharded LRU Cache – Provides high-concurrency memoization across 16‑32 shards, each with independent locking, TTL support, and LRU eviction.
- Episode-Storage Cache – Reduces filesystem I/O at the persistence layer with timestamp-based TTL and hit/miss tracking.
These layers are orchestrated through MemorySystem to create a thread-safe, observable, and highly configurable performance optimization pipeline.
Frequently Asked Questions
What is the default TTL for cached entries in Nemori?
The default time-to-live (TTL) for entries in the PerUserCache is 600 seconds (10 minutes), as defined in src/services/cache.py. The episode-storage cache uses a shorter default of 5 minutes (300 seconds) to balance freshness with I/O reduction. Both values are configurable via config.cache_ttl_seconds.
How does Nemori prevent cache contention in multi-threaded environments?
Nemori employs lock sharding to minimize contention. The OptimizedLRUCache distributes entries across 16‑32 shards, each protected by its own lock. The SemanticEmbeddingCache uses a separate lock per user, ensuring that concurrent operations for different users never block each other. Global locks are used only for the PerUserCache where cross-user consistency is required.
Can I disable caching or adjust cache sizes without modifying source code?
Yes. Nemori exposes caching controls through src/config.py. You can set enable_cache to False to disable functionality globally, or tune cache_size (entries per shard), cache_ttl_seconds, and num_cache_shards to match your workload. These settings are consumed by MemorySystem during initialization in src/core/memory_system.py.
What types of operations benefit most from Nemori’s caching strategies?
The sharded LRU cache accelerates expensive function calls like search result computation (search_*), semantic memory loading (semantic_*), and data-loading flag checks. The semantic-embedding cache specifically targets vector embedding generation, which is computationally intensive. The episode-storage cache optimizes filesystem reads for episodic data, while the per-user TTL cache speeds up retrieval of arbitrary user-bound objects like loaded conversation histories.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →