# Nemori Caching Strategies: How the AI Memory System Optimizes Performance with Four Thread-Safe Layers

> Discover Nemori's four thread-safe caching strategies: per-user TTL, semantic-embedding, sharded LRU, and episode-storage. Optimize AI memory performance by reducing redundant computations and I/O.

- Repository: [Nemori AI/nemori](https://github.com/nemori-ai/nemori)
- Tags: deep-dive
- Published: 2026-03-08

---

**Nemori employs four complementary caching strategies—per-user TTL cache, semantic-embedding cache, sharded LRU cache, and episode-storage cache—to eliminate redundant disk I/O, embedding computation, and repeated function calls.**

The `nemori-ai/nemori` repository implements a sophisticated, multi-layered caching architecture designed to keep expensive operations fast and thread-safe. These caching strategies work in concert to minimize latency across user sessions, semantic memory retrieval, and episodic data access. Below is a detailed breakdown of each layer, its implementation, and how they integrate within the system.

## Four Caching Strategies in Nemori

### 1. Per-User TTL Cache

The **per-user TTL cache** stores arbitrary objects—such as loaded episodes and semantic memories—for a configurable duration. This prevents redundant data fetching when the same user interacts with the system repeatedly within a short window.

In [`src/services/cache.py`](https://github.com/nemori-ai/nemori/blob/main/src/services/cache.py), the `PerUserCache` class provides thread-safe access via a global lock. Each entry automatically expires after a time-to-live (TTL) period, defaulting to 600 seconds.

```python
from nemori.services.cache import PerUserCache

# Create a cache that expires entries after 10 minutes

user_cache = PerUserCache(ttl_seconds=600)

# Store a value for a specific user

user_cache.put(user_id="alice", value={"last_seen": "2026‑03‑08"})

# Retrieve it later (returns None if expired or missing)

session = user_cache.get("alice")

```

*Source:* [`src/services/cache.py`](https://github.com/nemori-ai/nemori/blob/main/src/services/cache.py) – class definition starts at line 12.

### 2. Semantic-Embedding Cache

Computing vector embeddings for semantic memories is computationally expensive. The **semantic-embedding cache** avoids recomputation by storing the mapping between `memory_id` and its pre-calculated embedding vector.

Implemented in [`src/services/cache.py`](https://github.com/nemori-ai/nemori/blob/main/src/services/cache.py) as `SemanticEmbeddingCache`, this layer uses a separate lock per user to minimize contention. The cache maintains an in-memory dictionary mapping `memory_id → embedding`.

```python
from nemori.services.cache import SemanticEmbeddingCache

embed_cache = SemanticEmbeddingCache()

# After computing an embedding for a memory

embed_cache.set(user_id="bob", memory_id="mem123", embedding=[0.1, 0.2, …])

# Later fetch without recomputation

cached = embed_cache.get(user_id="bob", memory_id="mem123")

```

*Source:* [`src/services/cache.py`](https://github.com/nemori-ai/nemori/blob/main/src/services/cache.py) – embedding cache definition lines 43‑66.

### 3. Sharded LRU Cache

For generic function memoization across the system—such as search results, loading flags, and semantic memory lookups—Nemori uses a **sharded LRU cache**. This design distributes entries across multiple shards, each with its own lock, to reduce lock contention in high-concurrency scenarios.

Located in [`src/utils/performance.py`](https://github.com/nemori-ai/nemori/blob/main/src/utils/performance.py), the `OptimizedLRUCache` (aliased as `LRUCache`) supports configurable TTL per entry and automatic LRU eviction when a shard exceeds its size limit. It also exposes statistics including size, hit-rate, and shard distribution.

```python
from nemori.utils.performance import PerformanceOptimizer

optimizer = PerformanceOptimizer(
    cache_size=2000,
    cache_ttl=1800,
    max_workers=8,
    num_cache_shards=32,
)

def heavy_computation(x, y):
    # some expensive work...

    return x ** y

# First call – runs the function and stores the result

result1 = optimizer.cached_call(heavy_computation, "heavy_computation", 2, 30)

# Second call – returns instantly from the sharded LRU cache

result2 = optimizer.cached_call(heavy_computation, "heavy_computation", 2, 30)

```

*Source:* [`src/utils/performance.py`](https://github.com/nemori-ai/nemori/blob/main/src/utils/performance.py) – `OptimizedLRUCache` (lines 29‑41) and `cached_call` (lines 61‑78).

### 4. Episode-Storage Internal Cache

The **episode-storage cache** operates at the persistence layer to reduce filesystem reads when loading episodic data for a user. This is particularly effective when the same user session repeatedly accesses recent episodes.

Implemented in [`src/storage/episode_storage.py`](https://github.com/nemori-ai/nemori/blob/main/src/storage/episode_storage.py), this cache uses a timestamp-based TTL of five minutes and an `RLock` to protect both the cache dictionary and its timestamps. It tracks cache hit and miss counters for performance monitoring.

```python

# Inside Nemori, you typically just call the storage API:

episodes = memory_system.storage["episode"].load(owner_id="alice")

# The storage layer will hit its internal cache if the data was loaded recently.

```

*Source:* [`src/storage/episode_storage.py`](https://github.com/nemori-ai/nemori/blob/main/src/storage/episode_storage.py) – TTL and lock setup lines 57‑66.

## How Nemori Caching Layers Work Together

The caching strategies are not isolated; they are orchestrated through the `MemorySystem` class in [`src/core/memory_system.py`](https://github.com/nemori-ai/nemori/blob/main/src/core/memory_system.py) to create a cohesive performance optimization pipeline.

1. **Initialization**: `MemorySystem` instantiates a `PerformanceOptimizer` with the sharded LRU cache, configuring it via `config.cache_size` and `config.cache_ttl_seconds` (lines 137‑140).

2. **Memoization of Heavy Operations**: The same optimizer memoizes expensive calls:
   * Data-loading flags (`user_data_loaded_*`) using `cache.contains` and `cache.put` (lines 288‑295, 346‑353).
   * Search results (`search_*`) (lines 1055‑1061).
   * Semantic-memory loading (`semantic_*`) (lines 1510‑1521).

3. **Per-User Storage**: `MemorySystem` holds `PerUserCache` instances for:
   * Semantic memories (`semantic_memory_cache`) (lines 177‑179).
   * Episodes (`episode_cache`) (lines 183‑185).

4. **Embedding Optimization**: `SemanticEmbeddingCache` is created once per `MemorySystem` and accessed during the generation pipeline. Its per-user lock ensures that concurrent threads never corrupt the same user’s embedding dictionary (lines 51‑56 in [`src/services/cache.py`](https://github.com/nemori-ai/nemori/blob/main/src/services/cache.py)).

## Configuration and Monitoring

Nemori exposes caching behavior through [`src/config.py`](https://github.com/nemori-ai/nemori/blob/main/src/config.py), allowing operators to tune performance without modifying code. Key configuration options include:

* `enable_cache` – Global toggle for caching functionality.
* `cache_size` – Maximum entries per sharded LRU cache shard.
* `cache_ttl_seconds` – Default expiration time for cached entries.
* `num_cache_shards` – Concurrency level for the LRU cache (16‑32 recommended for high-load scenarios).

The sharded LRU cache and episode storage cache both expose statistics (hit rates, size, shard distribution) enabling runtime performance analysis and bottleneck identification.

## Summary

Nemori’s caching architecture combines four specialized strategies to minimize latency and computational overhead:

* **Per-User TTL Cache** – Stores user-specific objects like episodes and memories with configurable expiration, protected by a global thread-safe lock.
* **Semantic-Embedding Cache** – Eliminates redundant vector embedding computation using per-user locks to minimize contention.
* **Sharded LRU Cache** – Provides high-concurrency memoization across 16‑32 shards, each with independent locking, TTL support, and LRU eviction.
* **Episode-Storage Cache** – Reduces filesystem I/O at the persistence layer with timestamp-based TTL and hit/miss tracking.

These layers are orchestrated through `MemorySystem` to create a thread-safe, observable, and highly configurable performance optimization pipeline.

## Frequently Asked Questions

### What is the default TTL for cached entries in Nemori?

The default time-to-live (TTL) for entries in the `PerUserCache` is **600 seconds** (10 minutes), as defined in [`src/services/cache.py`](https://github.com/nemori-ai/nemori/blob/main/src/services/cache.py). The episode-storage cache uses a shorter default of **5 minutes** (300 seconds) to balance freshness with I/O reduction. Both values are configurable via `config.cache_ttl_seconds`.

### How does Nemori prevent cache contention in multi-threaded environments?

Nemori employs **lock sharding** to minimize contention. The `OptimizedLRUCache` distributes entries across 16‑32 shards, each protected by its own lock. The `SemanticEmbeddingCache` uses a separate lock per user, ensuring that concurrent operations for different users never block each other. Global locks are used only for the `PerUserCache` where cross-user consistency is required.

### Can I disable caching or adjust cache sizes without modifying source code?

Yes. Nemori exposes caching controls through [`src/config.py`](https://github.com/nemori-ai/nemori/blob/main/src/config.py). You can set `enable_cache` to `False` to disable functionality globally, or tune `cache_size` (entries per shard), `cache_ttl_seconds`, and `num_cache_shards` to match your workload. These settings are consumed by `MemorySystem` during initialization in [`src/core/memory_system.py`](https://github.com/nemori-ai/nemori/blob/main/src/core/memory_system.py).

### What types of operations benefit most from Nemori’s caching strategies?

The **sharded LRU cache** accelerates expensive function calls like search result computation (`search_*`), semantic memory loading (`semantic_*`), and data-loading flag checks. The **semantic-embedding cache** specifically targets vector embedding generation, which is computationally intensive. The **episode-storage cache** optimizes filesystem reads for episodic data, while the **per-user TTL cache** speeds up retrieval of arbitrary user-bound objects like loaded conversation histories.