# Understanding the Session Bank and SSD Cache Mechanism in MTPLX for Multi-Turn Chats

> Explore MTPLX's session bank and SSD cache mechanism that accelerates multi-turn chats using a two-tier KV-cache system for efficient data retrieval and storage.

- Repository: [Youssof Altoukhi/MTPLX](https://github.com/youssofal/MTPLX)
- Tags: internals
- Published: 2026-09-13

---

**MTPLX accelerates multi-turn conversations by storing reusable KV-cache snapshots in a two-tier system: an in-memory Session Bank for rapid exact-prefix retrieval and an SSD-backed Cold Tier for persistent overflow storage.**

MTPLX eliminates redundant computation during extended dialogues by caching the model’s internal state between turns. The **session bank and SSD cache mechanism** allows the framework to restore previous key-value (KV) tensors instead of recomputing them from scratch, drastically reducing latency for long conversational contexts. According to the `youssofal/MTPLX` source code, this architecture seamlessly blends high-speed RAM caching with durable disk persistence to support everything from brief exchanges to hours-long sessions.

## The Two-Tier Caching Architecture

MTPLX implements a hierarchical storage system that balances speed against capacity:

- **Session Bank** ([`mtplx/session_bank.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/session_bank.py)): An in-memory exact-prefix table storing snapshots of the model’s KV-cache, optional logits, hidden states, and GDN recurrent boundaries. This "warm tier" provides microsecond-level retrieval for active conversations.

- **SSD Cold Tier** ([`mtplx/cache_bank/cold_tier.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/cache_bank/cold_tier.py)): A persistent disk-based repository that offloads large snapshots when RAM budgets are exceeded. This "cold tier" uses content-addressed storage and streaming encoding to maintain performance without sacrificing conversation history.

## Session Bank: In-Memory Warm Cache

The Session Bank serves as the primary acceleration layer, maintaining live references to computed prompts for immediate reuse.

### Entry Structure and Metadata

Each cached turn is wrapped in a `SessionBankEntry` object (defined at lines 71–84 of [`session_bank.py`](https://github.com/youssofal/MTPLX/blob/main/session_bank.py)). These entries contain:

- `token_ids` and `token_hash`: The exact prompt prefix used for cache keys.
- `cache_snapshot`: A zero-copy view of the KV-cache tensors (`CacheSnapshot`).
- Optional `logits`, `hidden` states, and **GDN recurrent boundaries** (`gdn_boundaries`) for advanced model states.
- Metadata including `model_path`, `mtp_enabled` flags, `session_id`, and byte-size tracking (`nbytes`).

### Insertion and Size Management

When a conversation turn completes, `SessionBank.put()` (lines 375–440) captures the current KV-cache and creates a new entry. The method enforces strict memory governance:

```python
bank.put(
    runtime=runtime,
    token_ids=[101, 102, 103],
    cache=runtime.kv_cache,
    logits=runtime.logits,
    session_id="session-42",
    keep_live_ref=False,
)

```

The insertion process respects the `MTPLX_SESSION_BANK_PER_SESSION_MAX_BYTES` limit per session. For entries exceeding immediate encoding capacity, the system maintains a *live reference* (`keep_live_ref=True`) while asynchronously handing the payload to the cold tier.

### Prefix Matching Strategies

The bank employs three distinct retrieval strategies to maximize cache hits:

1. **`longest_shared_prefix_tokens()`**: Performs exact matching against stored `token_ids` sequences to find complete prefix reuse.
2. **`near_prefix_candidates()`** (line 560): Identifies entries where only a small suffix differs, allowing partial restoration plus minimal recomputation.
3. **`block_aligned_prefix_len()`**: Supports block-prefix restores for scenarios requiring specific alignment constraints, such as GDN boundary conditions.

### Eviction Policies and Memory Constraints

When the global RAM budget (`max_bytes`) or per-session limits are exceeded, `_evict_if_needed()` (around line 545) triggers intelligent cleanup:

- **Global Budget Enforcement**: Evicts oldest entries first until total memory consumption drops below threshold.
- **Protected-Terminal Policy**: When `MTPLX_SESSION_BANK_PROTECTED_TERMINAL` is enabled, the system preserves the newest extending entry even when fallback prompts would typically trigger eviction.
- **Boundary Shedding**: Optional `_shed_boundaries_to_fit` removes GDN recurrent boundaries from entries before evicting the KV-cache itself, preserving partial state.

### Lazy Snapshot Settlement

For environments with `MTPLX_SESSION_LAZY_SNAPSHOT` enabled, snapshots begin as zero-copy views of active GPU/CPU buffers. A background idle-lane thread later *settles* these snapshots via `_schedule_snapshot_settle()` (line 810), copying data into standalone buffers and marking `snapshot_settled_at` to free underlying model memory.

## SSD Cold Tier: Persistent Storage

When conversation histories grow beyond RAM capacity, the Cold Tier ensures persistence without performance degradation.

### Encoding and Write Queue Management

`SessionBankColdTier.put_entry()` (lines 780–850) handles serialization:

1. Encodes the `SessionBankEntry` using `TreeCodec.encode_payload()`.
2. Enqueues a `PendingWrite` to a background writer thread.
3. Writes content-addressed blobs to `~/.mtplx/session-bank/entries/<hash>/` using SHA-256 deduplication.

The writer thread operates asynchronously, ensuring model inference never blocks on disk I/O.

### Spill Path for Oversized Entries

For entries exceeding the writer-queue budget, `spill_entry()` (lines 1320–1380) streams tensors directly to disk without staging the complete payload in RAM. This *spill path* prevents memory pressure from massive snapshots (e.g., 100k+ token contexts) while maintaining data integrity.

### SSD Lookup and Restore Process

The cold tier mirrors the RAM bank’s retrieval logic with additional persistence guarantees:

- **`lookup()`** (lines 860–940): Retrieves exact-prefix matches from SQLite metadata.
- **`lookup_prefix_boundary()`** (lines 1070–1150): Implements near-prefix and block-prefix matching with *resident-duplicate* shadowing—if a RAM entry already covers the requested prefix, the SSD lookup aborts early to prioritize speed.

Restored entries are decoded and injected back into the Session Bank, becoming immediately available for subsequent turns.

### Write Budgeting and Safety Limits

To prevent SSD wear and pipeline stalls, the cold tier implements strict resource governance:

- **Global SSD Cap**: `MTPLX_SSD_MAX_BYTES` defaults to RAM-scaled values (see `default_cold_tier_max_bytes`).
- **Hourly Write Budget**: `MTPLX_SSD_WRITE_BUDGET_PER_HOUR` limits NAND flash wear.
- **Back-Pressure Handling**: `_admit_write()` and `_admit_hourly()` (lines 560–620) monitor `_backlog_budget_bytes`, while `MTPLX_SSD_WRITER_FOREGROUND_PAUSE` ensures the writer thread never chokes the decode pipeline.

## End-to-End Multi-Turn Chat Flow

The complete lifecycle demonstrates how both tiers cooperate:

1. **Turn N Completion**: `SessionBank.put()` stores the KV-cache snapshot in RAM. If the entry exceeds per-session limits, it maintains a live reference while the cold tier asynchronously persists the data.
2. **Turn N+1 Arrival**: The engine first queries the RAM bank via `restore()`. If no exact match exists, it queries `cold_tier.lookup_prefix_boundary()`.
3. **Cache Restoration**: The best candidate (exact, near, or block match) loads into the runtime’s KV-cache. The model then processes only the new token suffix rather than the full conversation history.
4. **Subsequent Access**: Once restored to RAM, the entry serves future turns at full speed until eviction returns it to cold storage.

## Implementation Examples

### Initializing the Session Bank

```python
from mtplx.session_bank import SessionBank
from mtplx.runtime import MTPLXRuntime

runtime = MTPLXRuntime(model_path="models/qwen4", mtp_enabled=False)

bank = SessionBank(
    max_entries=24,
    max_bytes=24 * 1024**3  # 24 GiB global budget

)

# After processing a turn

bank.put(
    runtime=runtime,
    token_ids=[101, 102, 103],
    cache=runtime.kv_cache,
    logits=runtime.logits,
    session_id="session-42",
)

```

### Enabling SSD Persistence

```python
from mtplx.cache_bank.cold_tier import SessionBankColdTier

cold = SessionBankColdTier(
    base_dir="~/.mtplx/session-bank",
    mode="on",
    max_bytes=32 * 1024**3,  # 32 GiB SSD budget

)

bank.cold_tier = cold

# Large entries automatically spill to SSD

bank.put(
    runtime=runtime,
    token_ids=long_token_list,
    cache=huge_cache,
    keep_live_ref=True,  # Triggers cold tier encoding

)

```

### Restoring from Exact Prefix

```python
prefix = [101, 102, 103]
restore = bank.restore(
    runtime=runtime,
    token_ids=prefix,
    model_path="models/qwen4",
    mtp_enabled=False,
)

if restore:
    # KV-cache prefilled; generate remaining tokens

    generated = runtime.generate(tokens=[104, 105])

```

## Summary

- **Two-tier architecture**: The Session Bank provides high-speed RAM caching while the SSD Cold Tier handles overflow and persistence.
- **Intelligent prefix matching**: Exact, near-prefix, and block-aligned restoration strategies maximize cache utilization across conversation variations.
- **Memory safety**: Per-session byte caps, global budgets, and protected-terminal policies prevent memory exhaustion while preserving critical context.
- **Lazy settlement**: Zero-copy snapshots with background settlement minimize immediate memory pressure during active inference.
- **SSD durability**: Hourly write budgets, spill paths for large entries, and back-pressure handling protect hardware longevity without sacrificing functionality.

## Frequently Asked Questions

### What is the difference between the Session Bank and the SSD Cold Tier?

The **Session Bank** operates exclusively in RAM, providing microsecond-latency retrieval of KV-cache snapshots for active or recently active conversations. The **SSD Cold Tier** persists snapshots to disk when they exceed RAM budgets or when `keep_live_ref=True` is specified, enabling restoration across server restarts and supporting conversations larger than available memory (controlled via `MTPLX_SSD_MAX_BYTES`).

### How does MTPLX handle cache eviction when memory limits are reached?

When storage exceeds `max_bytes` or `per_session_max_bytes`, the `_evict_if_needed()` method triggers a least-recently-used eviction sequence. If `MTPLX_SESSION_BANK_PROTECTED_TERMINAL` is enabled, the system exempts the newest extending entry from eviction to prevent critical context loss. Additionally, the system may shed GDN boundaries (`_shed_boundaries_to_fit`) to shrink entries before full eviction.

### What is lazy snapshot settlement and why does MTPLX use it?

**Lazy snapshot settlement** (enabled via `MTPLX_SESSION_LAZY_SNAPSHOT`) defers the physical copying of KV-cache tensors from GPU/CPU working buffers to standalone storage. Initially, entries hold zero-copy views; a background thread later settles these via `_schedule_snapshot_settle()` (line 810), marking `snapshot_settled_at` upon completion. This prevents immediate memory spikes during turn transitions while ensuring long-term storage efficiency.

### How does the system prevent excessive SSD wear during high-traffic scenarios?

The cold tier implements multiple protective mechanisms: `MTPLX_SSD_WRITE_BUDGET_PER_HOUR` caps hourly NAND writes; `_admit_write()` and `_admit_hourly()` (lines 560–620) enforce these limits through back-pressure; and `MTPLX_SSD_WRITER_FOREGROUND_PAUSE` ensures disk I/O never blocks the model’s decode pipeline. For oversized entries, the spill path (lines 1320–1380) streams data directly to disk without queuing, preventing buffer saturation.