Understanding the Session Bank and SSD Cache Mechanism in MTPLX for Multi-Turn Chats

MTPLX accelerates multi-turn conversations by storing reusable KV-cache snapshots in a two-tier system: an in-memory Session Bank for rapid exact-prefix retrieval and an SSD-backed Cold Tier for persistent overflow storage.

MTPLX eliminates redundant computation during extended dialogues by caching the model’s internal state between turns. The session bank and SSD cache mechanism allows the framework to restore previous key-value (KV) tensors instead of recomputing them from scratch, drastically reducing latency for long conversational contexts. According to the youssofal/MTPLX source code, this architecture seamlessly blends high-speed RAM caching with durable disk persistence to support everything from brief exchanges to hours-long sessions.

The Two-Tier Caching Architecture

MTPLX implements a hierarchical storage system that balances speed against capacity:

  • Session Bank (mtplx/session_bank.py): An in-memory exact-prefix table storing snapshots of the model’s KV-cache, optional logits, hidden states, and GDN recurrent boundaries. This "warm tier" provides microsecond-level retrieval for active conversations.

  • SSD Cold Tier (mtplx/cache_bank/cold_tier.py): A persistent disk-based repository that offloads large snapshots when RAM budgets are exceeded. This "cold tier" uses content-addressed storage and streaming encoding to maintain performance without sacrificing conversation history.

Session Bank: In-Memory Warm Cache

The Session Bank serves as the primary acceleration layer, maintaining live references to computed prompts for immediate reuse.

Entry Structure and Metadata

Each cached turn is wrapped in a SessionBankEntry object (defined at lines 71–84 of session_bank.py). These entries contain:

  • token_ids and token_hash: The exact prompt prefix used for cache keys.
  • cache_snapshot: A zero-copy view of the KV-cache tensors (CacheSnapshot).
  • Optional logits, hidden states, and GDN recurrent boundaries (gdn_boundaries) for advanced model states.
  • Metadata including model_path, mtp_enabled flags, session_id, and byte-size tracking (nbytes).

Insertion and Size Management

When a conversation turn completes, SessionBank.put() (lines 375–440) captures the current KV-cache and creates a new entry. The method enforces strict memory governance:

bank.put(
    runtime=runtime,
    token_ids=[101, 102, 103],
    cache=runtime.kv_cache,
    logits=runtime.logits,
    session_id="session-42",
    keep_live_ref=False,
)

The insertion process respects the MTPLX_SESSION_BANK_PER_SESSION_MAX_BYTES limit per session. For entries exceeding immediate encoding capacity, the system maintains a live reference (keep_live_ref=True) while asynchronously handing the payload to the cold tier.

Prefix Matching Strategies

The bank employs three distinct retrieval strategies to maximize cache hits:

  1. longest_shared_prefix_tokens(): Performs exact matching against stored token_ids sequences to find complete prefix reuse.
  2. near_prefix_candidates() (line 560): Identifies entries where only a small suffix differs, allowing partial restoration plus minimal recomputation.
  3. block_aligned_prefix_len(): Supports block-prefix restores for scenarios requiring specific alignment constraints, such as GDN boundary conditions.

Eviction Policies and Memory Constraints

When the global RAM budget (max_bytes) or per-session limits are exceeded, _evict_if_needed() (around line 545) triggers intelligent cleanup:

  • Global Budget Enforcement: Evicts oldest entries first until total memory consumption drops below threshold.
  • Protected-Terminal Policy: When MTPLX_SESSION_BANK_PROTECTED_TERMINAL is enabled, the system preserves the newest extending entry even when fallback prompts would typically trigger eviction.
  • Boundary Shedding: Optional _shed_boundaries_to_fit removes GDN recurrent boundaries from entries before evicting the KV-cache itself, preserving partial state.

Lazy Snapshot Settlement

For environments with MTPLX_SESSION_LAZY_SNAPSHOT enabled, snapshots begin as zero-copy views of active GPU/CPU buffers. A background idle-lane thread later settles these snapshots via _schedule_snapshot_settle() (line 810), copying data into standalone buffers and marking snapshot_settled_at to free underlying model memory.

SSD Cold Tier: Persistent Storage

When conversation histories grow beyond RAM capacity, the Cold Tier ensures persistence without performance degradation.

Encoding and Write Queue Management

SessionBankColdTier.put_entry() (lines 780–850) handles serialization:

  1. Encodes the SessionBankEntry using TreeCodec.encode_payload().
  2. Enqueues a PendingWrite to a background writer thread.
  3. Writes content-addressed blobs to ~/.mtplx/session-bank/entries/<hash>/ using SHA-256 deduplication.

The writer thread operates asynchronously, ensuring model inference never blocks on disk I/O.

Spill Path for Oversized Entries

For entries exceeding the writer-queue budget, spill_entry() (lines 1320–1380) streams tensors directly to disk without staging the complete payload in RAM. This spill path prevents memory pressure from massive snapshots (e.g., 100k+ token contexts) while maintaining data integrity.

SSD Lookup and Restore Process

The cold tier mirrors the RAM bank’s retrieval logic with additional persistence guarantees:

  • lookup() (lines 860–940): Retrieves exact-prefix matches from SQLite metadata.
  • lookup_prefix_boundary() (lines 1070–1150): Implements near-prefix and block-prefix matching with resident-duplicate shadowing—if a RAM entry already covers the requested prefix, the SSD lookup aborts early to prioritize speed.

Restored entries are decoded and injected back into the Session Bank, becoming immediately available for subsequent turns.

Write Budgeting and Safety Limits

To prevent SSD wear and pipeline stalls, the cold tier implements strict resource governance:

  • Global SSD Cap: MTPLX_SSD_MAX_BYTES defaults to RAM-scaled values (see default_cold_tier_max_bytes).
  • Hourly Write Budget: MTPLX_SSD_WRITE_BUDGET_PER_HOUR limits NAND flash wear.
  • Back-Pressure Handling: _admit_write() and _admit_hourly() (lines 560–620) monitor _backlog_budget_bytes, while MTPLX_SSD_WRITER_FOREGROUND_PAUSE ensures the writer thread never chokes the decode pipeline.

End-to-End Multi-Turn Chat Flow

The complete lifecycle demonstrates how both tiers cooperate:

  1. Turn N Completion: SessionBank.put() stores the KV-cache snapshot in RAM. If the entry exceeds per-session limits, it maintains a live reference while the cold tier asynchronously persists the data.
  2. Turn N+1 Arrival: The engine first queries the RAM bank via restore(). If no exact match exists, it queries cold_tier.lookup_prefix_boundary().
  3. Cache Restoration: The best candidate (exact, near, or block match) loads into the runtime’s KV-cache. The model then processes only the new token suffix rather than the full conversation history.
  4. Subsequent Access: Once restored to RAM, the entry serves future turns at full speed until eviction returns it to cold storage.

Implementation Examples

Initializing the Session Bank

from mtplx.session_bank import SessionBank
from mtplx.runtime import MTPLXRuntime

runtime = MTPLXRuntime(model_path="models/qwen4", mtp_enabled=False)

bank = SessionBank(
    max_entries=24,
    max_bytes=24 * 1024**3  # 24 GiB global budget

)

# After processing a turn

bank.put(
    runtime=runtime,
    token_ids=[101, 102, 103],
    cache=runtime.kv_cache,
    logits=runtime.logits,
    session_id="session-42",
)

Enabling SSD Persistence

from mtplx.cache_bank.cold_tier import SessionBankColdTier

cold = SessionBankColdTier(
    base_dir="~/.mtplx/session-bank",
    mode="on",
    max_bytes=32 * 1024**3,  # 32 GiB SSD budget

)

bank.cold_tier = cold

# Large entries automatically spill to SSD

bank.put(
    runtime=runtime,
    token_ids=long_token_list,
    cache=huge_cache,
    keep_live_ref=True,  # Triggers cold tier encoding

)

Restoring from Exact Prefix

prefix = [101, 102, 103]
restore = bank.restore(
    runtime=runtime,
    token_ids=prefix,
    model_path="models/qwen4",
    mtp_enabled=False,
)

if restore:
    # KV-cache prefilled; generate remaining tokens

    generated = runtime.generate(tokens=[104, 105])

Summary

  • Two-tier architecture: The Session Bank provides high-speed RAM caching while the SSD Cold Tier handles overflow and persistence.
  • Intelligent prefix matching: Exact, near-prefix, and block-aligned restoration strategies maximize cache utilization across conversation variations.
  • Memory safety: Per-session byte caps, global budgets, and protected-terminal policies prevent memory exhaustion while preserving critical context.
  • Lazy settlement: Zero-copy snapshots with background settlement minimize immediate memory pressure during active inference.
  • SSD durability: Hourly write budgets, spill paths for large entries, and back-pressure handling protect hardware longevity without sacrificing functionality.

Frequently Asked Questions

What is the difference between the Session Bank and the SSD Cold Tier?

The Session Bank operates exclusively in RAM, providing microsecond-latency retrieval of KV-cache snapshots for active or recently active conversations. The SSD Cold Tier persists snapshots to disk when they exceed RAM budgets or when keep_live_ref=True is specified, enabling restoration across server restarts and supporting conversations larger than available memory (controlled via MTPLX_SSD_MAX_BYTES).

How does MTPLX handle cache eviction when memory limits are reached?

When storage exceeds max_bytes or per_session_max_bytes, the _evict_if_needed() method triggers a least-recently-used eviction sequence. If MTPLX_SESSION_BANK_PROTECTED_TERMINAL is enabled, the system exempts the newest extending entry from eviction to prevent critical context loss. Additionally, the system may shed GDN boundaries (_shed_boundaries_to_fit) to shrink entries before full eviction.

What is lazy snapshot settlement and why does MTPLX use it?

Lazy snapshot settlement (enabled via MTPLX_SESSION_LAZY_SNAPSHOT) defers the physical copying of KV-cache tensors from GPU/CPU working buffers to standalone storage. Initially, entries hold zero-copy views; a background thread later settles these via _schedule_snapshot_settle() (line 810), marking snapshot_settled_at upon completion. This prevents immediate memory spikes during turn transitions while ensuring long-term storage efficiency.

How does the system prevent excessive SSD wear during high-traffic scenarios?

The cold tier implements multiple protective mechanisms: MTPLX_SSD_WRITE_BUDGET_PER_HOUR caps hourly NAND writes; _admit_write() and _admit_hourly() (lines 560–620) enforce these limits through back-pressure; and MTPLX_SSD_WRITER_FOREGROUND_PAUSE ensures disk I/O never blocks the model’s decode pipeline. For oversized entries, the spill path (lines 1320–1380) streams data directly to disk without queuing, preventing buffer saturation.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →