MTPLX Two-Tier Session Caching Architecture: Accelerating Multi-Turn Inference

MTPLX accelerates multi-turn LLM interactions by pairing an in-process warm-prefix SessionBank with a persistent SSD cold tier, enabling zero-copy KV-cache restoration across requests and process restarts.

The MTPLX inference engine optimizes multi-turn conversations through a sophisticated two-tier session caching architecture. By reusing previously computed KV-cache states instead of recomputing them from scratch, MTPLX dramatically reduces latency for repeated prompts. This system combines high-speed in-memory caching with durable disk storage to deliver both performance and persistence.

Understanding the Two-Tier Caching System

MTPLX implements a hierarchical caching strategy that separates hot, frequently accessed data from cold, persistent storage. This separation allows the engine to restore session state instantaneously while surviving process restarts.

Tier 1: Warm-Prefix SessionBank (In-Memory)

The SessionBank serves as the primary caching layer, maintaining exact token-prefix entries in process memory. Implemented in mtplx/session_bank.py, this component stores SessionBankEntry objects that capture KV-cache snapshots using the snapshot_cache and restore_cache methods.

When a request arrives, the engine checks if the prompt shares a prefix with an existing entry. On a match, the system restores the cached KV state directly, processing only the new suffix tokens. This restoration path is invoked from both mtplx/engine_session.py and mtplx/server/openai.py, utilizing policy-matching helpers like _policy_uses_committed_history and _restore_identity_compatible to validate compatibility. The behavior is governed by the MTPLX_SESSION_BOUNDARY_TRUE_RESTORE environment variable.

Tier 2: SSD Session Cache (Cold Storage)

The cold tier persists warm-prefix entries to disk, ensuring sessions survive process restarts. Located in mtplx/cache_bank/cold_tier.py, this layer serializes SessionBankEntry objects to ~/.mtplx/session_cache/ by default.

During server startup (mtplx start), the SSD cache is scanned and entries are reloaded into the in-memory SessionBank. If a request misses the warm-prefix bank, the system checks the SSD cache, enabling instant restoration of entire sessions. Users can disable this feature using the --ssd-session-cache off CLI flag. Shutdown flushes are logged from mtplx/server/openai.py with the message [mtplx] shutdown: SSD session cache flushed.

How the Two Tiers Interact

The caching system operates through a coordinated workflow:

  1. First Request – When no entry exists, the model performs a full pre-fill. Upon completion, the KV state is snapshotted and stored in the SessionBank. If enabled, the cold tier simultaneously writes the entry to disk.

  2. Subsequent Requests – The SessionBank is consulted first. On a prefix match, the cached KV state is restored—using zero-copy semantics when MTPLX_SESSION_LAZY_SNAPSHOT is enabled—and only suffix tokens are processed.

  3. Process Restart – Although the in-memory SessionBank is empty initially, the SSD cache is scanned at startup. Reloaded entries repopulate the SessionBank, enabling warm-prefix restores as if the process had never stopped.

Configuration and Environment Controls

MTPLX provides granular control over caching behavior through environment variables and CLI flags:

  • MTPLX_SESSION_LAZY_SNAPSHOT – Defaults to 1 (enabled), enabling zero-copy KV snapshots. Set to 0 to disable.
  • MTPLX_SESSION_SNAPSHOT_SETTLE – Defaults to 0 (disabled). When enabled, forces copy-on-write after each put to avoid COW stalls.
  • MTPLX_SESSION_NEAR_PREFIX_MAX_TOKEN_GAP – Defaults to 8, setting the token-gap tolerance for near-prefix restores.
  • MTPLX_SESSION_BOUNDARY_TRUE_RESTORE – Defaults to 1 (enabled), enforcing that restores respect recurrent boundaries.
  • --ssd-session-cache off – Disables the cold-tier SSD cache for the current run.

Practical Implementation Examples

Running MTPLX with default two-tier caching:


# Start the server – SSD cache is enabled by default

mtplx start

Disabling the SSD cache to use only the warm-prefix SessionBank:

mtplx start --ssd-session-cache off

Checking cache hit status programmatically:

import requests

resp = requests.post(
    "http://127.0.0.1:8000/v1/chat/completions",
    json={"model": "mtplx", "messages": [{"role": "user", "content": "Hello"}]},
    headers={"Content-Type": "application/json"},
)

cache_mode = resp.json()["usage"]["cache_mode"]

# Returns: "warm_prefix" (SessionBank hit), "ssd" (cold tier hit), or "cold" (miss)

Inspecting the on-disk SSD cache location:

ls ~/.mtplx/session_cache/

# Displays a hierarchy of serialized SessionBankEntry files

Summary

  • MTPLX uses a two-tier session caching architecture combining in-memory warm-prefix storage with SSD persistence.
  • The SessionBank (mtplx/session_bank.py) provides zero-copy KV-cache restoration for repeated prefixes using snapshot_cache and restore_cache.
  • The cold tier (mtplx/cache_bank/cold_tier.py) serializes entries to ~/.mtplx/session_cache/, enabling cross-restart persistence.
  • Integration occurs through mtplx/engine_session.py and mtplx/server/openai.py, with behavior controlled by environment variables like MTPLX_SESSION_LAZY_SNAPSHOT and MTPLX_SESSION_BOUNDARY_TRUE_RESTORE.

Frequently Asked Questions

What is the difference between the warm-prefix SessionBank and the SSD cache?

The warm-prefix SessionBank operates in process memory for instantaneous KV-cache restoration during active sessions, while the SSD cache persists these entries to disk in ~/.mtplx/session_cache/ to survive process restarts. The warm tier provides immediate access to recent computations, whereas the cold tier ensures durability across server restarts.

How does MTPLX handle cache misses?

When a request misses the in-memory SessionBank, MTPLX checks the SSD cache before falling back to a full model pre-fill. If the cold tier contains a matching SessionBankEntry, it restores the session to the warm bank instantly; otherwise, the model computes the full context and stores the result in both tiers for future requests.

Can I disable the SSD session cache in MTPLX?

Yes, pass the --ssd-session-cache off flag when starting the server: mtplx start --ssd-session-cache off. This forces MTPLX to rely solely on the in-memory SessionBank, meaning sessions will not persist across process restarts.

What controls the zero-copy behavior in MTPLX session caching?

The MTPLX_SESSION_LAZY_SNAPSHOT environment variable controls zero-copy semantics, defaulting to 1 (enabled). When active, the system uses zero-copy KV snapshots during restoration. Disabling this variable forces explicit memory copies, which may impact performance but can improve stability in memory-constrained environments.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →