MTPLX Two-Tier Session Caching Architecture: Accelerating Multi-Turn Inference
MTPLX accelerates multi-turn LLM interactions by pairing an in-process warm-prefix SessionBank with a persistent SSD cold tier, enabling zero-copy KV-cache restoration across requests and process restarts.
The MTPLX inference engine optimizes multi-turn conversations through a sophisticated two-tier session caching architecture. By reusing previously computed KV-cache states instead of recomputing them from scratch, MTPLX dramatically reduces latency for repeated prompts. This system combines high-speed in-memory caching with durable disk storage to deliver both performance and persistence.
Understanding the Two-Tier Caching System
MTPLX implements a hierarchical caching strategy that separates hot, frequently accessed data from cold, persistent storage. This separation allows the engine to restore session state instantaneously while surviving process restarts.
Tier 1: Warm-Prefix SessionBank (In-Memory)
The SessionBank serves as the primary caching layer, maintaining exact token-prefix entries in process memory. Implemented in mtplx/session_bank.py, this component stores SessionBankEntry objects that capture KV-cache snapshots using the snapshot_cache and restore_cache methods.
When a request arrives, the engine checks if the prompt shares a prefix with an existing entry. On a match, the system restores the cached KV state directly, processing only the new suffix tokens. This restoration path is invoked from both mtplx/engine_session.py and mtplx/server/openai.py, utilizing policy-matching helpers like _policy_uses_committed_history and _restore_identity_compatible to validate compatibility. The behavior is governed by the MTPLX_SESSION_BOUNDARY_TRUE_RESTORE environment variable.
Tier 2: SSD Session Cache (Cold Storage)
The cold tier persists warm-prefix entries to disk, ensuring sessions survive process restarts. Located in mtplx/cache_bank/cold_tier.py, this layer serializes SessionBankEntry objects to ~/.mtplx/session_cache/ by default.
During server startup (mtplx start), the SSD cache is scanned and entries are reloaded into the in-memory SessionBank. If a request misses the warm-prefix bank, the system checks the SSD cache, enabling instant restoration of entire sessions. Users can disable this feature using the --ssd-session-cache off CLI flag. Shutdown flushes are logged from mtplx/server/openai.py with the message [mtplx] shutdown: SSD session cache flushed.
How the Two Tiers Interact
The caching system operates through a coordinated workflow:
-
First Request – When no entry exists, the model performs a full pre-fill. Upon completion, the KV state is snapshotted and stored in the SessionBank. If enabled, the cold tier simultaneously writes the entry to disk.
-
Subsequent Requests – The SessionBank is consulted first. On a prefix match, the cached KV state is restored—using zero-copy semantics when
MTPLX_SESSION_LAZY_SNAPSHOTis enabled—and only suffix tokens are processed. -
Process Restart – Although the in-memory SessionBank is empty initially, the SSD cache is scanned at startup. Reloaded entries repopulate the SessionBank, enabling warm-prefix restores as if the process had never stopped.
Configuration and Environment Controls
MTPLX provides granular control over caching behavior through environment variables and CLI flags:
MTPLX_SESSION_LAZY_SNAPSHOT– Defaults to1(enabled), enabling zero-copy KV snapshots. Set to0to disable.MTPLX_SESSION_SNAPSHOT_SETTLE– Defaults to0(disabled). When enabled, forces copy-on-write after eachputto avoid COW stalls.MTPLX_SESSION_NEAR_PREFIX_MAX_TOKEN_GAP– Defaults to8, setting the token-gap tolerance for near-prefix restores.MTPLX_SESSION_BOUNDARY_TRUE_RESTORE– Defaults to1(enabled), enforcing that restores respect recurrent boundaries.--ssd-session-cache off– Disables the cold-tier SSD cache for the current run.
Practical Implementation Examples
Running MTPLX with default two-tier caching:
# Start the server – SSD cache is enabled by default
mtplx start
Disabling the SSD cache to use only the warm-prefix SessionBank:
mtplx start --ssd-session-cache off
Checking cache hit status programmatically:
import requests
resp = requests.post(
"http://127.0.0.1:8000/v1/chat/completions",
json={"model": "mtplx", "messages": [{"role": "user", "content": "Hello"}]},
headers={"Content-Type": "application/json"},
)
cache_mode = resp.json()["usage"]["cache_mode"]
# Returns: "warm_prefix" (SessionBank hit), "ssd" (cold tier hit), or "cold" (miss)
Inspecting the on-disk SSD cache location:
ls ~/.mtplx/session_cache/
# Displays a hierarchy of serialized SessionBankEntry files
Summary
- MTPLX uses a two-tier session caching architecture combining in-memory warm-prefix storage with SSD persistence.
- The SessionBank (
mtplx/session_bank.py) provides zero-copy KV-cache restoration for repeated prefixes usingsnapshot_cacheandrestore_cache. - The cold tier (
mtplx/cache_bank/cold_tier.py) serializes entries to~/.mtplx/session_cache/, enabling cross-restart persistence. - Integration occurs through
mtplx/engine_session.pyandmtplx/server/openai.py, with behavior controlled by environment variables likeMTPLX_SESSION_LAZY_SNAPSHOTandMTPLX_SESSION_BOUNDARY_TRUE_RESTORE.
Frequently Asked Questions
What is the difference between the warm-prefix SessionBank and the SSD cache?
The warm-prefix SessionBank operates in process memory for instantaneous KV-cache restoration during active sessions, while the SSD cache persists these entries to disk in ~/.mtplx/session_cache/ to survive process restarts. The warm tier provides immediate access to recent computations, whereas the cold tier ensures durability across server restarts.
How does MTPLX handle cache misses?
When a request misses the in-memory SessionBank, MTPLX checks the SSD cache before falling back to a full model pre-fill. If the cold tier contains a matching SessionBankEntry, it restores the session to the warm bank instantly; otherwise, the model computes the full context and stores the result in both tiers for future requests.
Can I disable the SSD session cache in MTPLX?
Yes, pass the --ssd-session-cache off flag when starting the server: mtplx start --ssd-session-cache off. This forces MTPLX to rely solely on the in-memory SessionBank, meaning sessions will not persist across process restarts.
What controls the zero-copy behavior in MTPLX session caching?
The MTPLX_SESSION_LAZY_SNAPSHOT environment variable controls zero-copy semantics, defaulting to 1 (enabled). When active, the system uses zero-copy KV snapshots during restoration. Disabling this variable forces explicit memory copies, which may impact performance but can improve stability in memory-constrained environments.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →