How MTPLX Implements Persistent Session Restoration Across Restarts
MTPLX achieves persistent session restoration by maintaining a durable session-bank that stores KV-cache snapshots in RAM with optional spillover to SSD, enabling sub-millisecond warm-prefix restores after server restarts.
The MTPLX inference engine (youssofal/MTPLX) eliminates cold-start latency for long-running conversations through a sophisticated persistent session restoration mechanism. By serializing cache states to disk and reconstructing them on startup, MTPLX allows conversations to resume instantly without re-prefilling entire prompts.
Session Bank Architecture and Snapshot Creation
The core of MTPLX’s durability lies in the SessionBank class defined in mtplx/session_bank.py. When a request finishes, the bank captures the runtime state—including the KV cache, logits, and hidden states—into a snapshot structure. This snapshot serves as the foundation for restoration after a restart.
Hybrid Snapshot Strategies
MTPLX implements two snapshot modes to balance memory efficiency against copy overhead, controlled by the MTPLX_SESSION_LAZY_SNAPSHOT environment variable (lines 88‑92):
- Live-reference lease: Stores only a pointer to the live cache (zero bytes copied) for very large prefixes.
- Zero-copy snapshot: Creates a full copy via
snapshot_cache_lazy_hybrid(the default, lines 84‑86) orsnapshot_cachewhen lazy snapshots are disabled.
Because the session-bank retains these snapshots in RAM, subsequent requests with matching prefixes achieve instant warm restores. However, to survive process restarts, MTPLX must persist these entries beyond volatile memory.
Cold Tier Persistence for Cross-Process Survival
To ensure persistent session restoration across restarts, MTPLX optionally offloads snapshot entries to a cold-tier SSD (self.cold_tier). This tier stores CacheSnapshot objects on disk alongside GDN boundary records that enable sub-prefix restoration logic.
Configure the cold tier before starting the server:
export MTPLX_SESSION_BANK_COLD_TIER_PATH=/var/mtplx/session-bank
Background Spill to SSD
The offloading process runs asynchronously via the idle-lane background thread (self.cold_enqueue_dispatch, lines 98‑100). When RAM pressure rises or entries age out, this thread serializes cold entries to the SSD path configured above. Each entry contains the full cache snapshot and boundary metadata necessary for reconstruction.
Restoration Workflow After Restart
When MTPLX starts, the SessionBank constructor automatically re-indexes persisted state. As implemented in mtplx/session_bank.py (lines 91‑94), the constructor reads the SSD directory (if configured) and populates its in-memory index with the persisted entries, making them available for immediate restoration.
Exact and Near-Prefix Matching
Incoming requests trigger the restoration pipeline via SessionBank.restore():
- Exact match: The method searches
self._entriesfor an identical token sequence. - Fallback candidates: If no exact match exists, the system queries
near_prefix_candidatesandblock_prefix_candidatesto find reusable partial caches.
If a candidate resides on the SSD tier, the system invokes gdn_boundary_loader (lines 132‑140) to lazily read the snapshot and any boundary records. The loader then calls restore_cache to install zero-copy views of the saved KV cache into the new MTPLXRuntime instance.
Lazy Loading from SSD
The restoration process leverages boundary-aware restore logic (recurrent_boundary_at_or_below and restore_cache, lines 60‑71) to handle sub-prefix scenarios. When gdn_boundary_loader retrieves a cold-tier entry, it reconstructs the cache hierarchy without requiring the model to re-prefill the prompt, dramatically reducing time-to-first-token after a restart.
Configuration Options
Developers control persistence behavior through environment variables:
| Variable | Effect |
|---|---|
MTPLX_SESSION_LAZY_SNAPSHOT |
Toggles between snapshot_cache_lazy_hybrid (default) and eager snapshot_cache modes (lines 88‑92). |
MTPLX_SESSION_BANK_COLD_TIER_PATH |
Enables SSD persistence by specifying the spillover directory for the cold tier. |
In mtplx/server/openai.py, the request handler integrates these components by invoking bank.restore() with parameters including session_id, model_path, and mtp_enabled, then installing the returned cache directly into the runtime:
# From mtplx/server/openai.py request handling logic
bank = state.sessions.bank
restore = bank.restore(
token_ids=new_prompt_ids,
session_id=session_id,
model_path=runtime.model_path,
mtp_enabled=runtime.mtp_enabled,
# ... additional matching criteria
)
if restore:
runtime.cache = restore.cache
runtime.logits = restore.logits
runtime.hidden = restore.hidden
Summary
- Session-bank durability: MTPLX stores KV-cache snapshots in a
SessionBankthat survives restarts via optional SSD spillover. - Hybrid storage: Snapshots use either live-reference leases or zero-copy copies via
snapshot_cache_lazy_hybridto optimize memory usage. - Automatic re-indexing: On startup,
SessionBank.__init__(lines 91‑94) reloads SSD entries, making them immediately available for restoration. - Intelligent matching: The restore pipeline tries exact matches first, then falls back to near-prefix or block-prefix candidates.
- Lazy reconstruction: The
gdn_boundary_loaderreads cold-tier data on-demand, usingrestore_cacheto rebuild runtime state without re-prefilling.
Frequently Asked Questions
How does MTPLX ensure session data survives a complete server restart?
MTPLX writes snapshot entries to a cold-tier SSD directory configured via MTPLX_SESSION_BANK_COLD_TIER_PATH. The SessionBank constructor re-indexes these files on startup (lines 91‑94), while gdn_boundary_loader (lines 132‑140) handles lazy deserialization, ensuring cached states persist beyond process lifetime.
What is the performance impact of enabling SSD persistence?
The background spill to SSD runs on an idle-lane thread (self.cold_enqueue_dispatch) to minimize interference with active inference. Restoration from SSD incurs a one-time load penalty via gdn_boundary_loader, but eliminates the far larger cost of re-prefilling long prompts, resulting in net latency reduction for multi-turn conversations.
Can MTPLX restore partial conversations or only complete prefixes?
MTPLX supports sub-prefix restoration through GDN boundary records and recurrent_boundary_at_or_below logic (lines 60‑71). When SessionBank.restore() finds a candidate via near_prefix_candidates or block_prefix_candidates, it can reconstruct the cache from any valid boundary point, not just exact matches.
How do I disable lazy snapshots for immediate consistency?
Set the environment variable MTPLX_SESSION_LAZY_SNAPSHOT to false (lines 88‑92). This forces the system to use snapshot_cache instead of snapshot_cache_lazy_hybrid, creating full copies immediately rather than deferred zero-copy views, trading memory for deterministic persistence timing.
Have a question about this repo?
These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:
curl -s "https://instagit.com/install.md" Maintain an open-source project? Get it listed too →