# MTPLX Two-Tier Session Caching Architecture: Accelerating Multi-Turn Inference

> Explore MTPLX's two-tier session caching architecture. Accelerate multi-turn LLM inference with zero-copy KV-cache restoration using an in-process warm-prefix and persistent SSD cold tier.

- Repository: [Youssof Altoukhi/MTPLX](https://github.com/youssofal/MTPLX)
- Tags: architecture
- Published: 2026-09-05

---

**MTPLX accelerates multi-turn LLM interactions by pairing an in-process warm-prefix SessionBank with a persistent SSD cold tier, enabling zero-copy KV-cache restoration across requests and process restarts.**

The MTPLX inference engine optimizes multi-turn conversations through a sophisticated two-tier session caching architecture. By reusing previously computed KV-cache states instead of recomputing them from scratch, MTPLX dramatically reduces latency for repeated prompts. This system combines high-speed in-memory caching with durable disk storage to deliver both performance and persistence.

## Understanding the Two-Tier Caching System

MTPLX implements a hierarchical caching strategy that separates hot, frequently accessed data from cold, persistent storage. This separation allows the engine to restore session state instantaneously while surviving process restarts.

### Tier 1: Warm-Prefix SessionBank (In-Memory)

The **SessionBank** serves as the primary caching layer, maintaining exact token-prefix entries in process memory. Implemented in [`mtplx/session_bank.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/session_bank.py), this component stores `SessionBankEntry` objects that capture KV-cache snapshots using the `snapshot_cache` and `restore_cache` methods.

When a request arrives, the engine checks if the prompt shares a prefix with an existing entry. On a match, the system restores the cached KV state directly, processing only the new suffix tokens. This restoration path is invoked from both [`mtplx/engine_session.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/engine_session.py) and [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py), utilizing policy-matching helpers like `_policy_uses_committed_history` and `_restore_identity_compatible` to validate compatibility. The behavior is governed by the `MTPLX_SESSION_BOUNDARY_TRUE_RESTORE` environment variable.

### Tier 2: SSD Session Cache (Cold Storage)

The **cold tier** persists warm-prefix entries to disk, ensuring sessions survive process restarts. Located in [`mtplx/cache_bank/cold_tier.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/cache_bank/cold_tier.py), this layer serializes `SessionBankEntry` objects to `~/.mtplx/session_cache/` by default.

During server startup (`mtplx start`), the SSD cache is scanned and entries are reloaded into the in-memory SessionBank. If a request misses the warm-prefix bank, the system checks the SSD cache, enabling instant restoration of entire sessions. Users can disable this feature using the `--ssd-session-cache off` CLI flag. Shutdown flushes are logged from [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py) with the message `[mtplx] shutdown: SSD session cache flushed`.

## How the Two Tiers Interact

The caching system operates through a coordinated workflow:

1. **First Request** – When no entry exists, the model performs a full pre-fill. Upon completion, the KV state is snapshotted and stored in the SessionBank. If enabled, the cold tier simultaneously writes the entry to disk.

2. **Subsequent Requests** – The SessionBank is consulted first. On a prefix match, the cached KV state is restored—using zero-copy semantics when `MTPLX_SESSION_LAZY_SNAPSHOT` is enabled—and only suffix tokens are processed.

3. **Process Restart** – Although the in-memory SessionBank is empty initially, the SSD cache is scanned at startup. Reloaded entries repopulate the SessionBank, enabling warm-prefix restores as if the process had never stopped.

## Configuration and Environment Controls

MTPLX provides granular control over caching behavior through environment variables and CLI flags:

- **`MTPLX_SESSION_LAZY_SNAPSHOT`** – Defaults to `1` (enabled), enabling zero-copy KV snapshots. Set to `0` to disable.
- **`MTPLX_SESSION_SNAPSHOT_SETTLE`** – Defaults to `0` (disabled). When enabled, forces copy-on-write after each `put` to avoid COW stalls.
- **`MTPLX_SESSION_NEAR_PREFIX_MAX_TOKEN_GAP`** – Defaults to `8`, setting the token-gap tolerance for near-prefix restores.
- **`MTPLX_SESSION_BOUNDARY_TRUE_RESTORE`** – Defaults to `1` (enabled), enforcing that restores respect recurrent boundaries.
- **`--ssd-session-cache off`** – Disables the cold-tier SSD cache for the current run.

## Practical Implementation Examples

Running MTPLX with default two-tier caching:

```bash

# Start the server – SSD cache is enabled by default

mtplx start

```

Disabling the SSD cache to use only the warm-prefix SessionBank:

```bash
mtplx start --ssd-session-cache off

```

Checking cache hit status programmatically:

```python
import requests

resp = requests.post(
    "http://127.0.0.1:8000/v1/chat/completions",
    json={"model": "mtplx", "messages": [{"role": "user", "content": "Hello"}]},
    headers={"Content-Type": "application/json"},
)

cache_mode = resp.json()["usage"]["cache_mode"]

# Returns: "warm_prefix" (SessionBank hit), "ssd" (cold tier hit), or "cold" (miss)

```

Inspecting the on-disk SSD cache location:

```bash
ls ~/.mtplx/session_cache/

# Displays a hierarchy of serialized SessionBankEntry files

```

## Summary

- MTPLX uses a **two-tier session caching architecture** combining in-memory warm-prefix storage with SSD persistence.
- The **SessionBank** ([`mtplx/session_bank.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/session_bank.py)) provides zero-copy KV-cache restoration for repeated prefixes using `snapshot_cache` and `restore_cache`.
- The **cold tier** ([`mtplx/cache_bank/cold_tier.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/cache_bank/cold_tier.py)) serializes entries to `~/.mtplx/session_cache/`, enabling cross-restart persistence.
- Integration occurs through [`mtplx/engine_session.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/engine_session.py) and [`mtplx/server/openai.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/server/openai.py), with behavior controlled by environment variables like `MTPLX_SESSION_LAZY_SNAPSHOT` and `MTPLX_SESSION_BOUNDARY_TRUE_RESTORE`.

## Frequently Asked Questions

### What is the difference between the warm-prefix SessionBank and the SSD cache?

The warm-prefix SessionBank operates in process memory for instantaneous KV-cache restoration during active sessions, while the SSD cache persists these entries to disk in `~/.mtplx/session_cache/` to survive process restarts. The warm tier provides immediate access to recent computations, whereas the cold tier ensures durability across server restarts.

### How does MTPLX handle cache misses?

When a request misses the in-memory SessionBank, MTPLX checks the SSD cache before falling back to a full model pre-fill. If the cold tier contains a matching `SessionBankEntry`, it restores the session to the warm bank instantly; otherwise, the model computes the full context and stores the result in both tiers for future requests.

### Can I disable the SSD session cache in MTPLX?

Yes, pass the `--ssd-session-cache off` flag when starting the server: `mtplx start --ssd-session-cache off`. This forces MTPLX to rely solely on the in-memory SessionBank, meaning sessions will not persist across process restarts.

### What controls the zero-copy behavior in MTPLX session caching?

The `MTPLX_SESSION_LAZY_SNAPSHOT` environment variable controls zero-copy semantics, defaulting to `1` (enabled). When active, the system uses zero-copy KV snapshots during restoration. Disabling this variable forces explicit memory copies, which may impact performance but can improve stability in memory-constrained environments.