# MTPLX Cache Snapshots Detach Modes: The Complete Guide to Four Materialization Options

> Explore the four MTPLX cache snapshot detach modes: eval_only, contiguous_eval, selected_slice_contiguous_eval, and metal_copy_leaf. Understand how each materializes KV-cache leaves for efficient storage.

- Repository: [Youssof Altoukhi/MTPLX](https://github.com/youssofal/MTPLX)
- Tags: deep-dive
- Published: 2026-09-05

---

**MTPLX supports four distinct detach modes for cache snapshots—`eval_only`, `contiguous_eval`, `selected_slice_contiguous_eval`, and `metal_copy_leaf`—which control how KV-cache leaves are materialized before being written back to storage.**

MTPLX is an open-source framework for efficient key-value (KV) cache management in machine learning workloads. The **detach mode** configuration determines how cache snapshots transition from lazy lineage evaluation to concrete tensor storage, directly impacting memory layout and computational overhead.

## The Four Supported Detach Modes

According to the `SUPPORTED_DETACH_MODES` constant defined in [`mtplx/cache_state.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/cache_state.py), MTPLX implements four specific strategies for cache leaf detachment:

### eval_only

The **`eval_only`** mode performs no physical detachment. The cache leaf remains in its original state and is used directly during evaluation without materializing a separate contiguous copy. This minimizes memory overhead but may sacrifice memory locality during subsequent access patterns.

### contiguous_eval

The **`contiguous_eval`** mode detaches the leaf into a **contiguous tensor** before storage. This defragments potentially strided memory layouts into dense blocks, improving cache locality during retrieval at the cost of an immediate memory copy operation.

### selected_slice_contiguous_eval

The **`selected_slice_contiguous_eval`** mode implements selective materialization. Only a **specific slice** of the cache leaf is detached into contiguous memory, allowing fine-grained control over which portions of the KV cache enter dense storage while preserving the original tensor structure for unselected regions.

### metal_copy_leaf

The **`metal_copy_leaf`** mode invokes a Metal-specific copy routine optimized for Apple Silicon environments. According to the source implementation, this mode falls back to standard contiguous evaluation on non-Metal backends, providing hardware-accelerated detachment when available.

## Implementation in the MTPLX Source Code

The detach mode validation and application logic resides in [`mtplx/cache_state.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/cache_state.py). The `_normalize_detach_mode` function validates user input against `SUPPORTED_DETACH_MODES` and raises a `ValueError` if an unsupported value is supplied.

When initializing a `TailOwnedKVCache` instance, the mode parameter undergoes normalization:

```python
from mtplx.cache_state import TailOwnedKVCache

# Create a cache with the default detach mode (contiguous_eval)

cache = TailOwnedKVCache(mode="contiguous_eval")

# Use Metal-optimized detachment for Apple Silicon

cache = TailOwnedKVCache(mode="metal_copy_leaf")

```

During the actual detachment operation, the `_own_tail` method invokes `detach_array_leaf` with the configured mode parameter, executing the specific materialization strategy selected during initialization.

### Validation Behavior

Attempting to instantiate `TailOwnedKVCache` with an invalid mode triggers explicit validation:

```python

# Invalid mode raises an exception

try:
    TailOwnedKVCache(mode="unsupported_mode")
except ValueError as e:
    print(e)   # detach mode must be one of [...]

```

This validation occurs within the `_normalize_detach_mode` helper function at lines 39-46 of [`mtplx/cache_state.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/cache_state.py).

## Summary

- **Four modes available**: `eval_only`, `contiguous_eval`, `selected_slice_contiguous_eval`, and `metal_copy_leaf`
- **Source location**: Defined in `SUPPORTED_DETACH_MODES` constant within [`mtplx/cache_state.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/cache_state.py)
- **Validation**: Enforced by `_normalize_detach_mode` with clear error messaging for unsupported values
- **Default behavior**: Typically defaults to `contiguous_eval` for balanced memory layout and performance
- **Hardware optimization**: `metal_copy_leaf` provides Metal-specific acceleration with automatic fallback

## Frequently Asked Questions

### What is the default detach mode in MTPLX?

While the source code allows explicit mode specification during `TailOwnedKVCache` initialization, the standard configuration defaults to `contiguous_eval`. This provides optimal memory locality for most inference workloads without requiring backend-specific optimizations.

### How does MTPLX handle invalid detach modes?

The `_normalize_detach_mode` function validates all mode strings against the `SUPPORTED_DETACH_MODES` tuple. If a user supplies an unsupported value, the function raises a `ValueError` with a descriptive message listing the valid options, preventing silent failures during cache operations.

### When should I use metal_copy_leaf versus contiguous_eval?

Use **`metal_copy_leaf`** when running on Apple Silicon (M1/M2/M3) devices where Metal Performance Shaders can accelerate memory operations. For CUDA, CPU, or other non-Metal environments, **`contiguous_eval`** provides identical functionality without backend detection overhead, as `metal_copy_leaf` automatically falls back to standard contiguous copying on incompatible hardware.

### Are detach modes backend-specific?

Only **`metal_copy_leaf`** contains backend-specific logic targeting Apple's Metal framework. The other three modes (`eval_only`, `contiguous_eval`, and `selected_slice_contiguous_eval`) operate agnostically across CPU, CUDA, and other accelerator types, making them portable across different deployment environments.