MTPLX Cache Snapshots Detach Modes: The Complete Guide to Four Materialization Options

MTPLX supports four distinct detach modes for cache snapshots—eval_only, contiguous_eval, selected_slice_contiguous_eval, and metal_copy_leaf—which control how KV-cache leaves are materialized before being written back to storage.

MTPLX is an open-source framework for efficient key-value (KV) cache management in machine learning workloads. The detach mode configuration determines how cache snapshots transition from lazy lineage evaluation to concrete tensor storage, directly impacting memory layout and computational overhead.

The Four Supported Detach Modes

According to the SUPPORTED_DETACH_MODES constant defined in mtplx/cache_state.py, MTPLX implements four specific strategies for cache leaf detachment:

eval_only

The eval_only mode performs no physical detachment. The cache leaf remains in its original state and is used directly during evaluation without materializing a separate contiguous copy. This minimizes memory overhead but may sacrifice memory locality during subsequent access patterns.

contiguous_eval

The contiguous_eval mode detaches the leaf into a contiguous tensor before storage. This defragments potentially strided memory layouts into dense blocks, improving cache locality during retrieval at the cost of an immediate memory copy operation.

selected_slice_contiguous_eval

The selected_slice_contiguous_eval mode implements selective materialization. Only a specific slice of the cache leaf is detached into contiguous memory, allowing fine-grained control over which portions of the KV cache enter dense storage while preserving the original tensor structure for unselected regions.

metal_copy_leaf

The metal_copy_leaf mode invokes a Metal-specific copy routine optimized for Apple Silicon environments. According to the source implementation, this mode falls back to standard contiguous evaluation on non-Metal backends, providing hardware-accelerated detachment when available.

Implementation in the MTPLX Source Code

The detach mode validation and application logic resides in mtplx/cache_state.py. The _normalize_detach_mode function validates user input against SUPPORTED_DETACH_MODES and raises a ValueError if an unsupported value is supplied.

When initializing a TailOwnedKVCache instance, the mode parameter undergoes normalization:

from mtplx.cache_state import TailOwnedKVCache

# Create a cache with the default detach mode (contiguous_eval)

cache = TailOwnedKVCache(mode="contiguous_eval")

# Use Metal-optimized detachment for Apple Silicon

cache = TailOwnedKVCache(mode="metal_copy_leaf")

During the actual detachment operation, the _own_tail method invokes detach_array_leaf with the configured mode parameter, executing the specific materialization strategy selected during initialization.

Validation Behavior

Attempting to instantiate TailOwnedKVCache with an invalid mode triggers explicit validation:


# Invalid mode raises an exception

try:
    TailOwnedKVCache(mode="unsupported_mode")
except ValueError as e:
    print(e)   # detach mode must be one of [...]

This validation occurs within the _normalize_detach_mode helper function at lines 39-46 of mtplx/cache_state.py.

Summary

  • Four modes available: eval_only, contiguous_eval, selected_slice_contiguous_eval, and metal_copy_leaf
  • Source location: Defined in SUPPORTED_DETACH_MODES constant within mtplx/cache_state.py
  • Validation: Enforced by _normalize_detach_mode with clear error messaging for unsupported values
  • Default behavior: Typically defaults to contiguous_eval for balanced memory layout and performance
  • Hardware optimization: metal_copy_leaf provides Metal-specific acceleration with automatic fallback

Frequently Asked Questions

What is the default detach mode in MTPLX?

While the source code allows explicit mode specification during TailOwnedKVCache initialization, the standard configuration defaults to contiguous_eval. This provides optimal memory locality for most inference workloads without requiring backend-specific optimizations.

How does MTPLX handle invalid detach modes?

The _normalize_detach_mode function validates all mode strings against the SUPPORTED_DETACH_MODES tuple. If a user supplies an unsupported value, the function raises a ValueError with a descriptive message listing the valid options, preventing silent failures during cache operations.

When should I use metal_copy_leaf versus contiguous_eval?

Use metal_copy_leaf when running on Apple Silicon (M1/M2/M3) devices where Metal Performance Shaders can accelerate memory operations. For CUDA, CPU, or other non-Metal environments, contiguous_eval provides identical functionality without backend detection overhead, as metal_copy_leaf automatically falls back to standard contiguous copying on incompatible hardware.

Are detach modes backend-specific?

Only metal_copy_leaf contains backend-specific logic targeting Apple's Metal framework. The other three modes (eval_only, contiguous_eval, and selected_slice_contiguous_eval) operate agnostically across CPU, CUDA, and other accelerator types, making them portable across different deployment environments.

Have a question about this repo?

These articles cover the highlights, but your codebase questions are specific. Give your agent direct access to the source. Share this with your agent to get started:

Share the following with your agent to get started:
curl -s "https://instagit.com/install.md"

Works with
Claude Codex Cursor VS Code OpenClaw Any MCP Client

Maintain an open-source project? Get it listed too →