# How MTPLX Manages Memory Using Request-Sized Paged KV Cache

> Discover how MTPLX efficiently manages memory with its request-sized paged KV cache. Learn about its fixed-size block allocation, token offset tracking, and precise memory slicing for optimized Metal attention inference.

- Repository: [Youssof Altoukhi/MTPLX](https://github.com/youssofal/MTPLX)
- Tags: internals
- Published: 2026-09-06

---

**MTPLX stores key/value pairs for each generation request in a paged KV cache that allocates memory in fixed-size blocks, tracks token offsets per request, and feeds exact memory slices to Metal attention kernels for efficient inference.**

The MTPLX inference engine implements a sophisticated memory management strategy centered around a request-sized paged KV cache that mirrors the layout expected by vLLM-Metal paged-attention kernels. Unlike static allocation schemes that reserve maximum context length upfront, this design grows memory block-by-block while maintaining contiguous logical token positions. This approach balances the memory efficiency of paged attention with simplified indexing semantics.

## Core Architecture of the Request-Sized Paged KV Cache

The foundation of MTPLX's memory management lies in its three-dimensional tensor structure that organizes KV data into manageable pages.

### Block-Based Tensor Layout

The cache stores key and value tensors using a specific dimensional ordering:

```

[num_blocks, block_size, num_kv_heads, head_dim]

```

Each **block** holds a fixed number of tokens defined by `block_size` (defaulting to 16). This block structure allows the system to allocate memory granularly rather than reserving vast contiguous regions for maximum theoretical sequence lengths.

In [`mtplx/cache_state.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/cache_state.py), the `VllmMetalPagedKVCache.__init__` method (lines 803-813) creates empty `key_cache` and `value_cache` tensors sized to `num_blocks × block_size × num_kv_heads × head_dim`. This pre-allocation establishes the initial memory pool without populating it with actual KV data.

### Pre-Allocation and Dynamic Growth Strategy

MTPLX employs a two-phase memory strategy: initial reservation followed by conditional expansion. The `_grow_to_capacity` method (lines 28-48 in [`mtplx/cache_state.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/cache_state.py)) handles dynamic enlargement when a request requires more tokens than currently allocated.

Growth behavior activates based on the `MTPLX_DYNAMIC_PAGED_KV` environment variable and respects the configured context-window limit. When enabled, the cache automatically allocates additional blocks rather than failing with out-of-memory errors during long-context generation.

```python

# Enable dynamic growth – the cache will automatically allocate new blocks

import os
os.environ["MTPLX_DYNAMIC_PAGED_KV"] = "1"

# Create a paged KV cache for a request (default 16-token blocks, 1024 blocks)

from mtplx.cache_state import VllmMetalPagedKVCache

cache = VllmMetalPagedKVCache(block_size=16, num_blocks=1024)

# After a few updates, the cache may have grown beyond its initial `num_blocks`

print("Allocated blocks:", cache.allocated_blocks)
print("Growth events:", cache.grow_events)

```

## Memory Management Mechanics

Beyond simple allocation, the cache implements sophisticated tracking mechanisms to manage active context windows and quantization formats.

### Request-Sized Offset Tracking

The cache maintains precise position awareness through `self.offset`, which tracks the current token count for the active request. During each generation step, the `update_without_fetch` method (lines 89-91) increments this offset, while subsequent attention calls begin computation from this recorded position.

This request-sized offset ensures that attention kernels read only the relevant portion of the KV cache, avoiding unnecessary memory bandwidth consumption on uninitialized or stale entries. The logical token positions remain contiguous from zero, keeping the block table trivial while still exercising native paged-read paths.

### Sliding-Window Memory Control

To prevent unbounded memory consumption during extremely long conversations, MTPLX implements sliding-window constraints. The `_active_attention_arrays` method (lines 84-88) trims the visible KV cache portion to a configured window size, effectively masking older tokens without deallocating the underlying physical blocks.

This mechanism allows the system to maintain constant memory usage regardless of total conversation length, provided the active window fits within the allocated block capacity.

### Quantization-Aware Memory Layout

When **TurboQuant** or **KV-Quant** compression is enabled, MTPLX allocates supplementary buffers including `_turboquant_v_centroids` and `_quant_bank` (lines 47-66). These structures hold compressed representations of the KV data, but critically, the logical layout and offset handling remain identical to the full-precision path.

This design ensures that quantization introduces no additional complexity to the paging logic—the same offset calculations and block indices apply whether storing FP16, INT8, or custom quantized formats.

## Integration with Metal Attention Kernels

The request-sized paged KV cache interfaces directly with Apple's Metal performance shaders through carefully orchestrated memory slicing.

### Paged Attention Routing Logic

The `paged_attention` method (lines 219-258 in [`mtplx/cache_state.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/cache_state.py)) implements intelligent dispatch logic that selects between primitive, partitioned, or fallback implementations based on:

- Environment variable configuration
- Current token offset position
- Sequence length requirements

This routing ensures optimal kernel selection for different batch sizes and sequence lengths while maintaining the request-sized memory view.

### Kernel Interface and Memory Slicing

When invoking native Metal kernels, MTPLX passes the current offset explicitly, ensuring the GPU reads only allocated pages rather than the entire reserved tensor. The offset parameter effectively truncates the logical view of the cache without requiring expensive memory copies or reallocation.

```python

# Add new KV entries for a generation step

cache.update_without_fetch(keys=new_keys, values=new_values)

# Perform attention for the next token batch

out = cache.paged_attention(
    queries,               # [batch, heads, q_len, dim]

    scale=1.0 / (head_dim ** 0.5),
    sliding_window=2048,   # optional sliding-window trim

)

print("Current offset:", cache.offset)          # total tokens stored so far

print("Cache capacity (tokens):", cache.capacity)

```

## Monitoring and Diagnostics

MTPLX exposes detailed memory usage metrics through counter variables maintained within the cache state object. The system tracks `self.paged_attention_calls`, `self.offset`, and `self.capacity` (lines 28-34), providing precise visibility into:

- Cumulative attention operations per request
- Current token consumption relative to allocation
- Growth events and capacity utilization

These statistics integrate with the diagnostic infrastructure in [`mtplx/kpi/runtime_kpis.py`](https://github.com/youssofal/MTPLX/blob/main/mtplx/kpi/runtime_kpis.py), enabling runtime verification through utilities like `exact_paged_attention_env`.

## Summary

- **Block-structured allocation**: MTPLX organizes KV cache as `[num_blocks, block_size, num_kv_heads, head_dim]` tensors, allocating memory in 16-token chunks rather than maximum sequence length.
- **Dynamic growth capability**: The `_grow_to_capacity` method expands block allocations on-demand when `MTPLX_DYNAMIC_PAGED_KV` is enabled, respecting context-window limits.
- **Request-scoped offset tracking**: `self.offset` maintains contiguous logical positions despite physical paging, allowing Metal kernels to read precise memory slices without table indirection.
- **Memory protection mechanisms**: Sliding-window support via `_active_attention_arrays` prevents unbounded growth, while quantization buffers maintain layout compatibility.
- **Kernel integration**: The `paged_attention` method routes to optimized Metal implementations while passing offset parameters that restrict reads to active pages.

## Frequently Asked Questions

### How does MTPLX prevent memory waste with variable-length requests?

MTPLX allocates the request-sized paged KV cache with an initial modest block count and expands only when necessary through the `_grow_to_capacity` mechanism. This avoids the common pattern of pre-allocating maximum context length for every request, instead matching physical memory allocation to actual token generation needs.

### What is the relationship between logical token positions and physical block locations?

The cache maintains contiguous logical positions from zero through `self.offset` while storing data in non-contiguous physical blocks. The block table remains trivial because logical indices map directly to physical locations without complex remapping, as implemented in `update_without_fetch` and consumed by the Metal kernels.

### How does MTPLX handle KV cache quantization without breaking the paging structure?

When TurboQuant or KV-Quant is active, MTPLX allocates additional centroid and quantization bank buffers (`_turboquant_v_centroids`, `_quant_bank`) but preserves the identical logical layout and offset handling. This allows the same paging infrastructure to serve both compressed and uncompressed KV representations.

### Can the sliding window be adjusted dynamically during inference?

The `paged_attention` method accepts a `sliding_window` parameter that trims the visible KV cache through `_active_attention_arrays`. While the underlying physical allocation remains fixed, this runtime parameter dynamically restricts which cached tokens participate in attention computation, effectively bounding memory bandwidth and computation regardless of total history length.